Infrastructure Engineer — Managed Inference

Location
Tel Aviv-Yafo, San Francisco +1
Workplace
On-site
Compensation
$200k – $260k + equity
Visa
Visa Sponsorship Available

About this role

What we're looking for:

We need someone with 5+ years of infrastructure engineering experience who has deep hands-on expertise operating production Kubernetes at scale and working with LLM inference serving systems. You should be comfortable debugging NVIDIA GPU systems end-to-end (drivers, CUDA, NCCL, network fabric) and have a track record of building highly available, multi-cloud infrastructure for latency-sensitive AI workloads. Bonus points if you've deployed across heterogeneous accelerator types (NVIDIA GPUs, AWS Trainium, Google TPUs) or contributed to open-source inference frameworks like vLLM or SGLang.

What you'll do:

Design, operate, and scale highly available Kubernetes clusters across AWS, GCP, and specialized GPU cloud providers to serve global inference demand at 1, 000+ tokens per second

Take ownership of the production serving layer — debug NVIDIA systems end-to-end including drivers, CUDA, NCCL, node health, and network fabric

Work closely with the model optimization team to ensure kernel-level speed improvements survive contact with real traffic, real hardware, and real customers

Extend deployment tooling so models ship consistently across NVIDIA, Trainium, and TPU hosts through a unified workflow

Build observability infrastructure that keeps the platform honest: time-to-first-token, inter-token latency, throughput, and availability — measured per model, per chip, and per region

Own reliability engineering including alerting, automated failover, self-healing infrastructure, and intelligent traffic routing across models, chips, and regions

What happens next

Skip the application pile. I get you in front of the people who decide.

Confirm the fit

A few questions to make sure this role is the right shape for you. Two minutes.

I pitch you to the company

I write the intro, send it to the founder, and handle the back-and-forth.

A meeting lands on your calendar

When the company wants to meet, I get the call on your calendar. You just show up.

Know someone who'd be great for this?