Infrastructure Engineer — Managed Inference
- Location
- Tel Aviv-Yafo, San Francisco +1
- Workplace
- On-site
- Compensation
- $200k – $260k + equity
- Visa
- Visa Sponsorship Available
About this role
What we're looking for:
We need someone with 5+ years of infrastructure engineering experience who has deep hands-on expertise operating production Kubernetes at scale and working with LLM inference serving systems. You should be comfortable debugging NVIDIA GPU systems end-to-end (drivers, CUDA, NCCL, network fabric) and have a track record of building highly available, multi-cloud infrastructure for latency-sensitive AI workloads. Bonus points if you've deployed across heterogeneous accelerator types (NVIDIA GPUs, AWS Trainium, Google TPUs) or contributed to open-source inference frameworks like vLLM or SGLang.
What you'll do:
Design, operate, and scale highly available Kubernetes clusters across AWS, GCP, and specialized GPU cloud providers to serve global inference demand at 1, 000+ tokens per second
Take ownership of the production serving layer — debug NVIDIA systems end-to-end including drivers, CUDA, NCCL, node health, and network fabric
Work closely with the model optimization team to ensure kernel-level speed improvements survive contact with real traffic, real hardware, and real customers
Extend deployment tooling so models ship consistently across NVIDIA, Trainium, and TPU hosts through a unified workflow
Build observability infrastructure that keeps the platform honest: time-to-first-token, inter-token latency, throughput, and availability — measured per model, per chip, and per region
Own reliability engineering including alerting, automated failover, self-healing infrastructure, and intelligent traffic routing across models, chips, and regions
What happens next
Skip the application pile. I get you in front of the people who decide.
Confirm the fit
A few questions to make sure this role is the right shape for you. Two minutes.
I pitch you to the company
I write the intro, send it to the founder, and handle the back-and-forth.
A meeting lands on your calendar
When the company wants to meet, I get the call on your calendar. You just show up.
Know someone who'd be great for this?
