Member of Technical Staff - Model Optimization and Inference (Experienced)

Location
Seattle
Workplace
On-site
Compensation
$250k – $350k + equity
Visa
Visa Sponsorship Available

About this role

What we're looking for?

We are seeking an experienced ML Infrastructure/Systems Engineer with at least 2 years of full-time experience in building and maintaining production-level ML systems. You should be comfortable designing scalable infrastructure from scratch, making informed design decisions by comparing various technologies, and have a track record of optimizing systems for latency, throughput, and cost. We're looking for someone with a broad understanding of the ML infra space, including inference infrastructure, real-time video streaming, and data engineering, who can own complex projects and debug distributed systems. Bonus points if you have experience with video or audio models and low-level optimization techniques like CUDA kernels.

What you'll do:

Own end-to-end inference optimization across our model stack — LLMs, audio models, and diffusion-based components

Implement and tune KV cache strategies for long-context conversations, including eviction policies, compression, and memory-efficient attention

Evaluate, deploy, and extend inference serving frameworks (vLLM, SGLang, TensorRT-LLM, etc.) for our specific workloads

Profile and benchmark end-to-end latency and throughput; identify and systematically eliminate bottlenecks

Build internal tooling that makes optimization work faster and more rigorous — profiling viewers, end-to-end inference test harnesses, and other infrastructure that helps the team move quickly

Accelerate diffusion model inference — consistency models, step distillation, caching strategies, and custom kernel optimizations

Apply and develop quantization techniques (INT8, INT4, GPTQ, AWQ, and beyond) to reduce memory footprint and increase throughput without meaningfully degrading quality

Work closely with research and infrastructure to ensure new models ship with optimized serving from day one

What happens next

Skip the application pile. I get you in front of the people who decide.

Confirm the fit

A few questions to make sure this role is the right shape for you. Two minutes.

I pitch you to the company

I write the intro, send it to the founder, and handle the back-and-forth.

A meeting lands on your calendar

When the company wants to meet, I get the call on your calendar. You just show up.

Know someone who'd be great for this?