Member of Technical Staff - Research Scientist

Location
San Francisco
Workplace
On-site
Compensation
$150k – $200k + equity
Visa
Visa Sponsorship Available

About this role

In this role, you'll be at the forefront of developing benchmarks and evaluation methodologies for large language models. You'll help shape how the most advanced AI systems are tested and validated before being deployed in enterprise environments.

Key Responsibilities

Evaluate new AI models as they're released (such as DeepSeek, Gemini, etc.)

Create new benchmarks from scratch, including hiring labelers, constructing datasets, and writing white papers

Improve underlying methods for auto-evaluation of generated text

Work closely with engineering to implement and scale evaluation methodologies

Collaborate with top AI labs and enterprise customers to understand evaluation needs

What happens next

Skip the application pile. I get you in front of the people who decide.

Confirm the fit

A few questions to make sure this role is the right shape for you. Two minutes.

I pitch you to the company

I write the intro, send it to the founder, and handle the back-and-forth.

A meeting lands on your calendar

When the company wants to meet, I get the call on your calendar. You just show up.

Know someone who'd be great for this?