Member of Technical Staff - Research Scientist
- Location
- San Francisco
- Workplace
- On-site
- Compensation
- $150k – $200k + equity
- Visa
- Visa Sponsorship Available
About this role
In this role, you'll be at the forefront of developing benchmarks and evaluation methodologies for large language models. You'll help shape how the most advanced AI systems are tested and validated before being deployed in enterprise environments.
Key Responsibilities
Evaluate new AI models as they're released (such as DeepSeek, Gemini, etc.)
Create new benchmarks from scratch, including hiring labelers, constructing datasets, and writing white papers
Improve underlying methods for auto-evaluation of generated text
Work closely with engineering to implement and scale evaluation methodologies
Collaborate with top AI labs and enterprise customers to understand evaluation needs
What happens next
Skip the application pile. I get you in front of the people who decide.
Confirm the fit
A few questions to make sure this role is the right shape for you. Two minutes.
I pitch you to the company
I write the intro, send it to the founder, and handle the back-and-forth.
A meeting lands on your calendar
When the company wants to meet, I get the call on your calendar. You just show up.
Know someone who'd be great for this?
