About the role

We're partnering with a fast-growing (Series B) AI company that builds thedatasets and evaluation systemsfrontier AI labs use to train their models. You'd be working hand-in-hand with top research teams designing the tasks models practise on, and the scoring that decides whether they're really getting smarter. What you'd own: Design data that exposes where models fail across finance, code & enterprise workflows Build reward signals and evaluation rubrics for RLHF / RLVR training pipelines Develop frameworks to measure dataset quality and its real impact on model performance Turn ambiguous research goals into concrete, shippable systems You must have: Hands-on with modern ML/LLM workflows training, fine-tuning or evaluation Sharp instincts for data quality, edge cases and precision/recall trade-offs1–4 years shipping technical work (RLHF/RLVR or eval experience) Previous experience as a startup founder or co-founder is advantageous The details: $250 base + significant profit share + equity This is a rare chance to have direct, measurable impact on frontier AI on a small team where your work ships straight into the models defining the field.

Matching similar jobs

JOB OVERVIEW

Experience level

Senior

Location

Millbrae, CA

Occupation

Software Developers

Industry

Custom Computer Programming Services

Posted

2 days ago

Tired of running searches?

Rank the roles you'd take once, and matches like these arrive on their own.

CREATE PROFILE