We're hiring a Forward Deployed Machine Learning Engineer in our Benchmarks and Evaluations vertical. You'll be the first MLE dedicated to this vertical and will work directly with the GM and our researchers to scale Protege's position as a renowned leader in the space.
At Protege, we believe that real world data is one of the largest bottlenecks to AI progress. Our data and data expertise position us to be neutral arbiters for the market, helping model builders understand the current performance of their models, identify what data will improve performance, and show that improvement over time.
Benchmarks and evaluations power that cycle. As an early engineer in the Benchmarks and Evaluations vertical, this role is an opportunity to help build the technical foundation for a critical area that greatly benefits current and future customers.
- Partner with the GM and early customers to define what constitutes strong evals in different domains
- Work with Protege researchers to design and build benchmarks
- Build the standards on how different modalities should be processed
- Build the backend the vertical runs on which includes data pipelines, execution environments, storage, and orchestration
- Stand up sandboxed environments for agentic evals, where models need tools, code execution, or multi-step tasks
- Find repeatable eval patterns, infrastructure gaps, and product opportunities from live engagements
- Partner with DataLab (our research team) on domain-specific data and research questions
- Build an understanding of the evals landscape, the GM's strategy, and customer demand
- Build an understanding of what our platform and data partners can support today, and where the gap is for eval building
- Identify the largest technical bets and ship multiple iterations of the eval infrastructure
- Own the engineering portion of customer engagements end to end
- 4+ years of engineering experience
- Hands-on ML work evaluating models
- Have previously owned backend and infrastructure
- High ambiguity tolerance and bias to action
- Comfort working with urgency to meet the pace and volume of the market demands
- Strong written communication
- Prior experience building benchmarks, evals, or human data pipelines for LLMs
- Time at a frontier lab, an eval-focused team, or a research org
- Founding or early engineer experience at a fast-moving startup
- Familiarity with agentic systems, RL environments, code-execution sandboxes, TEE/TREs
We act with integrity and do the right thing - especially when it's hard and no one is watching.
We are resourceful, resilient builders who solve hard problems and push through obstacles.
Velocity matters. We move with urgency, learn quickly, and continuously improve as individuals and as a company.
We communicate directly and respectfully, building trust through honest feedback and genuine care for one another.
We win as one team. Collaboration, accountability, and shared ownership drive our success.
Own the Outcome. Hone the Craft.
We take pride in our work, sweat the details, and continuously raise the bar for excellence.
Search Forward Deployed Machine Learning Engineer jobs near Remote → Browse all live jobs
This posting was published by Protege on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.