Ando

Research Engineer

Full-time · San Francisco
✓ Verified live on the employer's own system · added 23 days ago
Save search

Skills & tools

Customer ServiceResearch
Apply on company site ↗ See your fit → free

Full job description

You'll be one of the first members of a small research team, working close to both the data and the product.

- Build evaluation from real data - Mine production workspace data (carefully, with consent and redaction pipelines you'll help design) into benchmarks and labeled datasets. Design label schemas, run labeling with real inter-rater rigor, and build the harnesses that make expert judgment cheap to capture and hard to corrupt.

- Run experiments that ship - Initial work happens on offline workspace data; the destination is production systems used by every Ando customer. The distance between "the benchmark improved" and "the feature shipped" should be weeks and you'll own both ends.

- Decide what we build versus who we partner with - We won't do everything in-house. Part of the job is evaluating frontier vendors and research teams (eval infrastructure, observability, continual-learning tooling) and choosing who we build with.

- Publish when we have something real - We expect the team to publish as we make meaningful progress. But ideas are not Ando's moat; execution, product quality, and customer experience are. Research here is in service of customers first, and the publications will be better for it: they'll describe things that actually worked on real data.

- Strong applied research background , with depth in model evaluation, benchmarking, and/or failure analysis. You've built evals you trusted enough to make decisions with.

- Evidence over credentials. Work samples or code that demonstrate the skills: eval frameworks, benchmark suites, failure-analysis reports or tooling, labeling infrastructure. Show us something you built to find out whether a system actually worked.

- Strong technical communication. You can explain complex ideas simply and hold high-bandwidth, generative technical conversations with researchers and with our product team.

- Comfort with mess. Real workspace data is incomplete, ambiguous, and full of edge cases that break clean abstractions. You treat that as signal, not noise.

- Bonus: familiarity with simulation ( Park et al. ), human-in-the-loop evaluation ( Scale HIL leaderboard ), Cartridges and related context/memory-compression work, memory for multi-party long-running settings, or agent observability standards (setting up Langsmith or similar).

More jobs at Ando

Similar jobs near San Francisco

Tell me when more Research Engineer, Machine Learning (Reinforcement Learning) jobs post near San Francisco, CA +1 more We re-check every listing against the employer’s own board — no résumé needed.

Search Research Engineer jobs near San Francisco → Browse all live jobs

This posting was published by Ando on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.