Experience: 8+ years
Hazel.ai is building the AI engine for wealth management that unlocks 10x growth, efficiency and value for financial advisors and their clients in a regulated industry. Since its launch last September, Hazel has organically and rapidly grown its user base.
Hazel is a part of Altruist's broader mission to make financial advice better, more affordable, and accessible to all.
This role is hybrid, with four in-office days per week at our San Francisco FiDi location.
Architect our evaluation platform from first principles - the observability, scoring, golden datasets, verification agents, and CI/CD integration that define standards of quality. You'll work shoulder-to-shoulder with backend engineers, product managers, and a growing bench of subject matter experts, including practicing CFPs, CPAs, and tax planners, to translate fiduciary-grade requirements into automated quality signals.
- Design and build Hazel's evals platform end-to-end - online scoring, offline benchmarks, regression suites, LLM-as-judge pipelines, and human-in-the-loop review workflows across every Hazel surface.
- Build production observability and monitoring for AI quality: hallucination rates, factual accuracy, refusal behavior, latency, cost, and domain-specific quality signals across tax planning, financial planning, investment analysis, and operational AI workflows.
- Architect data curation pipelines that turn real advisor interactions into evaluation datasets - with rigorous sampling strategies, labeling protocols, dataset versioning, and the privacy and consent controls required for regulated finance.
- Build and steward Hazel's golden datasets in close partnership with SMEs and a network of practicing advisors, CFPs, and tax professionals - translating their tacit expertise into precise, measurable eval criteria.
- Develop LLM verification agents that catch hallucinations, computational errors, and compliance violations before they ever reach an advisor or client.
- Integrate evals into our deployment pipeline so that every prompt change, model swap, harness modification, or RAG pipeline tweak runs against regression and acceptance criteria before shipping - making evals a first-class deployment gate, not a quarterly audit.
- Partner with the team building Hazel's model-agnostic orchestration harness to evaluate cross-model and cross-provider performance, surface tradeoffs, and inform routing decisions across Anthropic, OpenAI, and self-hosted models.
- Define quality SLOs for each Hazel surface and build alerting that catches regressions in production before our customers do - especially for high-stakes flows like tax and financial planning.
- Establish Hazel's eval methodology as a defensible competitive advantage - infrastructure good enough that model upgrades from frontier labs become accelerants for us, not threats.
- 8+ years of engineering experience, with at least 2 years focused on evaluation infrastructure, model quality, fine-tuning, or ML platform work for production systems.
- Deep familiarity with evaluation and scoring methodologies for modern AI systems - RAG evaluation, document processing, fine-tuned model assessment, agentic and tool-use system evaluation, LLM-as-judge frameworks, and human evaluation protocols.
- Experience designing and curating golden datasets - sampling strategies, inter-rater agreement, dataset versioning, and managing the long tail of edge cases.
- Comfort working across the stack - data engineering (SQL, dbt, warehouses), backend integration (APIs, async pipelines, queues), and observability tooling.
- Strong communication skills. You can translate fuzzy domain requirements from advisors and SMEs into precise, measurable, automatable eval criteria - and explain quality tradeoffs clearly to engineers, product managers, and leadership.
- A bias toward shipping. You believe great evals enable speed, not just safety, and you build tools that engineers actually want to use.
- Prior experience at an applied AI company building evals, model quality, or applied research infrastructure.
- Experience evaluating multi-step agentic workflows, tool-use systems, or RAG pipelines in production.
- Familiarity with frameworks like Braintrust, Langfuse or similar - including a clear point of view on when to use which.
- Background in regulated industries (financial services, healthcare, legal) where accuracy, auditability, and the cost of a wrong answer are unusually high.
- Experience building human-in-the-loop labeling workflows, annotation tooling, or red-teaming programs.
- Domain knowledge of wealth management, tax planning, or financial planning - or genuine excitement to learn it deeply alongside our SME bench.
Search Staff Back End Engineer, Evals - Hazel AI jobs near San Francisco, CA → Browse all live jobs
This posting was published by Altruist on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.