Altruist

Staff Back End Engineer, Evals - Hazel AI

$275K–$325KFull-time · San Francisco, CA
✓ Verified live on the employer's own system · added 75 days ago
Save search
Senior · 8+ yrs exp

Requirements

Experience: 8+ years

Skills & tools

DevopsMachine LearningOperationsSQLWarehouseCommunicationsTeam Leadership
Apply on company site ↗ See your fit → free

Full job description

Hazel.ai is building the AI engine for wealth management that unlocks 10x growth, efficiency and value for financial advisors and their clients in a regulated industry. Since its launch last September, Hazel has organically and rapidly grown its user base.

Hazel is a part of Altruist's broader mission to make financial advice better, more affordable, and accessible to all.

This role is hybrid, with four in-office days per week at our San Francisco FiDi location.

Architect our evaluation platform from first principles - the observability, scoring, golden datasets, verification agents, and CI/CD integration that define standards of quality. You'll work shoulder-to-shoulder with backend engineers, product managers, and a growing bench of subject matter experts, including practicing CFPs, CPAs, and tax planners, to translate fiduciary-grade requirements into automated quality signals.

- Design and build Hazel's evals platform end-to-end - online scoring, offline benchmarks, regression suites, LLM-as-judge pipelines, and human-in-the-loop review workflows across every Hazel surface.

- Build production observability and monitoring for AI quality: hallucination rates, factual accuracy, refusal behavior, latency, cost, and domain-specific quality signals across tax planning, financial planning, investment analysis, and operational AI workflows.

- Architect data curation pipelines that turn real advisor interactions into evaluation datasets - with rigorous sampling strategies, labeling protocols, dataset versioning, and the privacy and consent controls required for regulated finance.

- Build and steward Hazel's golden datasets in close partnership with SMEs and a network of practicing advisors, CFPs, and tax professionals - translating their tacit expertise into precise, measurable eval criteria.

- Develop LLM verification agents that catch hallucinations, computational errors, and compliance violations before they ever reach an advisor or client.

- Integrate evals into our deployment pipeline so that every prompt change, model swap, harness modification, or RAG pipeline tweak runs against regression and acceptance criteria before shipping - making evals a first-class deployment gate, not a quarterly audit.

- Partner with the team building Hazel's model-agnostic orchestration harness to evaluate cross-model and cross-provider performance, surface tradeoffs, and inform routing decisions across Anthropic, OpenAI, and self-hosted models.

- Define quality SLOs for each Hazel surface and build alerting that catches regressions in production before our customers do - especially for high-stakes flows like tax and financial planning.

- Establish Hazel's eval methodology as a defensible competitive advantage - infrastructure good enough that model upgrades from frontier labs become accelerants for us, not threats.

- 8+ years of engineering experience, with at least 2 years focused on evaluation infrastructure, model quality, fine-tuning, or ML platform work for production systems.

- Deep familiarity with evaluation and scoring methodologies for modern AI systems - RAG evaluation, document processing, fine-tuned model assessment, agentic and tool-use system evaluation, LLM-as-judge frameworks, and human evaluation protocols.

- Experience designing and curating golden datasets - sampling strategies, inter-rater agreement, dataset versioning, and managing the long tail of edge cases.

- Comfort working across the stack - data engineering (SQL, dbt, warehouses), backend integration (APIs, async pipelines, queues), and observability tooling.

- Strong communication skills. You can translate fuzzy domain requirements from advisors and SMEs into precise, measurable, automatable eval criteria - and explain quality tradeoffs clearly to engineers, product managers, and leadership.

- A bias toward shipping. You believe great evals enable speed, not just safety, and you build tools that engineers actually want to use.

- Prior experience at an applied AI company building evals, model quality, or applied research infrastructure.

- Experience evaluating multi-step agentic workflows, tool-use systems, or RAG pipelines in production.

- Familiarity with frameworks like Braintrust, Langfuse or similar - including a clear point of view on when to use which.

- Background in regulated industries (financial services, healthcare, legal) where accuracy, auditability, and the cost of a wrong answer are unusually high.

- Experience building human-in-the-loop labeling workflows, annotation tooling, or red-teaming programs.

- Domain knowledge of wealth management, tax planning, or financial planning - or genuine excitement to learn it deeply alongside our SME bench.

More jobs at Altruist

Similar jobs near San Francisco, CA

Tell me when more Senior Full Stack Engineer, Hazel AI jobs post near San Francisco, CA We re-check every listing against the employer’s own board — no résumé needed.

Search Staff Back End Engineer, Evals - Hazel AI jobs near San Francisco, CA → Browse all live jobs

This posting was published by Altruist on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.