Experience: 7+ years
- Design, build, and maintain secure, maintainable, self-serve core infrastructure that engineering teams can rely on and operate independently
- Lead the architecture and evolution of a modern ML training infrastructure — scalable, reproducible, and built for rapid experimentation
- Build and operate a modern model serving architecture with a focus on reliability, cost efficiency, and low latency
- Lead and own the low-latency voice interface and audio processing pipeline — a technically demanding, performance-sensitive system at the core of Sesame's product
- Build developer tooling, server infrastructure, and data infrastructure that is high leverage and low maintenance — the kind that makes other engineers faster without creating new dependencies on you
- Lead technical direction within your domain, bring others along through clear communication and well-reasoned proposals, and raise the engineering bar across the team
- A strong systems thinker who is equally comfortable leading technical direction and getting hands-on with implementation
- 7+ years of software engineering experience, with significant time in infrastructure, platform, or ML systems roles
- Hands-on reliability engineering experience — you have well-formed convictions about observability, monitoring, deployment systems, and loosely coupled architectures, and you've put them into practice at scale
- Proven track record of building and shipping services at scale, with all the operational complexity that comes with it
- Kubernetes — significant production experience building, operating, and scaling Kubernetes clusters
- Experience designing and shipping flexible domain models and APIs — you think carefully about boundaries, contracts, and long-term maintainability
- A default toward automation — you've consistently built efficiency gains through automation and have the track record to show it
- Strong communication skills — you can lead your own direction, write clearly about tradeoffs, and bring engineers and stakeholders along with you
We'd love to hear about experience in any of these areas — we don't expect any one person to have all of them:
- Infrastructure as Code at scale — significant IaC experience, preferably Terraform; CloudFormation, Pulumi, or Kubernetes-based approaches also welcome. Ideally you've led, architected, or contributed to a multi-stack, self-serve IaC system and understand the challenges of building infra that teams can own independently
- PyTorch experience, especially model optimisation for serving
- Experience building ML serving and/or training infrastructure (TorchServe, Seldon, KServe, Ray Serve, or similar)
- Experience building and leading large-scale distributed training and serving systems
- Data engineering — pipeline design, dataset management, or data platform experience
- Database design — complex schema design, query optimisation, and hard data modelling decisions across relational and non-relational stores
- Real-time communication systems — low-latency audio, video, or streaming infrastructure
Search Staff Software Engineer, Backend Infrastructure jobs near San Francisco → Browse all live jobs
This posting was published by Sesame on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.