Experience: 3+ years
we are looking for people who want to build it with us. If you have already changed how you work because of AI - or you are ready to - and you care more about shipping something great than following a prescribed process, we should talk.
Babylist's Platform team is the foundation every engineering team builds on - and this role is at the center of keeping it reliable, fast, and scalable. As a Staff SRE, you'll own the infrastructure and reliability practices that support 9 million+ users and the engineers who build for them. Babylist started as an e-commerce and registry platform, and we're actively growing beyond that - into health, media, mobile, and new product surfaces that don't exist yet.
The Platform team is the foundation that makes all of it possible. This isn't a maintenance role - you'll be actively evolving how we build and operate AWS infrastructure, CI systems, and developer tooling. You'll work cross-functionally across all of Babylist Engineering, which means your decisions have wide leverage.
- Deep hands-on Terraform expertise - you own IaC, not just contribute to it
- Proven AWS experience at scale - EKS, RDS, cloud networking, DNS, CDNs, load balancers - you know the gotchas
- Experienced operating Kubernetes in production - you've debugged the hard stuff, not just deployed the easy stuff
- Comfortable designing and improving CI/CD systems - CircleCI, GitHub Actions, or similar; you care about developer velocity, not just pipeline uptime
- Strong observability instincts - Datadog, Sentry, PagerDuty, Cronitor - you build alerting that's actionable, not noisy
- Experienced with on-call and incident management - you've run the post-mortems and actually changed things afterward
- Comfortable supporting developers across local, staging, and production - you're a resource, not a gatekeeper
- You naturally reach for AI in your work - at Babylist, every team uses AI daily. You're already using it to move faster and improve your output, and you stay curious about what's coming next.
- Infrastructure ownership - manage and evolve our AWS environment using Terraform, keeping EKS clusters, databases, and core services current and performant
- CI/CD reliability - own the speed and reliability of our CI systems for the full Engineering org - every deploy starts here
- Developer support - be the person engineers turn to when environments break; unblock them fast across local, staging, and production
- Monitoring & alerting standards - establish and socialize best practices so the right people get paged for the right reasons
- Incident response - lead or support incident response, drive post-incident reviews, and close the loop so the same thing doesn't happen twice
- Platform strategy - contribute to architectural decisions that shape how Babylist's infrastructure evolves over the next several years
- Platform is the team every engineering team depends on - your work has outsized leverage across the entire product org, not just one area
- The infrastructure is solid but actively evolving - you're not inheriting chaos, you're shaping what comes next
- This is a staff-level role with real cross-team visibility - you'll influence how Babylist engineers build and ship, not just keep the lights on
- You'll work on systems that support millions of families at a high-stakes life moment - the scale is real and the product context makes the reliability work matter
$226,673 to $271,991 + target 20% annual bonus and competitive equity
Search Staff Engineer, Site Reliability jobs → Browse all live jobs
This posting was published by Babylist on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.