Experience: 10+ years
Alloy's Infrastructure Team is a small team (6 engineers) responsible for a large and growing infrastructure footprint: 15+ Kubernetes clusters, 100+ databases, dozens of services, and complex data organization.
Our challenge isn't just scale-it's making that scale reliable, secure, and operable with less manual work.
We're looking for engineers who enjoy turning complex, fragile systems into automated, self-service platforms with strong safety guarantees.
- Design and build systems to automate infrastructure management at scale (provisioning, upgrades, migrations)
- Reduce operational toil by turning manual processes into reliable, repeatable workflows
- Build internal tooling and platforms that enable safe self-service changes for other engineers
- Improve the reliability and resilience of our infrastructure (Kubernetes, databases, services)
- Implement and evolve systems for deploying and running applications in Kubernetes
- Contribute to architecture decisions across infrastructure, reliability, and security
- Participate in on-call rotations-but focus on building systems that prevent incidents, not just respond to them
- 10+ years of experience in infrastructure, SRE, or software engineering roles
- Strong software engineering skills-you build systems, not just scripts
- Experience managing production infrastructure at scale (cloud + containerized systems)
- Experience running and troubleshooting distributed systems (Docker/Kubernetes)
- Experience with observability and debugging tools (Datadog, CloudWatch, ELK/EFK, etc.)
- Proficiency in at least one programming language (Python, Go, JavaScript, etc.)
- Experience participating in on-call rotations and improving systems based on incidents
- Think in terms of systems, failure modes, and long-term scalability
- Care about building infrastructure that other engineers can use safely and confidently
- Enjoy working in a small team with high ownership and impact
- Experience building internal platforms or developer tooling
- Background in distributed systems or large-scale data systems
Search Lead Site Reliability Engineer jobs near New York City → Browse all live jobs
This posting was published by Alloy on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.