Alloy

Lead Site Reliability Engineer

$179K–$226KFull-time · New York City
✓ Verified live on the employer's own system · added 125 days ago
Save search
Senior · 10+ yrs exp

Requirements

Experience: 10+ years

Skills & tools

Project ManagementManagementDevopsOperationsSecurityProgrammingTroubleshootingDistributed Systems
Apply on company site ↗ See your fit → free

Full job description

Alloy's Infrastructure Team is a small team (6 engineers) responsible for a large and growing infrastructure footprint: 15+ Kubernetes clusters, 100+ databases, dozens of services, and complex data organization.

Our challenge isn't just scale-it's making that scale reliable, secure, and operable with less manual work.

We're looking for engineers who enjoy turning complex, fragile systems into automated, self-service platforms with strong safety guarantees.

Reporting to the Engineering Manager of Infrastructure, you'll:

- Design and build systems to automate infrastructure management at scale (provisioning, upgrades, migrations)

- Reduce operational toil by turning manual processes into reliable, repeatable workflows

- Build internal tooling and platforms that enable safe self-service changes for other engineers

- Improve the reliability and resilience of our infrastructure (Kubernetes, databases, services)

- Implement and evolve systems for deploying and running applications in Kubernetes

- Contribute to architecture decisions across infrastructure, reliability, and security

- Participate in on-call rotations-but focus on building systems that prevent incidents, not just respond to them

- 10+ years of experience in infrastructure, SRE, or software engineering roles

- Strong software engineering skills-you build systems, not just scripts

- Experience managing production infrastructure at scale (cloud + containerized systems)

- Experience running and troubleshooting distributed systems (Docker/Kubernetes)

- Experience with observability and debugging tools (Datadog, CloudWatch, ELK/EFK, etc.)

- Proficiency in at least one programming language (Python, Go, JavaScript, etc.)

- Experience participating in on-call rotations and improving systems based on incidents

- Think in terms of systems, failure modes, and long-term scalability

- Care about building infrastructure that other engineers can use safely and confidently

- Enjoy working in a small team with high ownership and impact

- Experience building internal platforms or developer tooling

- Background in distributed systems or large-scale data systems

More jobs at Alloy

Similar jobs near New York City

Tell me when more Site Reliability Engineer II jobs post near New York, NY We re-check every listing against the employer’s own board — no résumé needed.

Search Lead Site Reliability Engineer jobs near New York City → Browse all live jobs

This posting was published by Alloy on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.