Coinbase

Senior Site Reliability Engineer, Core AI Infrastructure

$186K–$219KFull-time · Remote - USA
✓ Verified live on the employer's own system · added 63 days ago
Save search
Mid-level · 5+ yrs exp

Requirements

Experience: 5+ years

Skills & tools

OperationsTeam LeadershipCloud PlatformsRoot Cause AnalysisDevopsSecurityRecordkeepingPython
Apply on company site ↗ See your fit → free

Full job description

You'll join a high-performing team of engineers driving AI transformation at Coinbase as a Senior Site Reliability Engineer on the IT Operations team. This team builds and scales the infrastructure powering Coinbase's AI products, with direct exposure to senior leadership in a fast-paced, incubator-style environment. You'll own the reliability and automation of critical AI infrastructure, ensuring our systems are resilient, observable, and secure at scale.

  • Own the reliability, monitoring, and incident response lifecycle for AI infrastructure services, including on-call support for AWS deployment pipelines, root cause analysis, and blameless retros.
  • Build automation and tooling to streamline operational IT workflows, eliminate manual tasks, and improve deployment velocity across CI/CD frameworks and Kubernetes environments.
  • Partner with the Coinbase Infrastructure team to extend CI/CD frameworks supporting IT services and enterprise network platforms, and with Security and Compliance to integrate surveillance tooling into deployment pipelines.
  • Strengthen observability and documentation standards across IT engineering by defining metrics, implementing monitoring solutions, and maintaining technical documentation that sets a standard of excellence.
  • Develop full-stack applications that power internal AI products and infrastructure with Go or Python.
  • 5+ years of experience automating and supporting cloud infrastructure (AWS) and network environments, with hands-on use of infrastructure-as-code tools (Terraform, Ansible, Chef, Puppet, or Salt).
  • Proven experience deploying, managing, and troubleshooting containerized workloads using Docker and Kubernetes in production environments.
  • Proficiency in at least one scripting or programming language (Python, Bash, Ruby, or Go) and version control workflows using Git-based CI/CD pipelines.
  • Track record of leading incident response in environments with strict SLAs, including root cause analysis, blameless retros, and measurable reliability improvements.

More jobs at Coinbase

Similar jobs near Remote - USA

Tell me when more Senior Site Reliability Engineer, Fleet Management jobs post near Remote - USA We re-check every listing against the employer’s own board — no résumé needed.

Search Senior Site Reliability Engineer, Core AI Infrastructure jobs near Remote - USA → Browse all live jobs

This posting was published by Coinbase on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.