Cisco

Senior Site Reliability Engineer, Production Engineer - ThousandEyes

$165KFull-time · San Francisco, CA
✓ Verified live on the employer's own system · added 9 days ago
Save search
Mid-level · 5+ yrs exp

Requirements

Experience: 5+ years

Skills & tools

OperationsCloud PlatformsDevopsProgrammingPythonSecurityLinuxDistributed Systems
Apply on company site ↗ See your fit → free

Full job description

We are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background in SaaS and operations. You will design and manage large-scale, highly available distributed systems in the cloud, collaborating directly with application development teams to enhance the reliability, performance, and security of our platform.

- Collaborate with software engineers to optimize architecture and services for availability, latency, performance, and reliability using cloud-native tools. - Design and implement scalable operations tooling to support platform growth and scaling across multiple regions. - Design, deploy, and maintain AWS cloud-native services that are elastic and resilient to failure. - Participate in and improve our 24x7 incident response and on-call rotation. - Use and expand our existing CNCF solutions like Kubernetes, Service Mesh, Prometheus, OpenTelemetry, and ArgoCD to increase platform reliability. - Automate production operations to provide guardrails and continuous platform operation. - Develop automation solutions for scalable service and platform operations, including deployment, scale testing, graceful failure, and chaos testing. - Stay updated on industry best practices for scalability and reliability to improve the scalability of the ThousandEyes platform. - Identify and provide solutions to common obstacles hindering operational excellence across engineering teams. - Generalize and standardize solutions and processes to enable repeated success across our microservice-based multi-region platform. - Play a key role in the ThousandEyes platform by leveraging scale testing, additional environments, and working with application teams to improve system reliability. - Manage a rapidly growing infrastructure capable of handling substantial daily data volumes, emphasizing operations/infrastructure/everything as code.

- 5+ years of experience in a related role - Proficiency in software development with languages such as Python or Go - Shown ability to build and implement scalable, well-tested, and security-focused solutions that integrate security protocols throughout the development and deployment lifecycle - Strong understanding of Unix/Linux systems, including kernel, system libraries, file systems, and client-server protocols - Knowledge of Site Reliability principles: Incident Response, Change Management, Distributed Systems, Deployment Strategies, and SLOs

- Familiarity with procedures for operating a large-scale, highly available enterprise platform - Excellent communication and documentation skills - Strong sense of ownership, drive, and attention to detail - Expert-level knowledge of Kubernetes and its ecosystem - In-depth knowledge of cloud providers, preferably AWS

The starting salary range posted for this position is $165,000.00 to $241,400.00 and reflects the projected salary range for new hires in this position in U.S. and/or Canada locations, not including incentive compensation*, equity, or benefits.

Non-Metro New York state & Washington state: $146,700.00 - $247,000.00

More jobs at Cisco

Similar jobs near San Francisco, CA

Tell me when more Production Engineer jobs post near San Francisco, CA We re-check every listing against the employer’s own board — no résumé needed.

Search Senior Site Reliability Engineer, Production Engineer - ThousandEyes jobs near San Francisco, CA → Browse all live jobs

This posting was published by Cisco on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.