Experience: 5+ years
We are looking for a Staff Site Reliability Engineer to join our team. This is a hybrid role going into the San Jose, CA office 3 days a week, reporting to the Chief Architect, REIS in the Cloud Infrastructure & Operations department. You are an SRE with proven experience in Linux/UNIX System Administration and hands-on expertise in building infrastructure and managing platforms like Kubernetes using automation tools and high security standards.
In this role, you will troubleshoot complex Linux networking and security issues, apply a strong understanding of firewall technologies, and manage access across various systems, platforms, and applications.
Create and maintain highly scalable solutions based on KVM Linux, Kubernetes, and Public Cloud Providers
Analyze and troubleshoot systems performance and complex issues across operating systems and applications
Maintain platform security and observability through nftables and comprehensive monitoring
Manage and deploy systems and software in diverse environments, including Kubernetes and Openstack
Write and maintain custom tools using Python, Golang, and BASH
You thrive in ambiguity. You're comfortable building the path as you walk it. You thrive in a dynamic environment, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful.
You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome.
True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution.
You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact.
You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback—knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust.
You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose.
Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain
5+ years of experience in a Linux/UNIX System Administration or SysAdmin role with a strong understanding of web security, protocols, and SSH
Technical security aptitude including PGP, SSH, PKI, and Multi-factor authentication
Mastery of network fundamentals including DHCP, ARP, subnetting, routing, NAT, firewalls, and IPv4/IPv6
Experience in container orchestration services including Docker and Kubernetes, alongside automation tools like Ansible
Experience implementing AI-driven anomaly detection, intelligent log parsing, or AIOps tools to automate root-cause analysis and predictive auto-scaling across Linux and Kubernetes infrastructure
Strong Openstack and CEPH experience paired with advanced Kubernetes expertise
Extensive professional history in dedicated Linux System Administration roles along with direct experience with Secrets Management Systems such as Hashicorp Vault or similar technologies
Search Staff Site Reliability Engineer jobs near San Jose, CA → Browse all live jobs
This posting was published by Zscaler on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.