NVIDIA | NVIDIA

Principal Engineer, Cloud Site Reliability Engineering

Full-time · US, CA, Santa Clara
✓ Verified live on the employer's own system · added 3 days ago
Save search
Senior · 15+ yrs exp

Requirements

Education: Bachelor's degree or related field

Experience: 15+ years

Skills & tools

DevopsQuality AssuranceProgrammingCloud PlatformsData AnalysisElectricalJavaPython
Apply on company site ↗ See your fit → free

Full job description

What you'll be doing: Serve as an SRE Architect part of GPU Private Cloud team used by thousands of NVIDIANs globally for interactive development, centralized CI/CD, and QA testing. Evaluating, identifying and developing software solutions to optimize critical software development workflows across various organizations within NVIDIA.

Architecting, implementing, and supporting end-to-end CI/CD system using open-source and NVIDIA proprietary software. Customer (NVIDIA Internal development teams) onboarding to Private cloud infrastructure with a good discovery of the use case and available solutions within the cloud. Identify performance bottlenecks and optimize the speed and cost efficiency of AI development and testing systems.

Leading software development projects and technically direct a team of brilliant engineers and guide them to provide efficient and impactful solutions. Looking for problems within software systems and resolving the issues Craft and implement critical metrics using various analytics methods and dashboards. What we need to see: BS or MS in Electrical Engineering, Computer Science, or relevant field (or equivalent experience).

15+ years of systems software development including at least 1 year dedicated to developing/exploring AI. Experience of maintaining cloud infrastructure and highly available production environment. Strong programming and software development skills in JAVA, Python, Shell-script along with good understanding of distributed systems and REST APIs.

Experience in working with SQL/NoSQL database systems such as MySQL, Cassandra, MongoDB or Elasticsearch. Excellent knowledge and working experience with Docker containers and Virtual Machines. Good background of Cloud technologies like: OpenStack, Docker, Kubernetes, Chef/Puppet, Hadoop/Ceph/SwiftStack, LXC, Git, Perforce, JFrog, Kafka.

Ability to work across organizational boundaries effectively to improve alignment and productivity between teams in a multi-national, multi-time-zone corporate environment. Ways to stand out from the crowd: Depth in AI, Machine Learning and Deep Learning algorithms and techniques. Strong collaborative and interpersonal skills, with a consistent record of guiding and influencing others in dynamic environments.

Experience developing large-scale software systems using modular architecture under real-time performance requirements. Background in designing high-performance, scalable software systems with a strong focus on hardware cost optimization. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.

The base salary range is 272,000 USD - 431,250 USD. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until August 9, 2026.

This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer.

As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

NVIDIA pioneered accelerated computing. Today, our AI infrastructure powers global intelligence, transforming every industry. Learn more about NVIDIA.

More jobs at NVIDIA | NVIDIA

Similar jobs near US, CA, Santa Clara

Tell me when more Senior Software Engineering Manager, Reliability jobs post near Mountain View, CA We re-check every listing against the employer’s own board — no résumé needed.

Search Principal Engineer, Cloud Site Reliability Engineering jobs near US, CA, Santa Clara → Browse all live jobs

This posting was published by NVIDIA | NVIDIA on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.