NVIDIA | NVIDIA

Senior Manager, Storage Production Engineering

US, CA, Santa Clara
✓ Verified live on the employer's own system · added 11 days ago
Save search
Senior · 6+ yrs exp

Requirements

Education: Bachelor's degree or related field

Experience: 6+ years

Skills & tools

CoachingData AnalysisManagementRoot Cause AnalysisDevopsMachine LearningOperationsTeam Leadership

Benefits — mentioned in this posting

Health, dental & vision401(k) / retirementEquity / stockPaid time off
Apply on company site ↗ See your fit → free

Full job description

- Leading and coaching a team of Storage Production Engineers, creating a collaborative, inclusive, and learning focused environment.

- Designing, deploying, and improving large scale storage systems, including distributed storage, parallel file systems, and object storage.

- Using automation, monitoring, and analytics to make storage services more reliable, efficient, and easier to operate.

- Owning capacity planning, data lifecycle management, cost awareness, and high availability and disaster recovery plans for storage.

- Evaluating and adopting modern storage approaches such as NVMe over Fabrics, RDMA, high speed interconnects, and cloud based storage.

- Guiding incident response and root cause analysis for storage issues, and putting in place changes that prevent repeat problems.

- Partnering with engineering, DevOps, and AI/ML teams to improve data pipelines, access patterns, and workflow performance.

- BS or MS in Computer Science, Storage Systems, or a related technical field, or equivalent experience.

- 12+ overall years of experience in large scale storage architecture, operations, production engineering, or infrastructure.

- 6+ years of people management or technical leadership experience with storage, infrastructure, or site reliability teams.

- Direct experience managing infrastructure operations including on-call rotations, incident response, ongoing maintenance, troubleshooting, and optimization of production systems, managing SLOs and operational KPIs

- Hands on experience with parallel file systems (such as Lustre or GPFS), distributed storage (such as Ceph or MinIO), and enterprise object or NAS platforms (such as S3 compatible systems, NetApp, or Pure Storage).

- Strong knowledge of block, file, and object storage, including how to tune performance, protect data, and design for high availability.

- Experience with storage networking and protocols like NFS, SMB, iSCSI, Fibre Channel, RDMA, and NVMe-oF.

- Practical experience with automation and infrastructure as code using tools such as Terraform, Ansible, or Puppet.

- Strong knowledge of monitoring and observability tools (for example Prometheus, InfluxDB, or Elastic stack), logging, and alerting used to operate and improve storage systems.

- Experience managing and scaling SRE/Production Engineering teams in large scale, mission critical environments with focus on operational excellence, service availability, and performance

- A track record of improving reliability, simplicity, and day to day operations for large, business critical storage systems.

- Experience building or scaling storage for AI/ML or HPC workloads, including hybrid or multi cloud setups (for example AWS S3, Azure Blob, or Google Cloud Storage), as well as on-prem infrastructure

- Experience with software defined storage, cloud native storage, and Kubernetes based storage orchestration.

- A visible passion for mentoring, career development, and building a supportive, high performing team culture.

NVIDIA offers a comprehensive benefits package designed to support your physical, mental, and financial well-being. This may include medical, dental, and vision coverage, mental health resources, retirement and 401(k) plans, an employee stock purchase plan, paid time off and holidays, family and caregiving leave, and a range of wellness and development programs.

Specific benefits vary by location, but all are built to help you do your life's work and grow your career at NVIDIA.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.

Applications for this job will be accepted at least until August 2, 2026.

More jobs at NVIDIA | NVIDIA

Similar jobs near US, CA, Santa Clara

Tell me when more Senior Manager, Event Content Program Management & Operations jobs post near Santa Clara, CALIFORNIA (Remote) We re-check every listing against the employer’s own board — no résumé needed.

Search Senior Manager, Storage Production Engineering jobs near US, CA, Santa Clara → Browse all live jobs

This posting was published by NVIDIA | NVIDIA on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.