Intuitive

Sr MLOps Engineer

Full-time · Sunnyvale, CA
✓ Verified live on the employer's own system · added 4 days ago
Save search
Mid-level · 3+ yrs exp

Requirements

Education: Bachelor's degree or related field

Experience: 3+ years

Skills & tools

Machine LearningDevopsRecordkeepingSecurityTroubleshootingScriptingPythonLinux
Apply on company site ↗ See your fit → free

Full job description

In this role, you will be responsible for designing, building, and maintaining the infrastructure and tools necessary to support the entire machine learning lifecycle, from development to deployment. You will work closely with ML engineers and software developers across Intuitive to ensure that machine learning models are seamlessly integrated into our systems and deliver value at scale.

The ideal candidate is an independent and fast-paced engineer with excellent problem-solving skills and practical working knowledge of modern ML development techniques.

- Bootstrap and maintain a production-grade Kubernetes cluster, including CNI networking and storage integration - Deploy and configure ML orchestration tooling (e.g., Metaflow) and artifact/dataset storage solutions to support reproducible ML workflows - Validate GPU node health and configuration across heterogeneous hardware (B200, L40S, A6000, V100), including driver/CUDA standardization and topology checks - Design and execute team migration playbooks, working directly with engineering teams to port workflows, migrate datasets/artifacts, and roll out tool updates - Write and maintain runbooks, architecture documentation, and disaster recovery procedures - Participate in on-call rotation and incident response for platform-level issues - Collaborate with IT/Security on identity integration, access control, and compliance

requirements - Continuously evaluate and adopt infrastructure best practices for reliability, cost, and developer experience

- 3+ years of experience in infrastructure, DevOps, or MLOps roles, or equivalent practical experience - Demonstrated experience operating Kubernetes in production (networking, storage, RBAC, troubleshooting) - Strong scripting/automation skills in Python and/or Bash; comfort with Infrastructure-as-Code tools (Ansible, Helm, Terraform, or similar) - Hands-on experience with at least one distributed storage system (S3, MinIO, NetApp, or similar) - Experience building or maintaining CI/CD pipelines (GitLab CI, ArgoCD, or equivalent) - Solid understanding of Linux systems administration and networking fundamentals - Excellent communication and documentation skills, with the ability to write clear runbooks and migration guides - High degree of autonomy and comfort working across the full stack, iteratively building solutions

- Bachelor's or Master's degree in Computer Science, Engineering, or a related field; or equivalent experience

- Experience with ML orchestration frameworks (Metaflow, MLflow, Kubeflow, or similar) - Familiarity with GPU infrastructure (NVIDIA drivers, CUDA, NVLink/NUMA topology, MIG partitioning) - Prior experience in a regulated industry (healthcare, finance, or similar) where auditability and access control are critical - Experience leading or supporting large-scale infrastructure migrations with multiple stakeholder teams

We provide market-competitive compensation packages, inclusive of base pay, incentives, benefits, and equity. It would not be typical for someone to be hired at the top end of range for the role, as actual pay will be determined based on several factors, including experience, skills, and qualifications. The target compensation ranges are listed.

Required Skills and Experience

  • 3+ years of experience in infrastructure, DevOps, or MLOps roles, or equivalent practical experience
  • Demonstrated experience operating Kubernetes in production (networking, storage, RBAC, troubleshooting)
  • Strong scripting/automation skills in Python and/or Bash; comfort with Infrastructure-as-Code tools (Ansible, Helm, Terraform, or similar)
  • Hands-on experience with at least one distributed storage system (S3, MinIO, NetApp, or similar)
  • Experience building or maintaining CI/CD pipelines (GitLab CI, ArgoCD, or equivalent)
  • Solid understanding of Linux systems administration and networking fundamentals
  • Excellent communication and documentation skills, with the ability to write clear runbooks and migration guides
  • High degree of autonomy and comfort working across the full stack, iteratively building solutions
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field; or equivalent experience
  • Experience with ML orchestration frameworks (Metaflow, MLflow, Kubeflow, or similar)
  • Familiarity with GPU infrastructure (NVIDIA drivers, CUDA, NVLink/NUMA topology, MIG partitioning)
  • Prior experience in a regulated industry (healthcare, finance, or similar) where auditability and access control are critical
  • Experience leading or supporting large-scale infrastructure migrations with multiple stakeholder teams Primary Function of Position

In this role, you will be responsible for designing, building, and maintaining the infrastructure and tools necessary to support the entire machine learning lifecycle, from development to deployment. You will work closely with ML engineers and software developers across Intuitive to ensure that machine learning models are seamlessly integrated into our systems and deliver value at scale.

The ideal candidate is an independent and fast-paced engineer with excellent problem-solving skills and practical working knowledge of modern ML development techniques.

  • Bootstrap and maintain a production-grade Kubernetes cluster, including CNI networking and storage integration
  • Deploy and configure ML orchestration tooling (e.g., Metaflow) and artifact/dataset storage solutions to support reproducible ML workflows
  • Validate GPU node health and configuration across heterogeneous hardware (B200, L40S, A6000, V100), including driver/CUDA standardization and topology checks
  • Design and execute team migration playbooks, working directly with engineering teams to port workflows, migrate datasets/artifacts, and roll out tool updates
  • Write and maintain runbooks, architecture documentation, and disaster recovery procedures
  • Participate in on-call rotation and incident response for platform-level issues
  • Collaborate with IT/Security on identity integration, access control, and compliance requirements
  • Continuously evaluate and adopt infrastructure best practices for reliability, cost, and developer experience

More jobs at Intuitive

Similar jobs near Sunnyvale, CA

Tell me when more Sr. Automation Manufacturing Engineer jobs post near Sunnyvale, CA We re-check every listing against the employer’s own board — no résumé needed.

Search Sr MLOps Engineer jobs near Sunnyvale, CA → Browse all live jobs

This posting was published by Intuitive on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.