- Experience operating large-scale GPU clusters and container orchestration frameworks (e.g. Kubernetes, Slurm, Docker)
- Strong systems background: Linux, networking, storage, infrastructure-as-code
- Knowledge of cloud platforms (GCP, AWS, or Azure) and their ML/AI service offerings
- Understanding of monitoring, logging, observability, and version control best practices for ML systems
- Familiarity with CUDA/NCCL and performance profiling for distributed workloads
- Owns deliverables end-to-end, from requirements through autonomous execution
Search Member of Technical Staff - Compute Cluster jobs near San Francisco → Browse all live jobs
This posting was published by Causal Labs on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.