Experience: 5+ years
You'll design and scale the compute, storage, and data infrastructure that powers advanced AI research. Your work will ensure the speed, reliability, and reproducibility of every experiment and deployment.
- Develop and maintain scalable compute infrastructure (cloud, bare metal, clusters).
- Automate build, test, and deploy pipelines for research and production.
- Build tools for experiment tracking, monitoring, and reproducibility.
- Ensure security, reliability, and rapid iteration across all environments.
- Deep experience with cloud-native platforms (AWS, GCP, OCI), containerization, and orchestration (Kubernetes, Docker).
- Strong proficiency in Python and infrastructure-as-code tools (Terraform, Ansible, etc.).
- Experience scaling infra for ML, data, or research workflows.
- Relentless focus on reliability, observability, and developer velocity.
Search Member of Technical Staff (Infrastructure) jobs near New York City (Remote) → Browse all live jobs
This posting was published by ATG on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.