Experience: 8+ years
- Own and evolve our Kubernetes infrastructure, including cluster management, service mesh configuration, and container security policies.
- Design and implement progressive delivery pipelines with canary deployments, automated rollbacks, and deployment health validation.
- Build and maintain our observability infrastructure in Datadog, including dashboards, monitors, SLOs, and distributed tracing.
- Drive incident response for high-severity outages and proactively model capacity needs for low-latency AI inference.
- Architect and automate secure infrastructure using Infrastructure-as-Code for VPCs, IAM policies, Kubernetes manifests, and private cloud deployments.
- Maintain and improve the infrastructure controls that support our SOC 2 compliance posture.
- Lead customer engagements for enterprise rollouts and mentor mid-level engineers on infrastructure best practices.
- 8+ years in infrastructure engineering or DevOps at high-growth or hyperscale companies.
- Experience with Docker and Kubernetes, including production cluster management, Helm, and service mesh technologies.
- A proven track record of architecting and operating AWS (preferred), GCP, or Azure at an enterprise scale.
- Experience with observability platforms, preferably Datadog (metrics, logs, APM, distributed tracing).
- A strong background in Infrastructure-as-Code (Terraform, Helm, Kustomize) and safe deployment practices (progressive delivery, canary deployments, GitOps, automated rollbacks).
- "Battle scars" from leading outages, capacity events, and large-scale incident reviews.
- Direct involvement in SOC 2 or other compliance audit preparation or remediation.
- Direct experience with private-cloud or on-premises deployments for regulated customers.
- Previous experience at startups scaling infrastructure from the early stages to the enterprise level.
- A background in fintech or building systems for highly regulated industries.
- Experience with AI/ML infrastructure and model deployment at scale.
- Build for Scale: You thrive at the intersection of technical leadership and customer impact, building systems that enable rapid development while maintaining the highest standards of security, compliance, and reliability.
- Infrastructure as a Product: You see infrastructure as a product for your engineering peers and understand the value of platform automation in enabling developer velocity.
- High-Impact Work: Your contributions will have a direct, measurable impact on how financial institutions adopt AI to fight crime.
- Mentorship and Leadership: You are comfortable balancing technical excellence with mentoring others and leading customer engagements.
Search Software Engineer, Infrastructure jobs near San Francisco, CA → Browse all live jobs
This posting was published by Bretton AI on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.