Tiger Analytics Inc

Sr. Site Reliability Engineer

Washington, DC
✓ Verified live on the employer's own system · added 94 days ago
Save search
Senior

Skills & tools

DevopsProgrammingManagementMachine LearningFinancial AnalysisGitPythonOperations
Apply on company site ↗ See your fit → free

Full job description

<h3><strong>Role Overview</strong></h3><p>We are seeking a high-caliber <strong>Site Reliability Engineer (SRE)</strong> to join our Forward Engineering team.

You will be the guardian of our production ecosystems, ensuring that our complex, data-driven AI platforms remain resilient, scalable, and highly performant.

This role is a hybrid of software engineering and systems architecture, with a specialized focus on <strong>MLOps</strong>—bridging the gap between model development and production-grade reliability.</p><p></p><h3><strong>Key Responsibilities</strong></h3><h3><strong>1.

Reliability &amp; Performance Engineering</strong></h3><ul><li><strong>SLA/SLO Management:</strong> Define, monitor, and maintain Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for critical AI/ML services.</li><li><strong>Error Budgeting:</strong> Manage error budgets to balance the velocity of feature releases from the ML team with the stability of the production environment.</li><li><strong>Scalability:</strong> Architect and manage auto-scaling strategies for <strong>Kubernetes (GKE)</strong> to handle fluctuating workloads during model training and high-volume inference.</li></ul><h3><strong>2.

MLOps &amp; AI Infrastructure</strong></h3><ul><li><strong>Model Serving Reliability:</strong> Ensure the high availability of <strong>Vertex AI endpoints</strong> and custom inference services.</li><li><strong>GPU/TPU Optimization:</strong> Monitor and optimize compute resource utilization (accelerators) to ensure cost-efficient performance for Large Language Models (LLMs).</li><li><strong>Pipeline Resilience:</strong> Support and stabilize ML pipelines (Vertex AI Pipelines/Kubeflow) to ensure seamless data flow from ingestion to model retraining.</li></ul><h3><strong>3.

Automation &amp; Orchestration (Eliminating "Toil")</strong></h3><ul><li><strong>Infrastructure as Code (IaC):</strong> Use <strong>Terraform</strong> or Pulumi to provision and manage consistent, version-controlled cloud environments.</li><li><strong>CI/CD &amp; GitOps:</strong> Design and optimize robust deployment pipelines for both application code and ML models using GitHub Actions, Cloud Build, or ArgoCD.</li><li><strong>Task Automation:</strong> Develop custom Python or Go scripts to automate repetitive operational tasks, self-healing mechanisms, and resource cleanup.</li></ul><h3><strong>4.

Monitoring, Alerting &amp; Incident Response</strong></h3><ul><li><strong>Observability:</strong> Build and manage comprehensive dashboards using <strong>Prometheus, Grafana, or Google Cloud Operations Suite (Stackdriver)</strong>.</li><li><strong>Incident Management:</strong> Act as a primary responder in on-call rotations, leading the technical resolution of production outages.</li><li><strong>Blameless Post-Mortems:</strong> Conduct deep-dive root cause analysis (RCA) to ensure systemic issues are identified and permanently remediated through code.</li></ul><p><strong>Requirements</strong></p><p><strong>Orchestration:</strong> Expert-level knowledge of <strong>Kubernetes (K8s)</strong> and Docker.</p><p><strong>MLOps Stack:</strong> Familiarity with tools such as <strong>Kubeflow, Vertex AI, MLflow, or DVC</strong>.</p><p><strong>Scripting:</strong> Strong proficiency in <strong>Python</strong> (for automation) and Bash; knowledge of Go is a plus.</p><p><strong>Data Systems:</strong> Experience managing the reliability of data-heavy services (BigQuery, Pub/Sub, or Vector Databases like Pinecone/Milvus).</p><p><strong>Networking:</strong> Solid understanding of VPCs, Load Balancers, DNS, and secure service mesh (Istio/Anthos).</p><p><strong>Benefits</strong></p><p>Benefits</p><p>Significant career development opportunities exist as the company grows.

The position offers a unique opportunity to be part of a small, fast-growing, challenging and entrepreneurial environment, with a high degree of individual responsibility.</p><p></p><p><em><strong>Tiger Analytics provides equal employment opportunities to applicants and employees without regard to race, color, religion, age, sex, sexual orientation, gender identity/expression, pregnancy, national origin, ancestry, marital status, protected veteran status, disability status, or any other basis as protected by federal, state, or local law.</strong></em></p>

More jobs at Tiger Analytics Inc

Similar jobs near Washington, DC

Search Sr. Site Reliability Engineer jobs near Washington, DC → Browse all live jobs

This posting was published by Tiger Analytics Inc on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.