Category Labs

Senior DevOps / Infrastructure Engineer

$180K–$250KFull-time · New York (Hybrid) (Remote)
✓ Verified live on the employer's own system · added 4 days ago
Save search
Mid-level · 5+ yrs exp

Requirements

Education: Bachelor's degree or related field

Experience: 5+ years

Skills & tools

DevopsSecurityOperationsMachine LearningLinuxTroubleshootingProgrammingScripting
Apply on company site ↗ See your fit → free

Full job description

We're looking for a Senior DevOps / Infrastructure Engineer to operate the infrastructure behind Monad, and to push how much of that operation can be driven by AI. You'll keep our globally-distributed validator, full node, and archive fleet healthy across mainnet and testnet, own our infrastructure-as-code and observability, and build the agentic tooling and guardrails that let a small team safely operate a large fleet.

As more of our engineering shifts toward AI, this role is central to designing the workflows, deterministic guardrails, and security perimeters within which autonomous agents operate our infrastructure; you'll also help stand up and operate the infrastructure behind our own growing model workloads.

- Operate the Monad node fleet: health, sync, upgrades, and recovery across validators, full nodes, archive/historical, and indexer nodes on mainnet and testnet, including safe, staged rollouts and incident response.

- Own our infrastructure-as-code: Ansible for fleet configuration, Terraform + Atlantis for cloud and DNS, and Kubernetes/Flux (GitOps) for platform services.

- Build and operate observability and alerting (Prometheus, Grafana, Loki); create dashboards and alerts that catch problems before they page while minimizing false positives.

- Automate the release pipeline: node upgrades, canary rollouts, snapshot/restore, and the guardrails that bound blast radius (e.g., protecting validators from automated changes).

- Design and build agentic operations: develop AI agents, tooling (e.g., MCP servers), and runbooks-as-code that let agents safely investigate, diagnose, and execute routine operations, with deterministic guardrails and human oversight.

- Codify operational knowledge into tools and automation that the whole team, and its agents, can reuse.

- Harden nodes and services, manage secrets, and continuously drive down manual toil.

- You have 5+ years in DevOps, SRE, or Infrastructure Engineering, operating production systems at scale.

- You have strong Linux, systemd, networking, and shell fundamentals, and you're comfortable debugging live systems over SSH.

- You have deep, hands-on infrastructure-as-code experience with Ansible and Terraform.

- You have experience with observability stacks (Prometheus, Grafana, Loki, or equivalents).

- You have hands-on fluency with AI-assisted engineering: you use coding agents and LLM tooling in your daily workflow and have judgment on where it helps and where it's risky.

- You have experience designing automation with safe guardrails, and you bring calm, methodical incident response.

- You have programming and scripting experience (e.g., Python, bash).

- Experience with Kubernetes and GitOps (Flux or Argo) is a plus.

- Experience building AI agent tooling, MCP servers, or agent orchestration frameworks is a plus.

- Experience serving inference, either locally or as a service is a plus.

- Previous experience with blockchain clients or node operations is a plus.

- A Bachelor of Science in Computer Science, Engineering, or a related field is a plus.

More jobs at Category Labs

Similar jobs near New York (Hybrid) (Remote)

Tell me when more DevOps Engineer jobs post near Remote (US) We re-check every listing against the employer’s own board — no résumé needed.

Search Senior DevOps / Infrastructure Engineer jobs near New York (Hybrid) (Remote) → Browse all live jobs

This posting was published by Category Labs on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.