Education: Bachelor's degree or related field
Experience: 5+ years
We're looking for a Senior DevOps / Infrastructure Engineer to operate the infrastructure behind Monad, and to push how much of that operation can be driven by AI. You'll keep our globally-distributed validator, full node, and archive fleet healthy across mainnet and testnet, own our infrastructure-as-code and observability, and build the agentic tooling and guardrails that let a small team safely operate a large fleet.
As more of our engineering shifts toward AI, this role is central to designing the workflows, deterministic guardrails, and security perimeters within which autonomous agents operate our infrastructure; you'll also help stand up and operate the infrastructure behind our own growing model workloads.
- Operate the Monad node fleet: health, sync, upgrades, and recovery across validators, full nodes, archive/historical, and indexer nodes on mainnet and testnet, including safe, staged rollouts and incident response.
- Own our infrastructure-as-code: Ansible for fleet configuration, Terraform + Atlantis for cloud and DNS, and Kubernetes/Flux (GitOps) for platform services.
- Build and operate observability and alerting (Prometheus, Grafana, Loki); create dashboards and alerts that catch problems before they page while minimizing false positives.
- Automate the release pipeline: node upgrades, canary rollouts, snapshot/restore, and the guardrails that bound blast radius (e.g., protecting validators from automated changes).
- Design and build agentic operations: develop AI agents, tooling (e.g., MCP servers), and runbooks-as-code that let agents safely investigate, diagnose, and execute routine operations, with deterministic guardrails and human oversight.
- Codify operational knowledge into tools and automation that the whole team, and its agents, can reuse.
- Harden nodes and services, manage secrets, and continuously drive down manual toil.
- You have 5+ years in DevOps, SRE, or Infrastructure Engineering, operating production systems at scale.
- You have strong Linux, systemd, networking, and shell fundamentals, and you're comfortable debugging live systems over SSH.
- You have deep, hands-on infrastructure-as-code experience with Ansible and Terraform.
- You have experience with observability stacks (Prometheus, Grafana, Loki, or equivalents).
- You have hands-on fluency with AI-assisted engineering: you use coding agents and LLM tooling in your daily workflow and have judgment on where it helps and where it's risky.
- You have experience designing automation with safe guardrails, and you bring calm, methodical incident response.
- You have programming and scripting experience (e.g., Python, bash).
- Experience with Kubernetes and GitOps (Flux or Argo) is a plus.
- Experience building AI agent tooling, MCP servers, or agent orchestration frameworks is a plus.
- Experience serving inference, either locally or as a service is a plus.
- Previous experience with blockchain clients or node operations is a plus.
- A Bachelor of Science in Computer Science, Engineering, or a related field is a plus.
Search Senior DevOps / Infrastructure Engineer jobs near New York (Hybrid) (Remote) → Browse all live jobs
This posting was published by Category Labs on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.