Filevine

Senior Site Reliability engineer

$175K–$195KFull-time · Remote
✓ Verified live on the employer's own system · added yesterday
Save search
Senior · 8+ yrs exp

Requirements

Experience: 8+ years

Skills & tools

OperationsTeam LeadershipProgrammingCloud PlatformsCloud EngineeringDevopsDistributed SystemsPython

Benefits — mentioned in this posting

Health, dental & visionFamily / parental leave
Apply on company site ↗ See your fit → free

Full job description

Responsibilities

• Design and improve the monitoring, logging, distributed tracing, dashboards, alerting, SLIs,

and SLOs that give teams meaningful visibility into production health and customer impact.

• Build and maintain automation, internal tools, and CI/CD systems that increase engineering

efficiency, reduce toil, and support reliable deployments at scale. Take responsibility for the

quality and reliability of tools and services you support.

• Drive the implementation and continuous improvement of reliable systems for building,

deploying, testing, and operating Filevine products, proactively identifying and resolving

reliability, performance, scalability, and security risks before they impact customers.

• Own complex production incidents through detection, triage, communication, resolution,

and follow-up. Turn incident learning into durable corrective actions, stronger runbooks and

operating practices, and improvements that reduce recurring incidents and operational

burden.

Lead significant technical initiatives from problem definition and design through

implementation and adoption. Coordinate work across engineers and teams, communicate

tradeoffs and risks, and help ensure the work delivers the intended results.

• Mentor other Site Reliability Engineers through design reviews, incident follow-ups, paired

problem-solving, and meaningful delegation. Help engineers develop stronger technical

judgment and become increasingly capable of handling complex production work

independently.

• Participate in the shared on-call rotation and help ensure production systems are prepared

to operate reliably at scale through capacity planning, operational readiness, and

continuous improvements to resilience and recovery.

• Apply AI and machine learning to analyze operational signals, identify patterns, forecast

reliability and capacity risks, and implement improvements that make systems more

reliable, efficient, and easier to operate.

Qualifications

• 8+ years of hands-on experience in software engineering, cloud infrastructure, platform

engineering, DevOps, or related technical roles, including at least 5 years in a Site Reliability

Engineering or reliability-focused role.

• Strong knowledge of distributed systems and hands-on experience operating Kubernetes

workloads and cloud infrastructure in AWS or a comparable platform, with proficiency in

Infrastructure as Code, monitoring, logging, alerting, distributed tracing, SLIs, and SLOs.

• Strong proficiency with Python, Go, Bash, or a similar language, with demonstrated

experience building and maintaining production tooling, automation, CI/CD pipelines, and

deployment systems that reduce toil, improve reliability, and simplify ongoing operations.

• Demonstrated ability to lead troubleshooting, incident response, root cause analysis, and

long-term reliability improvements for complex production systems, including the

elimination of recurring incidents and operational work.

• Proven ability to mentor Site Reliability Engineers, help others build stronger technical

judgment, communicate clearly with technical and business stakeholders, and lead complex

initiatives from planning through delivery.

• Demonstrated experience applying AI and machine learning to operational data and

engineering workflows to identify patterns, forecast reliability or capacity risks, and

implement measurable improvements with appropriate safeguards.

$175,000 - $195,000 a year

Cool Company Benefits:

- A dynamic, rapidly growing company, focused on helping organizations thrive

- Medical, Dental, & Vision Insurance (for full-time employees)

- Competitive & Fair Pay

- Maternity & paternity leave (for full-time employees)

- Short & long-term disability

- Opportunity to learn from a dedicated leadership team

- Top-of-the-line company swag

Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. Grounded in a singular system of truth, Filevine brings together data, documents, workflows, and teams into one unified platform—where modern legal work happens with clarity and consistency.

Powered by LOIS, the Legal Operating Intelligence System, Filevine connects context across every matter to transform legal operations from reactive to proactive. LOIS reads, understands, and reasons across your data to surface insight, automate complexity, and give professionals the clarity and confidence to see more, know more, and do more. Fueled by a team of exceptional collaborators and innovators, Filevine’s rapid growth has earned AI awards and recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.

More jobs at Filevine

Similar jobs near Remote

Search Senior Site Reliability engineer jobs near Remote → Browse all live jobs

This posting was published by Filevine on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.