LSEG (formerly Refinitiv)

Senior Engineer - Site Reliability Engineering

USA-Allen-700 Central Expressway
✓ Verified live on the employer's own system · added 4 days ago
Save search
Mid-level · 5+ yrs exp

Requirements

Experience: 5+ years

Skills & tools

DevopsOperationsRecordkeepingSecurityManagementTeam LeadershipCloud PlatformsFinancial Analysis
Apply on company site ↗ See your fit → free

Full job description

We are looking for a person who is passionate about reliability engineering and who bring a continuous improvement approach to everything they do!

Lead the establishment of SRE foundations for new projects building environments, monitoring, alerting, and ensuring operational readiness from day one.

Define, implement, and champion observability standards, tooling, and guidelines across metrics, logs, traces, and SLIs/SLOs.

Design and evolve monitoring and alerting solutions that improve visibility, reduce toil, and strengthen system health.

Continuously drive reliability improvements across our environments through incident reduction, performance tuning, and building resilient patterns.

Partner with Security teams to ensure our platforms meet compliance, security, and risk-management expectations.

Lead seamless handovers from project delivery into BAU SRE operations by ensuring documentation, readiness, and strong operational practices.

Be a technical leader and mentor supporting engineers, shaping engineering standards, and fostering a culture of learning and development.

- 5+ years of hands-on technical experience in SRE, Platform Engineering, Infrastructure, or related roles

- Strong experience with Azure, including services such as AKS, Azure Container Apps, Virtual Machines, virtual networking (VNet), Azure Active Directory (Entra ID), and Azure managed services

- Hands-on experience with Kubernetes and containerized platforms

- Proven experience designing and operating observability platforms , including monitoring, logging, and alerting

- Hands-on experience with Datadog for metrics, logs, APM, and alerting

- Strong understanding of SRE principles , including SLOs, error budgets, incident management, and reliability engineering

- Understanding of cloud security principles and experience collaborating with security teams

- Experience with cloud cost optimization strategies and tooling

- Exposure to Infrastructure as Code (e.g., Terraform, CloudFormation)

- Experience in large-scale, complex, or regulated environments

- Knowledge of vector databases and RAG architectures for building internal SRE knowledge assistants.

- Knowledge of Generative AI and LLM platforms (e.g., Claude, Amazon Bedrock)

- Experience integrating AI with observability stacks (Prometheus, Grafana, ELK, OpenTelemetry) for proactive issue detection.

- Strong technical authority with the ability to influence design and operational decisions

- Highly collaborative, comfortable working across architecture, engineering, security, and operations teams

- Calm and methodical under pressure, especially during incidents and critical issues

- Pragmatic problem-solver who balances reliability, security, cost, and delivery speed

- Clear communicator, able to explain complex technical concepts to diverse audiences

More jobs at LSEG (formerly Refinitiv)

Similar jobs near USA-Allen-700 Central Expressway

Search Senior Engineer - Site Reliability Engineering jobs near USA-Allen-700 Central Expressway → Browse all live jobs

This posting was published by LSEG (formerly Refinitiv) on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.