We’re seeking a future team member for the role of Sr. Site Reliability Automation Engineer to join our Technology team. This role is located in Lake Mary, FL and Pittsburgh, PA.
Design and implement end-to-end observability (logs, metrics, traces) across distributed systemsBuild Observability & Monitoring
Integrate and optimize tools such as AppDynamics, Dynatrace, Grafana, and Splunk
Develop dashboards, alerts, and telemetry frameworks to provide real-time visibility
Identify gaps in monitoring and drive adoption of best practices
Identify repetitive operational work and automate it using code and tooling
Enable scalable, reliable processes through automation and engineering rigor
Improve operational efficiency across production environments
Troubleshoot and resolve complex production issues across distributed systems
Participate in incident management, triage, and root cause analysis
Improve monitoring and automation based on recurring incident patterns
Collaborate with support and engineering teams to improve system stability
Define and measure service health using SLIs/SLOs and key performance metrics
Contribute to performance optimization and capacity planning
Provide input into system architecture to improve resilience and scalability
3–6 years of experience in Site Reliability Engineering, Software Engineering
Strong programming background in Java (preferred) or another modern language
Hands-on experience supporting and troubleshooting production systems
Experience with distributed systems or microservices architectures
Knowledge of SRE concepts like observability, incident management, and reliability engineering
This posting was published by BNY Wealth Newport Beach CA on their own careers system and is shown here with a direct
link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.