We are seeking a Lead Site Reliability Engineer to embed with the team and drive reliable operations across a broad infrastructure landscape. You will own monitoring, observability, and logging with Dynatrace and Splunk, and help evolve dashboards, alerting, and analysis practices. You will also join the on-call rotation and partner on incident response. Apply now.
Responsibilities
-
Monitor and sustain the health, performance, and reliability of the client's applications and services
-
Operate and continuously enhance Dynatrace and Splunk, including dashboards, alerting, anomaly detection, and log analysis
-
Analyze alerts, determine likely root causes, and deliver actionable recommendations to engineering and incident management teams
-
Support incident response activities and participate in the on-call rotation
-
Perform Root Cause Analysis (RCA) and drive measurable post-incident improvements
-
Define and monitor SLOs, SLAs, and error budgets
-
Detect and remediate observability and monitoring gaps across services
-
Maintain operational runbooks and supporting documentation
Requirements
-
Proven background with 5+ years in Site Reliability Engineering, Production Operations, DevOps, or a closely related role
-
Deep expertise in Dynatrace and Splunk, including APM, alerting, dashboards, RUM, synthetic monitoring, service flow analysis, SPL queries, and log analysis
-
Hands-on experience with production incident management and on-call support, covering alert triage, incident response, RCA, and post-incident reviews
-
Strong troubleshooting skills and RCA capability across distributed applications and services
-
Solid understanding of application architecture, service dependencies, integrations, performance analysis, dependency mapping, and bottleneck identification
-
Practical experience with AWS services, including CloudWatch, ECS, EC2, ALB, Route53, RDS, and VPC
-
Working knowledge of CI/CD pipelines and release validation processes
-
English proficiency at B2 (Upper-Intermediate) level or higher
Nice to have
-
Familiarity with AI-assisted observability capabilities
-
Experience optimizing monitoring and alerting strategies
-
Exposure to Infrastructure as Code (Terraform or equivalent)
-
Travel/Airline industry experience
-
Background supporting modernization and cloud transformation initiatives
We offer
-
International projects with top brands
-
Work with global teams of highly skilled, diverse peers
-
Healthcare benefits
-
Employee financial programs
-
Paid time off and sick leave
-
Upskilling, reskilling and certification courses
-
Unlimited access to the LinkedIn Learning library and 22,000+ courses
-
Global career opportunities
-
Volunteer and community involvement opportunities
-
EPAM Employee Groups
-
Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn
EPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.