We are looking for a visionary SRE Architect to define and scale reliability architecture for cloud-native, distributed platforms. You will set SRE standards across design, observability, and automation on AWS—apply now to help shape how teams build and run resilient services.
Responsibilities
-
Design self-healing, highly available, fault-tolerant architectures on AWS
-
Define standards for disaster recovery, multi-region failover, and traffic protection patterns
-
Establish organization-wide SRE practices for SLOs, SLIs, and error budget governance
-
Architect a global observability strategy for logs, metrics, traces, and APM
-
Build internal tools and automation to reduce operational toil and speed up delivery
-
Lead capacity planning to ensure predictable performance and scalability
-
Introduce chaos engineering practices to continuously validate system resilience
-
Partner with engineering teams to embed reliability requirements into design patterns
-
Create reference architectures and guidance that improve consistency across services
-
Review system designs and recommend improvements for resiliency and operability
Requirements
-
3+ years of solution architecture experience for cloud-native or distributed systems
-
7+ years of site reliability engineering experience applying SRE principles in production
-
Technical leadership experience guiding architecture decisions and standards across teams
-
Architecture design expertise for high availability, fault tolerance, and multi-region resilience on AWS
-
Observability platform experience with logging, metrics, tracing, and APM tooling
-
Strong software engineering skills in at least one language such as Go, Python, Java, or Rust
-
Hands-on AWS experience across services such as EKS or ECS, RDS, serverless, IAM, and networking
-
Reliability engineering knowledge of SLOs, SLIs, and error budgets
-
Chaos engineering experience with fault injection and resilience validation practices
-
Capacity planning skills for forecasting and scaling highly distributed workloads
-
Clear communication skills to align stakeholders and drive consistent SRE practices
-
Upper-Intermediate English proficiency (B2)
Nice to have
-
AWS Solutions Architect certification
-
AWS DevOps Engineer certification
-
OpenTelemetry implementation experience
-
Experience with tools such as Datadog, Dynatrace, New Relic, Prometheus, or ELK Stack
We offer
-
International projects with top brands
-
Work with global teams of highly skilled, diverse peers
-
Healthcare benefits
-
Employee financial programs
-
Paid time off and sick leave
-
Upskilling, reskilling and certification courses
-
Unlimited access to the LinkedIn Learning library and 22,000+ courses
-
Global career opportunities
-
Volunteer and community involvement opportunities
-
EPAM Employee Groups
-
Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn
EPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.