We are looking for a Lead Infrastructure Engineer to architect, build, automate, and manage infrastructure spanning our on-prem data centers and GCP environments. This role fits an engineer who developed their foundation through hands-on operations work (racking hardware, running networks, bare metal OS installs, keeping systems online at 2am) and has since cultivated deep skills in automation, infrastructure as code, and public cloud platforms. You'll define technical direction for infrastructure, mentor engineers, and take ownership of the systems that support our global-scale operations.
Responsibilities
-
Architect and manage large-scale, multi data center infrastructure supporting global operations across on-prem hardware, GCP, and AWS
-
Drive automation efforts for infrastructure deployment and configuration management, guiding the shift from legacy tooling toward modern Infrastructure as Code solutions like Terraform, Puppet, and Ansible
-
Design and sustain traffic management, load balancing, and DNS systems at scale
-
Construct and manage monitoring and observability systems spanning both on-prem and cloud environments
-
Create tooling that enables distributed, auditable systems administration
-
Author and maintain process, policy, and procedural documentation for the infrastructure team
Requirements
-
At least 5 years of relevant experience in infrastructure or systems engineering, including direct operational responsibility for production systems
-
Experience handling all aspects of remote management for physical hardware infrastructure hosted at colocation facilities
-
Extensive Linux systems administration experience, such as Red Hat or Debian, at scale
-
Practical experience with core networking, including Cisco hardware, DNS, load balancing, and traffic management
-
Experience with enterprise storage and backup systems
-
Established track record automating infrastructure using tools such as Ansible, Puppet, or comparable technologies
-
Skilled in Python or a similar language for infrastructure tooling and automation
-
Production experience with GCP, including Compute Engine, networking, storage, and IAM
-
Demonstrated capability to lead infrastructure projects and mentor fellow engineers
-
Excellent English proficiency (B2 level or higher)
Nice to have
-
Experience transitioning on-prem workloads to public cloud platforms
-
Experience with distributed monitoring and logging stacks, such as Grafana or Prometheus
-
Experience with container orchestration tools, such as Kubernetes or GKE
-
Experience guiding and mentoring junior engineers
We offer
-
International projects with top brands
-
Work with global teams of highly skilled, diverse peers
-
Healthcare benefits
-
Employee financial programs
-
Paid time off and sick leave
-
Upskilling, reskilling and certification courses
-
Unlimited access to the LinkedIn Learning library and 22,000+ courses
-
Global career opportunities
-
Volunteer and community involvement opportunities
-
EPAM Employee Groups
-
Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn
EPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.