Job Title: SRE Engineer
Location: Mexico City (Hybrid Position) 3day Onsite 2days Remote
Duration: Long term contract
The SRE Engineer will improve the reliability, observability, deployment safety, and operational readiness of cloud-native e-commerce services.
Required Qualifications
· Strong Site Reliability Engineering experience supporting highly available, production-scale systems.
· Strong hands-on Google Cloud Platform (GCP) experience.
· Deep understanding and practical application of SLIs, SLOs, error budgets, operability, and reliability engineering principles.
· Strong experience with observability and instrumentation, including metrics, logging, tracing, alerting, and production diagnostics.
· Experience with production troubleshooting, incident response, root-cause analysis, and operational readiness.
· Experience with Terraform or another Infrastructure as Code technology.
· Experience developing and improving CI/CD pipelines and deployment practices.
· Proficiency with Python, Bash, or another scripting/automation language.
· Ability to identify systemic reliability issues and drive engineering solutions rather than primarily responding to operational incidents.
· Strong technical leadership, collaboration, and communication skills across application, platform, and engineering teams.
Preferred Qualifications
· Experience supporting Kubernetes and GKE workloads in production.
· Experience with GitHub Actions.
· Experience with GCP Cloud Monitoring, OpenTelemetry, Prometheus, or Grafana.
· Experience with Istio or another service mesh.
· Experience with Apigee or another API gateway.
· Experience implementing canary, blue/green, or progressive delivery strategies.
· Experience improving engineering automation and automated testing practices.
· This is a hands-on technical leadership role with the opportunity to become a key technical authority for SRE practices across the organization.
GCP: Must show meaningful production GCP experience, not just certification or a skills-list mention.
SRE evidence: Look for actual examples of SLI/SLO, error budgets, reliability improvements, incident management, reducing MTTR/toil. “SRE” in the job title is insufficient.
Observability: Prefer statements like instrumented application/services, designed telemetry, implemented metrics/logs/traces, built SLO-based alerting. Be cautious when resumes only list Datadog, Dynatrace, Grafana, Splunk, etc.
Hands-on ownership: Look for “designed, implemented, instrumented, troubleshot, coded” with specific outcomes. Too much “led/managed/coordinated” is a concern for this role.
DevOps-heavy profiles: If most evidence is Terraform + CI/CD + Kubernetes + cloud infrastructure, but little application reliability/observability/SRE principles, do not treat that as strong SRE experience.
Interested candidates, please send your updated resume to [email protected]
Job Type: Contract
Work Location: Hybrid remote in Ciudad de México, CDMX