Location City
Guadalajara
Description
The opportunity
We are investing in the data foundation behind ATG's customer experiences and business decisions. As a Senior Data Engineer on the Data Enablement team, you will build production-grade data products that serve analytics, search, recommendations, personalization, and machine learning. You will work closely with product managers, analysts, data scientists, ML engineers, and software engineers to turn ambiguous needs into dependable, well-documented datasets and pipelines.
This is a hands-on engineering role for someone who cares about maintainability, data quality, and measurable outcomes. You will help shape standards and architecture while still writing code, reviewing designs, troubleshooting failures, and improving the platform.
Key Responsibilities
What you will do
Build durable data products
Design, build, and operate batch and event-driven pipelines for auction, inventory, customer, and transaction data.
Develop reusable transformation models and curated datasets in Snowflake and dbt for analytics and operational use cases.
Orchestrate complex dependencies with Airflow, Dagster, or a comparable workflow platform.
Design data models and interfaces that are clear, scalable, and easy for downstream teams to use.
Raise reliability and data quality
Define data contracts, validation rules, freshness expectations, lineage, and service-level objectives for critical datasets.
Implement automated testing, anomaly detection, alerting, and observability across the data lifecycle.
Own production issues through diagnosis, recovery, root-cause analysis, and prevention.
Improve query performance, warehouse efficiency, and cloud cost without compromising reliability.
Enable machine learning and customer experiences
Create versioned training, validation, and inference datasets for search, recommendations, personalization, and other ML products.
Partner with ML engineers and data scientists to make feature computation reproducible and consistent across experimentation and production.
Support experimentation by delivering trustworthy exposure, interaction, and outcome data for A/B testing and model evaluation.
Strengthen engineering practices
Apply software engineering practices to data work, including modular design, code review, automated testing, CI/CD, and infrastructure as code.
Improve documentation, discoverability, access controls, and governance for shared data products.
Contribute to architectural decisions, technical standards, and pragmatic platform improvements.
Mentor engineers and help the team make sound trade-offs among speed, scale, cost, and maintainability.
Key Requirements
What you bring
Five or more years of experience building and operating data pipelines or data platforms in production.
Strong Python and advanced SQL skills, including testing, debugging, performance tuning, and maintainable code design.
Hands-on experience with Snowflake or another modern cloud data platform, plus practical knowledge of dimensional and analytical data modeling.
Production experience with dbt or a comparable transformation framework and with Airflow, Dagster, Prefect, or similar orchestration tooling.
Experience with AWS data services and cloud storage; equivalent experience on another major cloud platform is welcome.
A working understanding of data quality, lineage, observability, data contracts, and operational ownership.
Comfort with Git, code review, automated testing, and CI/CD for data pipelines.
Clear communication and the ability to work across product, analytics, software engineering, data science, and ML teams.
A bachelor's degree in a relevant field or equivalent practical experience.
Useful, but not required
Event streaming or real-time processing with Kafka, Kinesis, Flink, Spark Structured Streaming, or similar technologies.
Distributed processing with Spark and familiarity with Parquet, Avro, JSON, and open table formats.
Search or retrieval systems such as Elasticsearch/OpenSearch, vector databases, or embedding pipelines.
Feature stores, ML data pipelines, model monitoring, or other MLOps capabilities.
Infrastructure as code and container platforms, including Terraform, Docker, or Kubernetes.
Data catalogs, semantic layers, master data management, or metadata-driven governance.
Experience with ecommerce, marketplaces, auctions, GDPR, privacy controls, or regulated data environments.
Employment Type
Permanent