We are seeking an experienced L3 Support Engineer to provide production support and operational excellence for enterprise data platforms built on Databricks, Apache Airflow, DAG-based orchestration, and DBT (Data Build Tool). The role involves incident management, problem resolution, monitoring, performance tuning, root cause analysis, and continuous improvement of data pipelines and analytics platforms.
The ideal candidate should have strong troubleshooting expertise, experience supporting large-scale data environments, and the ability to collaborate with engineering, business, and operations teams to ensure high availability and reliability of critical data workloads.
Key Responsibilities
Production & BAU Support
- Provide L3 production support for Data Engineering and Analytics platforms.
- Monitor and support critical data pipelines, ETL/ELT workflows, and reporting systems.
- Analyze and resolve production incidents within SLA timelines.
- Perform impact assessment and prioritization of incidents.
- Conduct troubleshooting of failed jobs, workflows, and data quality issues.
- Participate in on-call support rotations and major incident management.
Databricks Support
- Support Databricks jobs, notebooks, workflows, and clusters.
- Troubleshoot cluster performance, job failures, and resource utilization issues.
- Optimize Spark SQL, PySpark, and Databricks workloads.
- Monitor cluster health, execution performance, and cost optimization.
- Coordinate deployment and release activities in Databricks environments.
Airflow & DAG Support
- Monitor Apache Airflow environments and scheduler performance.
- Troubleshoot DAG failures, task retries, dependency issues, and scheduling conflicts.
- Analyze execution logs and identify bottlenecks.
- Ensure workflow reliability and performance optimization.
- Manage Airflow operational configurations and upgrades.
DBT Support
- Support DBT models, transformations, and deployments.
- Troubleshoot DBT execution failures and dependency issues.
- Validate data transformation logic and data quality checks.
- Monitor DBT jobs and lineage dependencies.
- Support CI/CD deployment of DBT projects.
Incident & Problem Management
- Perform Root Cause Analysis (RCA) for recurring issues.
- Create post-incident reports and preventive action plans.
- Manage service requests, change requests, and problem tickets.
- Collaborate with development teams for permanent fixes.
- Maintain knowledge base articles and operational runbooks.
Monitoring & Automation
- Implement monitoring dashboards and alerts.
- Identify opportunities for support automation and self-healing solutions.
- Develop scripts for operational efficiency.
- Improve observability and operational metrics.
Required Skills
Technical Skills
- Strong hands-on experience with:
- Databricks
- Apache Airflow
- DAG-based workflow orchestration
- DBT (Data Build Tool)
- Experience with:
- PySpark / Spark SQL
- Python
- SQL
- ETL/ELT pipelines
- Data Warehouse concepts
- Knowledge of cloud platforms:
- Azure (preferred)
- AWS
- GCP
- Experience with CI/CD tools and deployment processes.
- Understanding of monitoring and logging tools.
Support Skills
- Incident Management
- Problem Management
- Change Management
- SLA Management
- RCA Documentation
- Production Monitoring
- Troubleshooting & Debugging