a { text-decoration: none; color: #464feb; } tr th, tr td { border: 1px solid #e6e6e6; } tr th { background-color: #f5f5f5; }
The primary mandate of this role is to ensure production systems remain performant, healthy, secure, and reliable, with database and application performance serving as top priorities. This responsibility is supported through DevOps practices, including CI/CD, infrastructure as code, and deployment automation, as well as Site Reliability Engineering (SRE) ownership of observability, alerting, scalability, and incident prevention.
This individual will play a key role in driving the organization//'s on-premises to cloud migration strategy while improving the resilience, redundancy, and scalability of its platform. Operating within a product-focused and highly regulated environment, the ideal candidate values practical, proven solutions, prioritizing system stability, data integrity, and operational excellence over adopting technology trends for their own sake.
- Serve as the primary owner of production system health, availability, and performance.
- Build, maintain, and scale highly available production environments supporting APIs, services, and business-critical applications.
- Design and implement resilient architectures with redundancy, failover capabilities, and zero-downtime deployment strategies.
- Develop and enhance observability capabilities across infrastructure, networks, databases, and applications through centralized metrics, logs, and distributed tracing.
- Establish proactive monitoring and alerting frameworks to identify and resolve issues before customer impact occurs.
- Generate operational and SLA reporting for business and technical stakeholders.
- Diagnose and resolve database, infrastructure, and network performance issues, including SQL Server query optimization, indexing strategies, and execution plan analysis.
- Design and implement horizontal and vertical scaling strategies for APIs, web services, and database platforms.
- Conduct load, stress, and performance testing to validate scalability assumptions and identify bottlenecks.
- Partner with engineering teams to support capacity planning and growth initiatives across production workloads.
- Provision, manage, and optimize Azure cloud infrastructure using Terraform and infrastructure-as-code best practices.
- Administer cloud networking and security components, including VNets, VPN connectivity, Web Application Firewalls (WAFs), Network Security Groups (NSGs), and related controls.
- Maintain development, test, staging, and production environments while ensuring configuration consistency and operational health.
- Support, optimize, and troubleshoot CI/CD pipelines used for .NET/C# application delivery.
- Improve deployment automation, release reliability, and operational efficiency across development and production environments.
- Lead the migration of on-premises applications and infrastructure to Microsoft Azure as part of the organization//'s cloud transformation strategy.
- Assess existing dependencies and recommend appropriate modernization, cloud-native, or lift-and-shift approaches.
- Support the implementation of hybrid-cloud architectures where appropriate.
- 7+ years of experience in DevOps, Site Reliability Engineering (SRE), cloud infrastructure, or related disciplines supporting production environments.
- 10+ years of relevant experience preferred.
- Demonstrated success owning production system performance, availability, and operational health.
- Experience managing infrastructure and services against defined uptime, reliability, and SLA commitments.
- Strong SQL Server administration, performance tuning, and query optimization expertise.
- Hands-on experience with CI/CD platforms and deployment tooling such as Azure DevOps, GitHub Actions, Jenkins, or similar technologies.
- Strong experience with Microsoft Azure services, including compute, networking, security, monitoring, and automation capabilities.
- Proficiency with Terraform and infrastructure-as-code methodologies.
- Experience designing and supporting serverless, containerized, and cloud-native solutions using Azure services such as Azure Functions, App Services, Container Apps, and Logic Apps.
- Solid understanding of cloud networking and security fundamentals, including VPNs, VNets, WAFs, NSGs, TLS, and related technologies.
- Experience implementing scalable, resilient, and highly available distributed systems.
- Hands-on experience with API load and performance testing tools such as k6, JMeter, Azure Load Testing, or equivalent.
- Experience implementing observability and monitoring solutions utilizing Azure Monitor, Application Insights, Grafana, Prometheus, or similar platforms.
- Microsoft Azure certifications such as AZ-104, AZ-305, AZ-400, or equivalent.
- Experience supporting hybrid cloud environments and large-scale cloud migration initiatives.
- Experience working within distributed, nearshore, or globally distributed engineering teams.
- Experience supporting organizations operating in highly regulated industries.