Job Overview
We are seeking an experienced Sr. Principal Engineer, Site Reliability (SRE) to establish and scale the observability and reliability foundation for our global multi-cloud SaaS platform serving thousands of customers worldwide. This role will define enterprise standards for logging, metrics, tracing, alerting, dashboards, SLIs/SLOs, and service ownership, then work hands-on with Engineering and Operations teams to put those standards into practice. The successful candidate will bring deep experience operating observability at scale, modernizing telemetry and tooling, reducing alert noise and tool sprawl, and using actionable signals to improve detection, MTTR, performance, and reliability. Off hours support as needed
Success Metrics
- Customer Impact: Reduced MTTD/MTTR and improved customer experience through faster detection, diagnosis, and recovery
- Observability Adoption: Measurable adoption of common logging, metrics, tracing, alerting, dashboarding, and service ownership standards across critical services
- Reliability Engineering: Expanded use of meaningful SLIs, SLOs, and error budgets to drive service health and engineering priorities
- Signal Quality: Reduction in noisy, duplicate, and non-actionable alerts while improving coverage of critical customer journeys and dependencies
- Tooling & Cost Efficiency: Improved observability platform efficiency through governance, consolidation, telemetry optimization, and reduced tool sprawl
- Cross-functional Adoption: Strong partnership with Product and Engineering teams that translates standards into measurable production adoption
About Us
ICIMS is a leading enterprise hiring platform that combines the scale and reliability of enterprise software with the transformative power of AI. Thousands of organizations across more than 200 countries and territories trust ICIMS to find and hire the people who shape their future and drive their business forward. Powered by insights from billions of hiring interactions, continuous AI innovation, and a highly extensible platform, ICIMS helps organizations turn talent acquisition into a competitive advantage. For more than 25 years, ICIMS has delivered end-to-end hiring solutions that improve recruiting efficiency, reduce costs and create exceptional candidate experiences.
ICIMS helps solve one of the biggest challenges businesses face today: building a workforce that can adapt, scale, and perform in an increasingly competitive and unpredictable talent market. We uniquely do that by combining enterprise-grade hiring technology, AI-powered insights and automation, and connected talent experiences to help organizations improve hiring outcomes while driving measurable impact.
Responsibilities
Technical Leadership
- Provide strategic technical direction for a team of 5+ SRE engineers across one or more geographic regions (US, Ireland, or India)
- Own the technical strategy and roadmap for enterprise observability and reliability capabilities in partnership with SRE, Engineering, Cloud, and Product teams
- Define reference architectures, engineering patterns, and standards that teams can consistently apply in production
- Drive architecture reviews and technical decision-making for complex observability, reliability, scalability, and performance challenges
- Provide hands-on technical mentorship and guidance, raising observability and SRE engineering practices across teams
Incident Management & Response
- Participate in enterprise-wide incident management, ensuring rapid detection, response, restoration, and prevention of recurring issues
- Improve incident detection and triage through actionable telemetry, service health views, dependency context, and well-designed alerting
- Develop and maintain runbooks, emergency response procedures, and operational readiness practices for critical services
- Lead root cause and post-incident reviews, ensuring clear documentation and implementation of durable corrective actions
- Participate in 24/7 on-call and escalation procedures and serve as a senior technical leader with Engineering and Incident Management during critical incidents
Observability Strategy & Standards
- Establish and evolve enterprise standards for logs, metrics, traces, alerting, dashboards, instrumentation, and service ownership
- Champion OpenTelemetry-first, vendor-neutral instrumentation patterns with consistent context, correlation, naming, and metadata across services
- Implement meaningful SLIs, SLOs, error budgets, and service health views that connect technical signals to customer impact
- Drive practical adoption and governance of observability standards, measuring coverage, signal quality, and operational effectiveness across teams
Platform Reliability, Automation & Tooling
- Design and operate scalable observability platforms and telemetry pipelines using technologies such as Grafana, Prometheus, Sumo Logic, New Relic, and cloud-native services
- Lead observability platform modernization, migration, and consolidation while maintaining coverage and controlling ingestion, retention, cardinality, and overall tooling cost
- Use infrastructure-as-code, automation, self-service patterns, and automated remediation to make reliability practices repeatable and reduce operational overhead
- Monitor and optimize multi-cloud infrastructure and core services across AWS, Azure, and GCP for reliability, performance, capacity, and operational efficiency
Qualifications
- Bachelor’s degree in computer science, Engineering, Information Systems, or related technical field
- Equivalent combination of education and experience will be considered
- Cloud certifications (AWS, Azure, or Google Cloud)
Technical Experience
- 8+ years in SRE, DevOps, Infrastructure Engineering, or Observability Engineering roles with 4+ years in senior technical positions
- Proven hands-on experience designing, implementing, and operating observability capabilities at scale in large enterprise SaaS or cloud production environments
- Deep experience across logging, metrics, distributed tracing, alerting, dashboards, and OpenTelemetry, with platforms such as Grafana, Prometheus, Sumo Logic, and New Relic
- Strong multi-cloud and cloud-native experience, including AWS, containers, Kubernetes/ECS, Linux, and distributed application architectures
- Experience designing scalable telemetry pipelines and managing sampling, retention, cardinality, data quality, and cost tradeoffs in high-volume environments
Leadership & Communication
- Proven track record creating technical standards and successfully driving them from architecture into consistent production adoption across engineering teams
- Experience serving as a senior technical leader during critical incidents and complex cross-team reliability initiatives
- Strong communication and influencing skills with engineers, architects, product leaders, and senior stakeholders
- Demonstrated ability to mentor technical teams, build alignment across organizational boundaries, and lead through influence
SRE & Operations
- Demonstrated success implementing SRE principles in large-scale production environments, including practical use of SLIs, SLOs, and error budgets
- Strong background in incident management, root cause analysis, operational readiness, automation, and continuous reliability improvement
- Experience with ITIL frameworks and tools and with establishing service-level expectations for enterprise SaaS products
Preferred
- Experience leading large-scale observability platform migrations, consolidation initiatives, and telemetry cost/governance programs
- Infrastructure-as-code expertise with Terraform or CloudFormation; authentication and identity management systems knowledge is a plus
EEO Statement
iCIMS is a place where everyone belongs. We celebrate diversity and are committed to creating an inclusive environment for all employees. Our approach helps us to build a winning team that represents a variety of backgrounds, perspectives, and abilities. So, regardless of how your diversity expresses itself, you can find a home here at iCIMS. We prohibit discrimination and harassment of any kind based on race, color, religion, national origin, sex (including pregnancy), sexual orientation, gender identity, gender expression, age, veteran status, genetic information, disability, or other applicable legally protected characteristics. If you’d like to request an accommodation due to a disability, please contact us at
[email protected].
Compensation and Benefits
Competitive health and wellness benefits include medical insurance (employee and dependent family members), personal accident and group term life insurance, bonding and parental leave, lifestyle spending account reimbursements, wellness services offerings, sick and casual/emergency days, paid holidays, tuition reimbursement, retirals (PF - employer contribution) and gratuity. Benefits and eligibility may vary by location, role, and tenure. Learn more here: https://careers.icims.com/benefits