About the role
Lead the Site Reliability Engineering function to define and drive the organization's reliability strategy and roadmap. Manage the end-to-end reliability posture of production systems while mentoring a high-performing SRE team to reduce toil through automation.
What they look for
Requirements
Requires a degree in Computer Science or a related field with over 8 years of experience in software engineering or SRE, including at least 3 years of people leadership. Must be proficient in Python or Go and have expert-level experience with Dynatrace and cloud platforms like Azure.
Full description
Key Objective:
- Lead the Site Reliability Engineering function to define and drive the organisation’s reliability engineering strategy — bridging software development and operations through engineering discipline, not manual process.
- Own the end-to-end reliability posture of production systems: define SLO/SLI frameworks, govern error budgets, and enforce production-readiness standards to protect business continuity.
- Build, mentor, and scale a high-performing SRE team that prioritises engineering over toil — automating manual work, embedding reliability into the SDLC, and driving down mean time to recovery through systematic improvement.
- Champion observability-led engineering through full-stack Dynatrace adoption, AIOps integration, and data-driven reliability decision-making at every layer of the stack.
- Serve as the primary reliability engineering partner to development and platform leadership, shaping architecture decisions, release policies, and automation strategy.
Key Responsibilities:
- Define and drive the SRE strategy and multi-year roadmap aligned to business priorities.
- Lead and develop the SRE team, including hiring, onboarding, performance management, career development, and succession planning.
- Own incident management, including severity classification, escalation, response SLAs, and leadership of major incidents.
- Champion blameless postmortems, root cause analysis, and implementation of systemic fixes.
- Establish and govern SLOs, SLIs, and error budgets, ensuring reliability targets are aligned to business needs.
- Drive resilience engineering, including chaos engineering, GameDays, production readiness reviews, and failure mode analysis.
- Reduce toil through automation, improved runbooks, and continuous operational improvement.
- Own observability and alerting standards, including monitoring strategy, dashboards, and alert quality.
- Partner with engineering, architecture, product, and leadership teams to embed reliability into design and delivery.
- Represent the SRE function in senior forums and provide reporting on reliability, risk, and operational performance.
Qualifications
Qualifications:
- Degree in Computer Science, Software Engineering, IT, or a related technical field.
- 8+ years’ experience in software engineering, platform reliability, or SRE, including 3+ years in people leadership.
- Strong hands-on coding ability in Python, Go, or similar, with experience building automation and self-healing solutions.
- Proven experience leading enterprise-scale SRE or platform reliability functions.
- Experience defining and operating RTO/RPO and SLI/SLO frameworks.
- Strong background in observability, production readiness, error budgets, and chaos engineering.
- Experience leading on-call models, incident response, and executive stakeholder engagement.
- Solid understanding of SDLC, Agile, and DevOps delivery models.
- ITIL Foundation is desirable.
Managerial & Soft Skills:
- Proven people leader with experience building and coaching high-performing teams.
- Strategic thinker who can turn business priorities into reliability roadmaps.
- Strong communicator who can explain technical risk in business terms.
- Calm and decisive during major incidents.
- Influential partner across engineering, product, and leadership teams.
- Strong advocate for developer experience and sustainable on-call practices.
- Data-driven and able to balance reliability, speed, and cost.
- Champions psychological safety, continuous learning, and operational excellence.
Technical Skills:
- Expert in Dynatrace, with experience in observability, monitoring, SLOs, tracing, and log management.
- Proficient in Grafana, Prometheus, Splunk, ELK, Azure Monitor, and Log Analytics.
- Strong knowledge of OpenTelemetry and telemetry pipeline design.
- Experience with ServiceNow, CI/CD tools, Kubernetes, Docker, Terraform, and Bicep.
- Familiar with Java, .NET, databases, APIs, Kafka, and cloud platforms, especially Azure.
- Experience with AIOps, AI-assisted triage, and automation tooling.
- Able to support reliability engineering through scripting, auto-remediation, and operational automation.
Desired:
- Experience in insurance or financial services.
- Dynatrace, Azure, ITIL 4, or Google Cloud/SRE-related certifications.
- Experience with chaos engineering, AIOps, MLOps, and FinOps.
- Strong analytical skills and experience working with large operational datasets.
Similar roles
-
Site Reliability Engineer III (DBA)
Backblaze External Website United States · $125K–$150K/yr
-
Site Reliability Engineer II (AI Platform)
OpenTable Toronto, Ontario, Canada · CA$110K–CA$130K/yr
-
Site Reliability Engineer
Apple Hyderabad, Telangana, India
-
Senior Site Reliability Engineer (SRE) – Application Observability & Readiness (Azure)
Encora Perímetro Urbano Santiago de Cali, Valle del Cauca, Colombia
-
Senior Site Reliability Engineer
Salesforce Dublin, Leinster, Ireland
-
Sr Staff Site Reliability Engineer, AI Infrastructure
d-Matrix Santa Clara, California, United States · $175K–$265K/yr