JPMorgan Chase & Co.

Sr Lead SRE - Reliability Engineering & Problem Management

JPMorgan Chase & Co. Hyderabad, Telangana, India

Financial Services · 10,001+ employees

6 h ago
sre Senior (5-10 yrs) Full-time India
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will serve as the technical authority for post-incident investigations and lead deep-dive root cause analysis sessions to ensure systemic improvements. Additionally, you will drive reliability maturity across platforms by governing service level objectives and integrating AI-assisted workflows into the incident lifecycle.

What they look for

Site Reliability Engineering Problem Management Root Cause Analysis Incident Management Cloud Infrastructure Distributed Systems AIOps SLO/SLI Engineering Network Engineering DevOps Toolchains Observability Automation System Architecture Data Analytics Executive Communication

Requirements

Candidates must have 5+ years of applied experience in Site Reliability, Infrastructure, or Software Engineering with a strong background in conducting technical root cause analysis. Expertise in cloud infrastructure, distributed systems, and enterprise monitoring tools is required, along with the ability to influence stakeholders and communicate technical findings to executives.

Full description

As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Reliability Engineering & Problem Management team, you will serve as the technical authority for post-incident investigations and operational resilience. You will lead deep technical reviews, validate causal analysis, and ensure corrective actions address systemic causes. You will partner with cross-functional teams to drive measurable improvements in reliability, resilience, and operational excellence. You will help ensure incidents translate into lasting engineering improvements and reduction of repeat failures.

Job responsibilities

  • Serve as the technical authority for Problem Management-led Root Cause Analysis reviews across major incidents and service-impacting events
  • Lead deep-dive RCA challenge sessions, validating technical findings and causal chains through evidence-based analysis. Assess the quality and accuracy of root cause investigations, ensuring conclusions are technically sound and defensible. Challenge assumptions, unsupported conclusions, symptom-based findings, and ineffective corrective actions. Drive a culture of accountability focused on systemic learning and long-term reliability improvements
  • Evaluate detection gaps, monitoring effectiveness, observability shortcomings, automation opportunities, resilience weaknesses, process breakdowns, and human factors. Validate that corrective actions address the true root cause and reduce recurrence likelihood and impact. Review corrective actions for closure and effectiveness, ensuring intended reliability and operational outcomes. Provide technical challenge and independent review of vendor, third-party, and internal investigation reports
  • Use enterprise-authorized AI capabilities to accelerate reliability design and operational decisioning, validating outputs and handling operational data according to sensitivity and security requirements
  • Lead reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices, ensuring traceability, auditability, resiliency, and security controls. Define and govern Service Level Objectives, Service Level Indicators, Reliability Metrics, and Error Budgets
  • Evaluate system architecture against reliability design principles
  • Recommend resilience patterns including graceful degradation, dependency isolation, rate limiting, circuit breakers, fault tolerance, capacity management, auto-remediation, and self-healing capabilities
  • Lead reliability maturity assessments across platforms and services
  • Lead enterprise-authorized AI capabilities for RCA generation, incident analysis, log analytics, pattern discovery, problem trend analysis, and corrective action recommendations
  • Establish governance for explainability, auditability, data handling, and validation of AI-generated findings
  • Drive AI-assisted reliability workflows across the incident lifecycle

Required qualifications, capabilities and skills

  • Formal training or certification on security engineering concepts and 5+ years applied experience. Hands on experience in Infrastructure Engineering, Site Reliability Engineering, Production Engineering, Systems Engineering, or Software Engineering. Strong exposure to leading or supporting critical incident investigations
  • Experience conducting deep technical RCAs for enterprise-scale environments, working in highly regulated and mission-critical environments
  • Advanced expertise in Network Engineering, Cloud Infrastructure, Linux/Windows Platforms, Middleware Technologies, Database Technologies, Storage Platforms, Application Architecture, DevOps Toolchains, Distributed Systems, and Enterprise Monitoring Platforms
  • Strong hands-on expertise in SLO/SLI Engineering, Distributed Tracing, Telemetry Design, Reliability Metrics, Error Budget Management, and AIOps Platforms
  • Experience with tools such as Splunk, Dynatrace, Grafana, Datadog, Prometheus, AppDynamics, Elastic, and Open Telemetry
  • Deep understanding of Root Cause Analysis methodologies, Five Whys, Fault Tree Analysis, Event Correlation, Human Factors Analysis, Systemic Cause Analysis, Problem Management Governance, and Major Incident Management
  • Demonstrated experience using enterprise-authorized AI capabilities to improve reliability engineering workflows with strong validation habits and awareness of data sensitivity
  • Ability to set team practices for safe AI usage in operations while maintaining resiliency, security, and auditability outcomes
  • Strong executive communication skills. Ability to challenge senior engineering stakeholders constructively. Proven ability to influence without direct authority
  • Ability to translate technical findings into executive-ready narratives

Preferred qualifications, capabilities and skills

  • Experience leading reliability engineering initiatives in large-scale, complex environments
  • Expertise in AI-enabled incident analysis and reliability workflows
  • Experience establishing governance for explainability and auditability of AI-generated findings
  • Demonstrated ability to drive measurable improvements in reliability, resilience, and operational excellence

Similar roles