About the role
The Site Reliability Engineer II will manage incident response, observability, and automation while providing technical guidance to engineering teams. They will operate within shift-based or follow-the-sun models to ensure system reliability and support cross-region collaboration.
What they look for
Requirements
Candidates must have at least 5 years of experience in SRE, DevOps, or production operations and proficiency in cloud platforms and observability tools. Strong scripting skills and an understanding of reliability engineering fundamentals are required to support global infrastructure.
Full description
- 2+ years of experience in SRE, DevOps, production support, or infrastructure engineering roles
- Hands-on experience with monitoring and observability tooling (e.g., Datadog, Prometheus, Grafana, CloudWatch, or similar)
- Working knowledge of at least one major cloud provider (AWS preferred)
- Proficiency in at least one scripting or programming language (e.g., Python, Bash, Go) for automation, with demonstrated examples of automating away manual operational work
- Experience participating in incident response and on-call or shift-based operations
- Understanding of SLI/SLO concepts and reliability engineering fundamentals
- Ability to work follow-the-sun shift rotations, including structured handoffs with teams in other regions
- Strong written and verbal English communication skills for cross-region collaboration
Responsibilities
- Own the Vietnam shift within GRE’s global follow-the-sun coverage model, including schedule design, coverage planning, and holiday/leave management
- Ensure clean, structured handoffs to and from US and India teams, with clear ownership transfer on open incidents and in-flight work
- Maintain shift readiness: runbooks current, alerts actionable, escalation paths clear
- Serve as escalation point for the Vietnam shift during complex or high-severity incidents
Qualifications
- Kubernetes, container orchestration, and infrastructure as code experience (e.g., Terraform)
- Familiarity with AI-assisted operations tooling and automation-first reliability approaches, including auto-healing and auto-remediation patterns
- Exposure to platform engineering and internal developer platform concepts: self-service tooling, developer portals (e.g., Port, Backstage), GitOps
- Experience in multi-region or globally distributed team models
- Relevant certifications (AWS, CKA, or similar)
Similar roles
-
Senior Manager, Site Reliability Engineering
Invoca Los Angeles, California, United States · $190K–$250K/yr
-
Snr. Site Reliability Engineer (Remote in Brazil)
KnowBe4 São Paulo, São Paulo, Brazil
-
Staff Site Reliability Engineer (Remote in Brazil)
KnowBe4 São Paulo, São Paulo, Brazil
-
Snr. Site Reliability Engineer (Remote)
KnowBe4 United States · $130K–$155K/yr
-
Senior Site Reliability Engineer
ADT Whitpain Township, Pennsylvania, United States · $129K–$193K/yr
-
Sr Cloud Platform & Site Reliability Engineering Lead
National Life Insurance Company Addison, Texas, United States · $137K–$201K/yr