About the role
The role involves providing reliability coverage and facilitating handoffs between global shifts for restaurant technology. It focuses on transforming reliability practices from reactive incident response to a platform engineering model using automation and self-service tools.
What they look for
Requirements
Candidates must have 5+ years of experience in site reliability engineering or production operations with proficiency in cloud providers and observability tooling. Strong communication skills and technical depth in automation, Kubernetes, and infrastructure as code are required.
Full description
- 2+ years of experience in SRE, DevOps, production support, or infrastructure engineering roles
- Hands-on experience with monitoring and observability tooling (e.g., Datadog, Prometheus, Grafana, CloudWatch, or similar)
- Working knowledge of at least one major cloud provider (AWS preferred)
- Proficiency in at least one scripting or programming language (e.g., Python, Bash, Go) for automation, with demonstrated examples of automating away manual operational work
- Experience participating in incident response and on-call or shift-based operations
- Understanding of SLI/SLO concepts and reliability engineering fundamentals
- Ability to work follow-the-sun shift rotations, including structured handoffs with teams in other regions
- Strong written and verbal English communication skills for cross-region collaboration
Responsibilities
- 2+ years of experience in SRE, DevOps, production support, or infrastructure engineering roles
- Hands-on experience with monitoring and observability tooling (e.g., Datadog, Prometheus, Grafana, CloudWatch, or similar)
- Working knowledge of at least one major cloud provider (AWS preferred)
- Proficiency in at least one scripting or programming language (e.g., Python, Bash, Go) for automation, with demonstrated examples of automating away manual operational work
- Experience participating in incident response and on-call or shift-based operations
- Understanding of SLI/SLO concepts and reliability engineering fundamentals
- Ability to work follow-the-sun shift rotations, including structured handoffs with teams in other regions
- Strong written and verbal English communication skills for cross-region collaboration
Qualifications
- Kubernetes, container orchestration, and infrastructure as code experience (e.g., Terraform)
- Familiarity with AI-assisted operations tooling and automation-first reliability approaches, including auto-healing and auto-remediation patterns
- Exposure to platform engineering and internal developer platform concepts: self-service tooling, developer portals (e.g., Port, Backstage), GitOps
- Experience in multi-region or globally distributed team models
- Relevant certifications (AWS, CKA, or similar)
Similar roles
-
Site Reliability Engineer
Cover Genius Sydney, New South Wales, Australia
-
Manager - Site Reliability Engineer
Frontier Airlines Denver, Colorado, United States · $123K–$164K/yr
-
Staff Site Reliability Engineer
Okta Bengaluru, Karnataka, India
-
Staff Site Reliability Engineer
Veeam Software United States · $172K–$442K/yr
-
Senior Site Reliability Engineer
TailorCare United States
-
Fielded Site Reliability Engineer
Anduril Industries Waltham, Massachusetts, United States · $112K–$149K/yr