Observability Engineer - Site Reliability Engineer
IT Services and IT Consulting · 501-1,000 employees
Applying here? Try the free cover letter tool — paste this posting and your résumé, no account needed.
About the role
Design, implement, and maintain comprehensive observability solutions including metrics, logs, and traces across the organization. Collaborate with cross-functional teams to embed observability by design while optimizing operational costs and system performance.
What they look for
Requirements
Requires 3+ years of experience in SRE or Observability Engineering with proficiency in Kubernetes, Terraform, and the Grafana stack. Candidates must possess a product-oriented mindset, strong automation skills, and fluency in English.
Full description
This is a remote position.
We are looking for a Site Reliability Engineer (SRE) with strong expertise in Observability Engineering to join our team. This role is pivotal to ensuring the reliability, visibility, and performance of our platforms and services. The ideal candidate will have hands-on experience with the Grafana Stack (Tempo, Loki, Mimir, Alloy), knowledge in Java development, a strong SRE mindset, and a passion for automation, scalability, and ownership.
You’ll be joining a motivated, cross-functional team responsible for implementing and scaling our new observability stack that is being built to be a platform for observability for the whole company. Your contributions will directly impact system performance, user experience, and operational cost efficiency.
Requirements
Responsibilities
· Design, implement, and maintain observability solutions covering metrics, logs, traces, and RUM.
· Work with tools such as Grafana Cloud, Tempo, Loki, Mimir, Alloy, and OpenTelemetry.
· Build reliable alerting and monitoring pipelines based on SLOs/SLAs, focusing on low-maintenance automation.
· Ensure the health and integrity of observability data flows from instrumentation to dashboards.
· Collaborate with development and operations teams to embed observability by design into the software lifecycle.
· Define and promote best practices and standards for observability across the organization.
· Support the modernization of observability by replacing and evolving legacy monitoring and alerting solutions.
· Monitor observability-related costs and contribute to FinOps efforts by identifying optimization opportunities.
Requirements
Must-have:
· 3+ years of experience as an SRE, Observability Engineer, or equivalent role.
· Practical experience with OpenTelemetry, or similar instrumentation tools.
· Experience in Kubernetes, Helm, Terraform, and ArgoCD.
· Experience designing and managing telemetry pipelines (metrics/logs/traces), exporters, and sidecars.
· Product-oriented mindset with a bias for automation and a “you build it, you run it” culture
· Fluency in English.
Nice-to-have:
· Knowledge of APM and distributed tracing solutions.
· Experience with FinOps practices applied to observability.
· Hands-on involvement in replacing legacy monitoring stacks.
· Experience with Cloud environments (Azure preferred)
· Contributions to open-source observability tools.
· Knowledge in Java development and applications instrumentation
· Expertise in performance monitoring, alerting, dashboarding, and root cause analysis.
Similar roles
-
Software Developer III, Site Reliability
Google Waterloo, Ontario, Canada · CA$150K–CA$153K/yr
-
Senior Software Engineer, Site Reliability Engineering
Google New York, New York, United States · $174K–$252K/yr
-
Senior Site Reliability Engineer, Platform & Reliability (M/F/X)
Crossbeam Paris, Ile-de-France, France
-
Site Reliability Engineer (SRE) – Building X
Siemens Pune, Maharashtra, India
-
Head of Infrastructure SRE in China
Apple Shanghai, Shanghai, China
-
Site Reliability Engineering Intern (Brno, Czech Republic)
Red Hat Brno, Southeast, Czechia