METRO/MAKRO

DevOps Engineer/SRE- Digital Commerce

METRO/MAKRO Bucureşti, Bucharest, Romania

Wholesale · 5,001-10,000 employees

9 h ago
devops Mid (2-5 yrs) Full-time Romania
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The DevOps Engineer/SRE will ensure the stability and reliability of cloud-native applications on GCP while automating infrastructure provisioning and operational tasks. They will also define SLOs, manage database systems, and collaborate with development teams to integrate reliability practices into CI/CD pipelines.

What they look for

Site Reliability Engineering DevOps Google Cloud Platform Kubernetes Docker Terraform Helm Kustomize Datadog PostgreSQL Cassandra CI/CD GitHub Actions Linux Bash Infrastructure as Code

Requirements

Candidates must have 4+ years of experience in SRE or DevOps with strong expertise in cloud-native technologies, Kubernetes, and infrastructure as code. Proficiency in scripting, monitoring tools, and troubleshooting complex distributed systems is essential for this role.

Benefits

Hybrid work model Agile work environment Professional development programs Individual training Leadership development Work-life balance support

Full description

Company Description

About us:   

Passion for food. Hunger for tech. We make METRO digital.   

Today technology is driving the world. And at METRO.digital we are driving the technology for one of the leading international wholesalers specializing in food - METRO. From e-commerce to checkout, to delivery software, we work on a wide range of products to make each day a success for our customers and colleagues. With passion and ownership, we build the future of wholesale.    

  You are driving to create smart solutions for customers around the globe? You want to grow in a flexible environment? Let the right career opportunity find you and join us!   

How you will make an impact? 

We are seeking a DevOps Engineer/SRE for Evaluate Pipeline & API system, the backbone of Digital Commerce platform.

We are consolidating and streamlining METROs article and assortment data to be consumed by other solutions along with reliably building  the business logic for digital commerce solutions and provide METRO's article and assortment data for internal consumers.Our new colleague will need hands-on SRE/DevOps, will take ownership of a critical GCP platform, will troubleshoot and solve infrastructure problems, will automate operational work, and will be confident making technical decisions.  

The ideal candidate will have hands-on expertise in cloud-native technologies, infrastructure as code, observability, and automation. 

 Please note that this role is open also for candidates in Brasov and Cluj. 

Job Description

Key Responsibilities: 

  • Ensure the stability and reliability of cloud-native applications deployed on GCP containerized with Docker and orchestrated via Kubernetes; 
  • Define, implement, and monitor SLOs, SLAs, and SLIs to measure system performance and user experience;
  • Automate infrastructure provisioning using Terraform and manage Kubernetes configurations with Kustomize and Helm;
  • Develop and maintain monitoring and alerting systems using Datadog and GCP-native tools; 
  • Conduct incident analysis and postmortems to drive continuous improvement;
  • Collaborate with development teams to integrate reliability practices into CI/CD pipelines using GitHub Actions;
  • Manage and troubleshoot database systems, particularly PostgreSQL and Cassandra;
  • Apply networking knowledge and Linux system administration skills to troubleshoot and optimize system connectivity and performance. 

Qualifications

Required key competencies and qualifications:   

  • 4+ years of experience in Site Reliability Engineering/DevOps; 
  • Proven experience designing and operating elastic, resilient systems in cloud environments; 
  • Strong understanding of GCP/AWS/Azure, Kubernetes, and container orchestration; 
  • Experience with cloud platforms (Google Cloud Platform - ideally, Amazon Web Services, Microsoft Azure), Kubernetes, and container orchestration; 
  • Proficiency in infrastructure as code and configuration management tools (Terraform, Helm, Kustomize); 
  • Experience with monitoring and observability tools (Datadog, GCP Monitoring); 
  • Solid scripting skills in bash and familiarity with automation frameworks; 
  • Experience with CI/CD pipelines, especially using GitHub Actions; 
  • Familiarity with networking fundamentals and troubleshooting; 
  • Strong coding skills and ability to develop reliability-focused tooling; 
  • Excellent communication skills in English (written and spoken); 
  • Strong problem-solving skills and a process-oriented mindset; 
  • Ability to work independently and collaboratively in a fast-paced environment; 
  • Passion for clean code, automation, and continuous improvement. 

Nice-to-Have: 

  • Familiarity with monitoring tools (e.g., DataDog, Prometheus, GCP Monitoring); 
  • Experience working in Agile/Scrum teams. 

Additional Information

What we offer at METRO.digital?   

  • Hybrid and agile work: thrive in a flexible, multicultural environment.  

At METRO.digital, we promote work-life balance through a hybrid working model. You’ll be part of self-organizing, multicultural teams that collaborate in an agile setup.  

  • People development: when you grow so do we!  

We want you to become the best version of yourself with individual and company-wide programs and trainings for people development. Focused among other on development, leadership, appreciation ... it´s time to upskill your career.  

  • Support with individual solutions: we are people-caring!  

We offer support whenever you need it - at every stage of your professional journey.  

Want to know more about all our benefits? Discover more here.   

Let´s connect soon. Apply for the role now!     

Position grade within our career framework: Site Reliability Engineer Grade 3 (Md7).

Similar roles