Site Reliability Engineer
Expleo Bucharest, Romania
IT Services and IT Consulting · 10,001+ employees
Applying here? Try the free cover letter tool — paste this posting and your résumé, no account needed.
About the role
You will design, deploy, and operate scalable services while ensuring high availability, performance, and security for an international client. You will also build infrastructure tooling, manage CI/CD pipelines, and collaborate with development teams to maintain platform reliability.
What they look for
Requirements
The role requires at least 3 years of experience in SRE, DevOps, or Linux platform engineering with strong administration and troubleshooting skills. Proficiency in Ansible, scripting languages, and experience with IT service management platforms like ServiceNow are essential.
Benefits
Full description
Overview
Expleo is a global engineering, technology, and consulting service provider that partners with leading organizations to guide them through their business transformation, helping them achieve operational excellence and future-proof their businesses. Expleo benefits from more than 50 years of experience developing complex products in automotive and aerospace, optimizing manufacturing processes, and ensuring the quality of information systems. Leveraging its deep sector knowledge and wide-ranging expertise in fields including AI engineering, digitalization, automation, cybersecurity and data science, the group’s mission is to fast-track innovation through each step of the value chain. With a worldwide presence in 30 countries, our global footprint includes excellence centers around the world, including Romania since 1994. Responsibilities
We are looking for a Site Reliability Engineer to support an international client in the payments and financial technology sector. You will help design, deploy and operate services at scale, ensuring high availability, performance and security. Working closely with development teams, you will build and maintain the infrastructure and tooling that underpin platform reliability and operational excellence throughout the software development lifecycle.
Key responsibilities:
- Develop and improve shared services and tooling to increase delivery speed, availability, scalability and operational efficiency.
- Integrate community and open-source products into the technical ecosystem.
- Build tooling using scripting and programming languages, and collaborate with the team through knowledge sharing.
- Support business teams in migrating existing solutions to new infrastructure using Infrastructure as Code, automation and CI/CD.
- Work across teams to build fast, reliable and resilient production systems.
- Develop, configure and enhance the automation stack and CI/CD pipelines.
- Create and maintain technical and user documentation that colleagues can use, improve and share.
- Perform daily operational activities, including change management, application monitoring and application deployment.
- Embed security into service design and operations from the outset, in line with payment industry requirements.
Qualifications
- At least 3 years of experience in SRE, DevOps, Linux platform engineering or production operations.
- Strong Linux administration and troubleshooting skills, preferably in Red Hat or CentOS environments.
- Hands-on experience monitoring service availability and responding to production incidents.
- Practical experience with PagerDuty or a comparable on-call and alerting platform.
- Practical experience with ServiceNow or a comparable IT service-management platform.
- Strong Ansible experience, including reusable playbooks, roles, inventories and controlled execution.
- Experience creating or maintaining Linux packages and repositories using RPM, YUM, DNF or equivalent tooling.
- Good Git knowledge and experience working with branches, merge requests, reviews and release workflows.
- Scripting ability in Bash, Go (Python) or another language used for operational automation.
- Strong troubleshooting, documentation and communication skills in English.
Desired skills
- Experience with Red Hat Satellite, Foreman, Katello or another repository and lifecycle-management platform.
- Experience integrating and maintaining open-source products in enterprise environments.
- Experience with Prometheus, Grafana, Datadog, ELK, Splunk, New Relic or similar monitoring tools.
- Experience with GitLab CI/CD, Jenkins, GitHub Actions, Azure DevOps or comparable pipelines.
- Exposure to Terraform, OpenTofu, Puppet or other infrastructure and configuration-management tooling.
- Experience operating services across multiple data centers or hybrid infrastructure.
- Knowledge of security hardening, vulnerability remediation or PCI-DSS and other regulated environments.
What do I need before I apply
- Hybrid, a few days per month at the office.
- CIM only
Benefits
- Benefit Platform
- Holiday Voucher
- Private medical insurance
- Performance bonus
- Easter and Christmas bonus
- Employee referral bonus
- Bookster subscription
- Work from home options depending on project
Similar roles
-
Site Reliability Engineer II
Mastercard Dublin, Leinster, Ireland
-
Site Reliability Engineer (MTS) GovCloud 24x7
Salesforce Denver, Colorado, United States · $117K–$177K/yr
-
2027 Technology Intern - Site Reliability Engineering
Charles Schwab Inc. Austin, Texas, United States · $63K–$70K/yr
-
Senior Site Reliability Engineer
Charles Schwab Inc. Southlake, Texas, United States · $130K–$155K/yr
-
Senior Site Reliability Engineer
Mastercard Ciudad de México, Mexico
-
Mainframe SRE
Zensar Gurgaon, Haryana, India