Site Reliability Engineer II
PROS Sofia, Sofia-City, Bulgaria
Software Development · 1,001-5,000 employees
Applying here? Try the free cover letter tool — paste this posting and your résumé, no account needed.
About the role
The Site Reliability Engineer II is responsible for monitoring service performance, maintaining infrastructure stability, and implementing reliability enhancements. They also collaborate with development teams to resolve performance bottlenecks and participate in on-call rotations to troubleshoot production incidents.
What they look for
Requirements
Candidates must have a university degree in computer science or a related field and proficiency in at least one high-level programming language like Ruby, Go, or Java. Strong skills in operating systems, networking, database management, and automation tools are required for this role.
Benefits
Full description
The Site Reliability Engineer II is a primary team member who works to administer, support, troubleshoot, and problem solve complex systems and services.
A Day in the Life of the Site Reliability Engineer II:
- Monitor service performance, reliability metrics, and infrastructure stability.
- Perform in-depth analysis of system performance and identify areas for improvement.
- Participate in disaster recovery testing and implement reliability enhancements.
- Define and maintain Service Level Objectives (SLOs) and related visualizations/alerts.
- Collaborate with product teams to resolve performance bottlenecks.
- Implement and maintain automated deployments and self-service tools.
- Create and troubleshoot automation scripts for operational tasks.
- Leverage automation to improve system scalability and efficiency.
- Participate in Follow-the-sun on-call rotations and respond to incidents promptly.
- Troubleshoot and resolve production incidents, identifying root causes and creating detailed post incident reports.
- Work with development teams to address reliability and performance concerns.
- Maintain and update documentation, including user stories and operational processes.
- Share knowledge through team sessions and contribute to continuous improvement.
- Implement automation for security auditing and vulnerability mitigation.
- Collaborate with security teams to enhance cloud security posture.
- Identify root causes of incidents and outages and participate in detailed post-incident analysis and documentation.
Required Qualifications - About you:
- Working knowledge of operating systems, networking and database management.
- Advanced scripting and automation for deployment, scaling and maintenance tasks.
- Proficiency in at least one high-level programming language (Ruby, Go, Java).
- Knowledge of infrastructure and configuration management via automation.
- Advanced skills in creating monitoring and alerting rules (Prometheus, Grafana).
- Implement and optimize Cloud environments.
- Knowledge of RESTful API design and development.
- Familiarity with API testing tools (e.g., Postman).
- Excellent communication skills.
- Excellent time management, organizational skills, crisis management and problem-solving skills.
- Ability to work in a team and independently.
- Willing to innovate, learn and share knowledge.
- University degree in computer science or related.
- Developing and implementing IT security best practices and procedures.
- Excellent command of English language.
Highly Preferred:
- Applicable IT Certifications.
- System administrator experience.
- Previous experience with cloud services - including open-source technology, software development, system engineering, scripting languages and multiple cloud provider environment.
AI Fluency & Growth Mindset- We welcome candidates who:
- Understand core AI concepts and apply them ethically to enhance productivity, insights, and decision-making.
- Craft effective prompts to optimize the quality and relevance of AI-generated outputs.
- Explore and apply agentic AI systems, using or managing autonomous agents to streamline workflows and automate tasks.
- Leverage AI tools to boost efficiency, creativity, and innovation in their daily work.
- Stay curious and adaptable, continuously experimenting with AI-driven solutions to elevate team performance and customer impact.
Why Join PROS?
PROS culture and its extraordinary people are at the core of our success. We are passionate about what we do and relentless in delivering on our promises.
Our commitment to customer success inspires us to think smarter and dream bigger, empowering airlines to achieve more than they ever imagined through intelligent offer and revenue optimization.
At PROS, we foster a culture of care, where people feel supported to grow, innovate, and bring their best selves to work—every day. From flexible ways of working to continuous learning, we empower our teams to thrive both personally and professionally.
Join PROS, a dedicated travel technology company with nearly 40 years of proven airline expertise and a long runway for future growth, now powering the future of AI-driven airline retailing. If you want to be part of something exceptional, help us shape how airlines compete, innovate, and win.
PROS Core Values
- We are Owners
We look for every opportunity to create a better PROS and a better experience for our customers – and we hold ourselves accountable.
- We are Innovators
We think creatively to find new paths to success – for our people, our customers and our business.
- We Care
We are centered on caring for the people, businesses, and communities we serve.
Similar roles
-
Sr. Staff Site Reliability Engineer-Federal, Security Clearance
Zscaler Arlington County, Virginia, United States · $164K–$205K/yr
-
SRE III
JPMorgan Chase & Co. Buenos Aires, Argentina
-
Sr. SRE Lead
Chubb Bogota, Capital District, RAP (Especial) Central, Colombia
-
Site Reliability Engineer II
KFC Europe Irvine, California, United States · $105K–$132K/yr
-
Sr Manager, Site Reliability Engineer (FedRAMP / AWS GovCloud)
BeyondTrust Broken Hill, New South Wales, Australia
-
Site Reliability Engineer
BeyondTrust United States