Lead Site Reliability Engineer
Sherwin-Williams Cleveland, Ohio, United States
Paint, Coating, and Adhesive Manufacturing · 10,001+ employees
Applying here? Try the free cover letter tool — paste this posting and your résumé, no account needed.
About the role
The Lead Site Reliability Engineer optimizes IT products and services by developing automation solutions and monitoring system health metrics. They also lead cross-functional teams to improve application reliability, performance, and operational readiness while mentoring staff on engineering best practices.
What they look for
Requirements
Candidates must have a Bachelor’s degree in Computer Science or Information Systems, or at least 9 years of relevant experience. A minimum of 6 years of experience in Site Reliability Engineering or a related technical discipline is required.
Benefits
Full description
The Lead Site Reliability Engineer role is responsible for optimizing the organization's IT products, services, systems, and digital products. The incumbent works to develop customized solutions to automate the administration, monitoring, and operation of business-critical applications and services, supervises tests for resiliency, redundancy, and failover to ensure uptime, and troubleshoots potential issues to ensure that IT products, services, systems, and digital products are running efficiently and effectively. In addition, the role is responsible for leading the design and implementation of scalable and reliable application and service solutions that can run across multiple environments and technologies. The incumbent fosters a culture of collaboration between cross-functional departments, including development teams, infrastructure teams, and support organizations, to enhance and improve system operability and provide training and knowledge transfer to team members. The role is also responsible for implementing leading practices and emerging technologies that will drive additional efficiencies across the IT organization. The role is also responsible for overseeing project planning, cost analysis, and vendor comparisons when assessing potential solutions and implementing technology and operational improvements.
Responsibilities
WHAT THE ROLE WILL DO:
- Optimize IT products, services, systems, and digital products by proactively analyzing application, service, and operational health metrics to identify potential inefficiencies and thereby ensuring improvements in performance, availability, and reliability.
- Develop customized solutions to automate the deployment, administration, monitoring, and operation of applications and services and train other team members on the use of automation tools.
- Develop a formal process for continuously reviewing and monitoring system SLIs, SLOs, SLAs, and OKRs and build optimization plans to address areas of improvement.
- Foster a culture of collaboration between cross-functional teams to ensure improvement in IT products, services, systems, and digital products.
- Optimize the use of applications, services, integrations, observability tools, infrastructure components, and load-balancing technologies by tracking performance, identifying potential issues, and ensuring optimal operation.
- Lead efforts to continuously improve application and service reliability, performance, resiliency, observability, and operational readiness and take steps to mitigate potential issues.
- Evaluate application and service requirements, lead cross-functional implementation teams, and conduct post-implementation reviews to share lessons learned from the project.
- Create detailed implementation plans for the integration of new technologies, products, and services into the existing environment that will improve application and service resilience, performance, and reduce costs.
- Provide leadership and knowledge-sharing to mentor engineers on reliability engineering, observability, automation, incident management, and operational excellence practices and assess their effectiveness.
This position is not hybrid/remote and will be located at our Cleveland Headquarters office.
This position is not eligible for sponsorship for work authorization now or in the future, including conversion to H1-B visa.
Job duties include contact with other employees and access confidential and proprietary information and/or other items of value, and such access may be supervised or unsupervised. The Company therefore has determined that a review of criminal history is necessary to protect the business and its operations and reputation and is necessary to protect the safety of the Company’s staff, employees, and business relationships.
Must be eighteen years or older
Qualifications
Required Qualifications:
- Bachelor’s degree in Computer Science or Information Systems, or in lieu of a degree, at least 9 years of experience in the field of site reliability engineering
- 6+ years of experience in Site Reliability Engineering, Software Engineering, DevOps Engineering, Platform Engineering, Systems Engineering, or a related technical discipline.
- Experience supporting and operating production applications, services, or enterprise technology platforms.
- Experience with monitoring, logging, observability, and incident response practices.
- Experience utilizing automation, scripting, or tooling to improve reliability, operational efficiency, and supportability.
- Experience collaborating with cross-functional teams to troubleshoot, resolve, and prevent production issues.
- Strong analytical and problem-solving skills
- Must be at least (18) eighteen years of age
- Must be legally authorized to work in the country of employment without now or in the future requiring sponsorship for employment visa status (e.g., OPT, CPT, H1B, EB-1, etc.)
•
Preferred Qualifications:
- Experience supporting customer-facing or business-critical digital products and services at a global level.
- Experience establishing and managing Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs).
- Experience leading major incident response, root cause analysis, and post-incident reviews.
- Experience with cloud-native architecture (Microsoft Azure, AWS, Kubernetes, etc.)
- Experience implementing observability, alerting, and operational excellence practices.
- Experience with CI/CD pipelines and Infrastructure as Code (IaC).
- Experience driving reliability, resiliency, and performance improvements in distributed systems.
- Experience with leadership and executive-level communication.
- Microsoft Certified: Azure Solutions Architect Expert
- Relevant Site Reliability Engineering, Azure, Cloud, DevOps, or Platform Engineering certifications
Technical Skills:
- Monitoring and Logging
- Automation and Scripting
- Reliability Engineering
- Incident Management
- Application and Service Operations
- Performance and Capacity Management
- Microsoft Azure
- Kubernetes
- Application Performance Monitoring
- Observability and Telemetry
- Infrastructure as Code
- Distributed Systems
- API and Integration Technologies
- DevOps Practices
At Sherwin-Williams, our purpose is to inspire and improve the world by coloring and protecting what matters. Our paints, coatings and innovative solutions make the places and spaces in our world brighter and stronger. Your skills, talent and passion make it possible to live this purpose, and for customers and our business to achieve great results. Sherwin-Williams is a place that takes its stability, growth and momentum and translates it to possibility for our people. Our people are behind the strength of our success, and we invest and support you in:
Life … with rewards, benefits and the flexibility to enhance your health and well-being Career … with opportunities to learn, develop new skills and grow your contribution Connection … with an inclusive team and commitment to our own and broader communities It's all here for you... let's Create Your Possible
At Sherwin-Williams, part of our mission is to help our employees and their families live healthier, save smarter and feel better. This starts with a wide range of world-class benefits designed for you. From retirement to health care, from total well-being to your daily commute—it matters to us. A general description of benefits offered can be found at http://www.myswbenefits.com/. Click on “Candidates” to view benefit offerings that you may be eligible for if you are hired as a Sherwin-Williams employee.
Compensation decisions are dependent on the facts and circumstances of each case and will impact where actual compensation may fall within the stated wage range. The wage range listed for this role takes into account the wide range of factors considered in making compensation decisions including skill sets; experience and training; licensure and certifications; and other business and organizational needs. The disclosed range estimate has not been adjusted for the applicable geographic differential associated with the location at which the position may be filled. The wage range, other compensation, and benefits information listed is accurate as of the date of this posting. The Company reserves the right to modify this information at any time, with or without notice, subject to applicable law.
Qualified applicants with arrest or conviction records will be considered for employment in accordance with applicable federal, state, and local laws including with the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act where applicable.
Sherwin-Williams is proud to be an Equal Employment Opportunity employer. All qualified candidates will receive consideration for employment and will not be discriminated against based on race, color, religion, sex, sexual orientation, gender identity, national origin, protected veteran status, disability, age, pregnancy, genetic information, creed, marital status or any other consideration prohibited by law or by contract.
As a VEVRAA Federal Contractor, Sherwin-Williams requests state and local employment services delivery systems to provide priority referral of Protected Veterans.
Please be aware, Sherwin-Williams recruiting team members will never request a candidate to provide a payment, ask for financial information, or sensitive personal information like national identification numbers, date of birth, or bank account numbers during the application process.
Similar roles
-
Senior Site Reliability Engineer (SRE)
LeoLabs, Inc. $171K–$192K/yr
-
Senior Site Reliability Engineer
2K Austin, Texas, United States
-
Senior Site Reliability Engineer (SRE)
Tradeweb United States · $170K–$210K/yr
-
Senior Software Engineer, Site Reliability Engineering
Google New York, New York, United States · $174K–$252K/yr
-
Software Developer III, Site Reliability
Google Waterloo, Ontario, Canada · CA$150K–CA$153K/yr
-
Senior Site Reliability Engineer - 2394503
UnitedHealth Group Eden Prairie, Minnesota, United States · $150K–$184K/yr