Sr. Manager, SRE, Operations & Product Support
American Bureau of Shipping Knoxville, Tennessee, United States
Maritime Transportation · 5,001-10,000 employees
About the role
The role involves establishing production reliability and support operating models for a next-generation Fleet Management System. Responsibilities include leading incident response, implementing observability, and automating operational tasks to ensure system scalability and performance.
What they look for
Requirements
Candidates must have a bachelor's degree and at least 10 years of experience in site reliability engineering or cloud operations. Proficiency in Azure cloud services and experience with production SaaS environments are essential requirements.
Benefits
Full description
ABS Group Digital Solutions is seeking a Lead, Site Reliability Engineering, Operations & Product Support to establish the production reliability and support operating model for its next-generation Fleet Management System (FMS). The role will help take FMS from development through beta, customer migration, and scaled SaaS operations. It combines hands-on reliability engineering with leadership of incident response, operational readiness, technical support escalation, and post-launch improvement.
Working with the Platform Engineering & Cloud Architecture lead, application engineering, QA, security, and customer support teams, this person will ensure the product is observable, recoverable, and supportable. The lead will use conventional and AI-assisted automation to streamline operations while establishing clear boundaries between technical SRE ownership and customer-facing support ownership.
What You Will Do:
- Define service health measures and reliability objectives for FMS, including availability, latency, error rates, recovery expectations, and appropriate service-level indicators and objectives.
- Establish end-to-end observability across applications, infrastructure, integrations, and AI-enabled services through actionable logs, metrics, traces, dashboards, health checks, and alerts; assess where AI-assisted anomaly detection can improve signal quality.
- Lead the technical incident-response model, including severity definitions, on-call and escalation practices, incident coordination, recovery procedures, and post-incident reviews; use AI-assisted summarization and evidence gathering where it improves response without replacing human judgment.
- Work with engineering and platform teams to design for resilience, performance, capacity, backup and recovery, and safe operation under expected customer and data growth.
- Define and implement operational-readiness criteria for beta and production releases, including monitoring, runbooks, rollback plans, ownership, support handoffs, and post-release validation.
- Establish the technical support escalation model and partner with customer support, product, and engineering to resolve issues and turn recurring incidents and tickets into permanent fixes; evaluate AI-assisted ticket categorization and knowledge retrieval to speed technical triage.
- Support customer migrations, go-lives, and post-launch stabilization by preparing technical monitoring and response plans, triaging production issues, and incorporating lessons into repeatable procedures.
- Automate routine operational tasks, health checks, deployment verification, incident triage, and recovery; use AI where it demonstrates improved speed or accuracy, with access controls, auditability, and human approval for production-impacting actions.
- Track reliability, incident, supportability, and operational-efficiency trends; communicate risks, corrective actions, and progress to engineering and program leadership, including evidence of whether AI-assisted workflows reduce toil or improve outcomes.
- Help build and mentor an SRE/production-operations capability as FMS moves from initial releases to scaled customer use.
What You Will Need:
Education and Experience
- Bachelor’s degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent relevant experience.
- 10+ years of relevant experience in site reliability engineering, production engineering, cloud operations, or software operations, including experience leading incident response or operational improvement across teams.
- Demonstrated experience operating production web or SaaS services and improving their reliability through software engineering and automation.
- Experience establishing observability, on-call practices, runbooks, and service-health or reliability measures for production systems.
- Experience partnering with software engineering, platform engineering, and customer-facing support teams during releases, incidents, and customer go-lives.
- Experience applying AI-assisted tools or workflows to technical operations, incident triage, monitoring analysis, support knowledge retrieval, or operational automation, with an understanding of how to validate results before use in production.
- Experience with Azure cloud preferred.
Knowledge, Skills, and Abilities
- Strong command of SRE practices, including SLIs/SLOs, incident response, root-cause analysis, performance, capacity, resilience, and disaster recovery.
- Experience operating production SaaS applications on Azure, including compute, networking, identity, storage, containers, databases, integrations, and security.
- Ability to build effective observability and on-call practices using telemetry, logs, metrics, traces, and tools such as Azure Monitor, Application Insights, and Log Analytics-without creating unnecessary alert noise.
- Ability to automate secure deployments and operational workflows using scripting, APIs, CI/CD, infrastructure as code, managed identities, and Key Vault.
- Sound judgment on release risk, rollback, customer impact, and the responsible use of AI-assisted operations, including data protection and human oversight of production-impacting actions.
- Clear communication and collaborative leadership across engineering, support, and business stakeholders, including the ability to drive improvements without direct ownership of every team or system.
Reporting Relationships:
Reports to the Sr Director, Platform Engineering & Cloud Architecture. The role will initially lead cross-functional operational practices; direct-report scope will be determined as the SRE and production-operations capability scales. Customer-facing support teams retain ownership of routine customer communications and frontline support.
About ABS Wavesight
ABS Wavesight is the new ABS Affiliate maritime software as a service (SaaS) company dedicated to helping shipowners and operators streamline compliance while maintaining competitive,more efficient, and sustainable operations. Our mission is to develop world-class software products that improve vessel performance for the health of our seas, environment and self. The ABS Wavesight portfolio is comprised of best-in-class proprietary technology and third-party integrations that offer unparalleled insight into every aspect of a fleet’s operations. Backed by ABS’s 160-year legacy of maritime innovation and experience, our products are collectively installed on more than 5,000 vessels across the global fleet. Learn more about ABS Wavesight by visiting www.abswavesight.com.
About Our Benefits
ABS Wavesight proudly offers a variety of industry-leading benefits designed to enhance the life and well-being of our employees and their families. These benefits include, but are not limited to, medical insurance (PPO and HD), dental and vision insurance, Health Savings account (HSA), Flexible Savings Account (FSA), life insurance, accidental death and dismemberment insurance, disability leave programs, parental leave program, paid holidays, and paid vacation time. The Company provides an Employee Assistance Plan (EAP) that offers additional support in personal wellness, including work-life services. ABS Wavesight also offers a 401K plan with a generous company match, subject to plan requirements.
Equal Opportunity
ABS Wavesight is committed to the equal employment opportunity of its employees and prohibits discrimination against any employee or qualified applicant based on race, color, creed, religion, national origin, sex, gender identity, age, disability, marital status, sexual orientation, citizenship status or veteran status, or other non-work-related characteristics that may be protected under the law of the Federal Government or specific state employment laws.
Notice
ABS and Affiliated Companies (ABS) will not pay a fee to any third-party agency without a valid ABS Master Service Agreement (MSA) authorized and signed by Human Resources. Any resume, CV, application, or other forms of candidate submission provided to any employee of ABS without a valid MSA on file will be considered property of ABS, and no fee will be paid.
Other
This job description is not intended, and should not be construed, to be an all-inclusive list of responsibilities, skills, efforts or working conditions associated with the job of the incumbent. It is intended to be an accurate reflection of the principal job elements essential for making a fair decision regarding the pay structure of the job. #ogjs
Similar roles
-
Site Reliability Engineer
InRule Gothenburg, Västra Götaland County, Sweden
-
Especialista em Engenharia de Plataforma | SRE Experience
C6 Bank São Paulo, São Paulo, Brazil
-
Principal, Staff Site Reliability
DigitalBridge Boca Raton, Florida, United States
-
Senior Manager, I&IT Site Reliability Engineering & Operations
Metrolinx Toronto, Ontario, Canada
-
Staff Site Reliability Engineer (Remote)
KnowBe4 United States · $170K–$210K/yr
-
Site Reliability Engineer
Ordergroove Jyväskylä, Central Finland, Finland