Senior Manager, Site Reliability Engineering - Cloud Infrastructure
Futurex Bulverde, Texas, United States
Computer and Network Security · 51-200 employees
Applying here? Try the free cover letter tool — paste this posting and your résumé, no account needed.
About the role
The Senior Manager will lead a global SRE team to ensure the availability, security, and operational excellence of Futurex Cloud Infrastructure. Responsibilities include managing incident response, defining service level objectives, and overseeing infrastructure operations across multiple regions.
What they look for
Requirements
Candidates must have over 10 years of experience in SRE or DevOps roles, including at least 5 years in a management capacity. A bachelor's degree in Computer Science or Engineering is required, along with hands-on expertise in cloud platforms and compliance frameworks.
Benefits
Full description
About Futurex
Futurex is a Texas-based leader in hardware security modules (HSMs), enterprise key management, and data protection, trusted by global financial institutions, payment processors, and cloud providers for more than four decades. Futurex Cloud Infrastructure delivers secure, highly available services to customers around the world.
Role overview
The Senior Manager, Site Reliability Engineering – Cloud Infrastructure owns the availability, security, and operational excellence of Futurex Cloud Infrastructure worldwide. You will lead a globally distributed SRE team that runs production infrastructure across multiple regions, cloud providers, and datacenters.
This role sits at the intersection of cloud operations, security, and compliance. Our customers depend on services that cannot go down, so reliability here is measured in customer trust as much as uptime.
You will report to executive leadership and partner closely with Engineering, Product, Security, Compliance, and Customer Support.
Key responsibilities
Team leadership
• Lead, hire, and develop a global SRE team across multiple time zones, with clear ownership and follow-the-sun coverage.
• Own the global on-call program, including rotation design, response time standards, escalation paths, and coverage across all regions.
• Hold the team accountable for incident response, remediation follow-through, and adherence to operational standards.
• Set the SRE roadmap, staffing plan, and operating budget.
Reliability and service levels
• Define and own SLIs, SLOs, and error budgets for Cloud Infrastructure services, and use them to balance stability against release velocity.
• Partner with Sales and Legal so customer SLA commitments are consistently met.
• Track SLA compliance and error budget consumption across all services.
Operational reporting
• Deliver regular reporting to executive leadership on uptime, SLA performance, incident trends, and on-call metrics such as response times and escalation volume.
• Maintain dashboards that give leadership and stakeholders real-time visibility into service health.
• Produce customer-facing incident reports and root cause analyses for significant events.
Incident management
• Own the incident response process, including severity definitions, escalation paths, and customer communications.
• Lead postmortems and drive remediation items to closure.
• Serve as the senior escalation point during major incidents.
Infrastructure and operations
• Review monitoring, observability, and alerting standards.
• Maintain and regularly test disaster recovery and business continuity plans against defined RTO and RPO targets.
Security and compliance
• Operate within the controls required by industry security and compliance frameworks.
• Oversee secure operational procedures, including dual control and chain of custody where required.
• Enforce strict tenant isolation, least-privilege access, and change management.
• Support internal and external audits with accurate operational evidence.
Required qualifications
• 10+ years in SRE, DevOps, or production infrastructure roles, including 5+ years managing engineering teams.
• Experience leading geographically distributed teams across multiple time zones.
• Proven ownership of a production SaaS or cloud platform with contractual SLAs and 24x7 operations.
• Hands-on depth in at least one major cloud provider (AWS, GCP, or Azure), with working knowledge of the others.
• Strong background in Linux, networking, infrastructure as code (Terraform or similar), and CI/CD.
• Experience operating under compliance frameworks such as PCI DSS, SOC 2, or ISO 27001.
• A track record of building incident management, observability, and DR programs.
• Clear written and verbal communication with executives, customers, and auditors.
• Bachelor's degree in Computer Science, Engineering, or equivalent experience.
Preferred qualifications
• Experience operating security-sensitive or cryptographic infrastructure.
• Familiarity with security certification standards such as FIPS 140.
• Experience running hybrid environments that combine public cloud with physical datacenter hardware.
• Experience leading datacenter migrations or regional expansions.
• Background in financial services or another highly regulated industry.
- Health, dental, vision, life, and short/long-term disability insurance
- Paid vacation, holidays, and sick leave
- Competitive compensation and opportunities for advancement
- Retirement plan with employer contribution match
- Welcoming, family-style corporate culture uniquely suited to fast-paced, entrepreneurial, and motivated individuals
- One of San Antonio’s “Best Places to Work” for nine consecutive years
Similar roles
-
Staff Software Engineer, Site Reliability Engineering, Raxium
Google Fremont, California, United States · $207K–$300K/yr
-
Sr Site Reliability Engineer
BeyondTrust Toronto, Ontario, Canada
-
Principal Site Reliability Engineer
UnitedHealth Group Eden Prairie, Minnesota, United States · $135K–$231K/yr
-
(EG0069) Senior Site Reliability Engineer (SRE) - Cassandra & AWS - Talent Connection
Nortal Modena, Emilia-Romagna, Italy
-
Senior Site Reliability Engineer
AbbVie Atlanta, Georgia, United States · $110K–$208K/yr
-
Site Reliability Engineer, Enterprise Technology Services
Apple Austin, Texas, United States