Wesco

Senior Site Reliability Engineer

Wesco Irving, Texas, United States · $82K–$139K/yr

Wholesale · 10,001+ employees

14 h ago
sre Senior (5-10 yrs) Full-time United States
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The Senior Site Reliability Engineer is responsible for deploying, validating, and operationalizing AI, HPC, and Kubernetes infrastructure environments. This role involves transforming hardware into production-ready platforms through automation, testing, and standardized provisioning.

What they look for

Kubernetes Infrastructure as Code Python PowerShell Bash Linux Administration VMware ESXi GPU Technologies CUDA High Performance Computing Networking Storage Integration Automation Troubleshooting Root-cause analysis Security hardening

Requirements

Candidates must have at least 5 years of experience in infrastructure or platform engineering with strong knowledge of Kubernetes and Linux. A degree in a technical discipline is preferred, along with proficiency in automation scripting and hardware configuration.

Benefits

Paid time off Medical coverage Dental coverage Vision coverage Retirement savings plans

Full description

As the Senior Site Reliability Engineer, you will serve as a trusted technical resource responsible for deploying, validating, and operationalizing AI, HPC, Kubernetes, and enterprise infrastructure environments. This role transforms newly installed hardware into production-ready platforms through standardized provisioning, automation, testing, and infrastructure validation activities. Working as part of a holistic team strategy, you will support large, complex customer deployments and ensure infrastructure environments are ready for operational handoff and long-term success.

Responsibilities:

  • Provide technical expertise and engagement to support infrastructure readiness, platform engineering, and deployment activities across customer environments.
  • Deploy, configure, and validate AI, GPU, and High Performance Computing (HPC) infrastructure solutions.
  • Prepare and administer Kubernetes platforms, container runtimes, storage integrations, networking components, and cluster infrastructure.
  • Install, configure, and validate NVIDIA technologies including GPU drivers, CUDA, GPU Operators, AI Enterprise prerequisites, and telemetry solutions.
  • Validate accelerated networking technologies including InfiniBand, RoCE, RDMA, and GPU-to-GPU communications.
  • Perform infrastructure readiness assessments, burn-in testing, operational acceptance testing, and performance validation activities.
  • Configure and support server infrastructure including iDRAC, iLO, BMC, firmware, storage, and networking components.
  • Deploy and administer Windows, Linux, VMware ESXi, Hyper-V, and KVM-based environments.
  • Apply security hardening standards, compliance requirements, and operational best practices throughout deployment and validation activities.
  • Develop and maintain automation workflows utilizing PowerShell, Python, Bash, and Infrastructure-as-Code methodologies.
  • Create customer-facing deployment documentation, technical reports, readiness assessments, and operational validation deliverables.
  • Troubleshoot complex hardware, operating system, virtualization, containerization, networking, and AI platform issues.
  • Participate in advanced technical training and continued education to maintain expertise in cloud, infrastructure, AI, and platform technologies.
  • Support technical engagements across customer environments and collaborate with internal engineering, architecture, and service delivery teams.

Qualifications:

  • Associate degree (U.S.)/College Diploma (Canada) or equivalent combination of education and technical experience required.
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related technical discipline preferred.
  • 5+ years of experience in Infrastructure Engineering, Platform Engineering, Site Reliability Engineering (SRE), Systems Administration, or related technical roles.
  • Experience deploying, supporting, or validating AI, GPU, HPC, or large-scale enterprise infrastructure environments.
  • Experience with Kubernetes, container platforms, and enterprise Linux administration.
  • Strong knowledge of server provisioning, virtualization, storage, networking, and infrastructure operations.
  • Experience with VMware ESXi, Hyper-V, KVM, or related virtualization technologies.
  • Experience developing automation and scripting solutions using PowerShell, Python, Bash, or similar tools.
  • Knowledge of Infrastructure-as-Code and automated deployment methodologies.
  • Experience with NVIDIA GPU technologies, CUDA, AI Enterprise, or related AI infrastructure platforms preferred.
  • Knowledge of InfiniBand, RDMA, RoCE, or high-performance networking technologies preferred.
  • Demonstrated troubleshooting, root-cause analysis, and problem-solving skills.
  • Possess a customer-centric mindset and strong written and verbal communication skills.
  • Possess intermediate computer skills, including proficiency with Microsoft Office applications.
  • Ability to travel up to 25%.

Preferred Certifications

  • Certified Kubernetes Administrator (CKA)
  • Red Hat Certified System Administrator (RHCSA) or equivalent Linux certification
  • NVIDIA certifications related to AI, GPU, or DGX platforms
  • VMware Certified Professional (VCP) or equivalent

#LI-VR1 #Hybrid

This amount is what we reasonably believe we will pay for the position; however, offer amounts may vary based on factors such as geographic location, relevant education, experience, qualifications, skills, shift, or any collective bargaining agreements.

For eligible positions, compensation may include participation in a bonus or sales incentive plan, subject to the terms and conditions of the applicable plan documents. For certain sales roles, Wesco also offers a commission structure that provides additional compensation based on sales results, as defined by the applicable commission plan.

In addition, Wesco offers a benefits program for eligible employees, which may include paid time off, medical, dental, and vision coverage, and retirement savings plans. Additional details about benefits are available here.

At Wesco, we build, connect, power and protect the world. As a leading provider of business-to-business distribution, logistics services and supply chain solutions, we create a world that you can depend on. ​

Our Company’s greatest asset is our people. Wesco is committed to fostering a workplace where every individual is respected, valued, and empowered to succeed. We promote a culture that is grounded in teamwork and respect. With a workforce of over 20,000 people worldwide, we embrace the unique perspectives each person brings. Through comprehensive benefits and active community engagement, we create an environment where every team member has the opportunity to thrive. ​

Learn more about Working at Wesco here and apply online today!​

Founded in 1922 and headquartered in Pittsburgh, Wesco is a publicly traded (NYSE: WCC) FORTUNE 500® company.​

Wesco International, Inc., including its subsidiaries and affiliates (“Wesco”) provides equal employment opportunities to all employees and applicants for employment. Employment decisions are made without regard to race, religion, color, national or ethnic origin, sex, sexual orientation, gender identity or expression, age, disability, or other characteristics protected by law. US applicants only, we are an Equal Opportunity Employer.​

Los Angeles Unincorporated County Candidates Only: Qualified applicants with arrest or conviction records will be considered for employment in accordance with the Los Angeles County Fair Chance Ordinance and the California Fair Chance Act.

This posting is for a current, active vacancy intended for immediate hire.

Similar roles