Senior Site Reliability Engineer, NetBox Delivery
NetBox Labs London, England, United Kingdom · $180K–$195K/yr
Computer Networking Products · 11-50 employees
About the role
You will own the NetBox build and release pipeline to ensure reliable delivery of software to Cloud and Enterprise customers. Additionally, you will act as an escalation point for performance issues and implement observability and security practices across the infrastructure.
What they look for
Requirements
Candidates must have 5+ years of experience in software or platform engineering with production-level expertise in Django and Postgres. Proficiency in containerization, cloud infrastructure (AWS), and modern CI/CD tooling is essential for this role.
Benefits
Full description
NetBox Labs is seeking a Senior Site Reliability Engineer for NetBox Delivery, a new team in our Applications group.
NetBox is a product that reaches people in a few ways: our open source community runs NetBox OSS; commercial customers use NetBox Cloud (SaaS) or NetBox Enterprise (self-managed). NetBox Delivery owns everything between a NetBox Core release and a healthy, running instance on Cloud and Enterprise. We ship the software, keep an eye on it in production, and when something breaks, we fix it at the source rather than working around it. You will be one of the first engineers on this team and help shape how it works.
In this role you will:
- Own the NetBox build and release pipeline, from base images to downstream availability on Cloud and Enterprise
- Build the release handoff between NetBox Core and the Cloud and Enterprise teams, so new releases reach customers quickly and predictably
- Make NetBox faster and more reliable in production, from application startup to Postgres performance
- Build real observability for the application and the release pipeline, including monitoring, alerting, and SLOs
- Act as the escalation point for performance and reliability issues and take fixes back to NetBox Core when the cause is in the code
- Strengthen supply chain security and support SOC 2 compliance for the build pipeline
- Share on-call duties and lead incident response and postmortems for your area
Requirements:
- 5+ years in software engineering, platform engineering, or SRE, with proven experience writing robust, maintainable code
- Production experience with Django and Postgres at scale, including schema design, migration risk, and query performance under real load
- Strong container build skills, including base image design, Python dependency management, and supply chain security practices like vulnerability scanning and image signing
- Hands-on experience with our stack or something close to it: AWS (EC2, VPC, IAM, RDS), Kubernetes and Helm, GitHub Actions, ArgoCD or FluxCD, Terraform, and Prometheus and Grafana
- Hands-on experience building inside an AI-augmented development harness, including Claude Code and the workflows that make agentic tooling reliable
- A track record of driving work across team boundaries, from writing the RFC to getting a cross-team migration done
Nice to haves:
- Familiarity with the NetBox ecosystem or network automation
- Open source contributions or project involvement
- Experience working in a B2B software startup or high-growth organization
- Deep experience with supply chain security tooling such as cosign, Sigstore, or SLSA
- Experience operating high-throughput or performance-sensitive systems for large enterprise customers
About NetBox Labs:
NetBox Labs helps companies build and manage complex networks. We help customers accelerate network automation by delivering open, composable products and supporting the network automation community.
NetBox Labs is the commercial steward of open source NetBox, the world’s most popular network source of truth, and Orb, the next-generation open source network observability platform. Our products include NetBox Enterprise, a fully supported self-managed NetBox with advanced features, and NetBox Cloud, a secure, scalable, and reliable SaaS edition of NetBox.
NetBox powers thousands of companies, and NetBox Labs is backed by investment from Notable Capital (formerly GGV), Grafana Labs CEO Raj Dutt, Flybridge, IBM, Salesforce Ventures, and Mango Capital.
Our culture and values:
- We own and solve problems with high attention to detail.
- Our open source contributors, users, customers & team are all part of our community. When our community wins, we win.
- We prioritize simplicity and think twice before adding complexity
- Clear communication helps keep our team aligned and collaborating smoothly.
NetBox Labs is proud to be an equal opportunity employer. We believe diverse teams build better software, and we welcome applicants of every race, color, religion, gender identity, sexual orientation, national origin, age, disability, and veteran status. If you need accommodation at any point in the process, just let us know.
Similar roles
-
Senior Site Reliability Engineer (SRE) – Application Observability & Readiness (Azure)
Encora Perímetro Urbano Santiago de Cali, Valle del Cauca, Colombia
-
Site Reliability Engineer III - Performance Engineer- Service now
JPMorgan Chase & Co. San Francisco, California, United States · $138K–$185K/yr
-
Site Reliability Engineering Lead
Graphcore Austin, Texas, United States
-
Site Reliability Engineer
Prolaio Chicago, Illinois, United States · $118K/yr
-
Senior Site Reliability Engineer
Wesco Irving, Texas, United States · $82K–$139K/yr
-
ASKUSR0147035 Site Reliability Engineer - NERSC Operations
Essnova Solutions, Inc. Berkeley, California, United States · $156K/yr