3M Consultancy

Senior Platform / Site Reliability Engineer (SRE)

3M Consultancy A$140K–A$150K/yr

Staffing and Recruiting · 2-10 employees

17 h ago
Remote sre Senior (5-10 yrs) Full-time
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The engineer will manage and optimize AWS production environments using Terraform while ensuring high availability and performance for SaaS applications. They are responsible for maintaining CI/CD pipelines, database reliability, and implementing robust disaster recovery and monitoring solutions.

What they look for

AWS Terraform PostgreSQL Amazon RDS CI/CD Infrastructure as Code Site Reliability Engineering Observability Disaster Recovery Capacity Planning Performance Tuning Security Patching Automation SaaS Linux Networking

Requirements

Candidates must have at least 5 years of experience in platform or site reliability engineering with a strong background in AWS and Infrastructure as Code. Proficiency in PostgreSQL, database management, and automated deployment processes is essential for this role.

Full description

This is a remote position.

Job Title: Senior Platform / Site Reliability Engineer (SRE)

Location: Remote – Australia

Employment Type: Full-Time.

ABOUT THE OPPORTUNITY

We are looking for an experienced Senior Platform / Site Reliability Engineer to take ownership of a growing cloud-based SaaS environment.

This is a hands-on technical role for someone who enjoys solving infrastructure challenges, improving system reliability, and working independently. The selected candidate will be responsible for maintaining a secure, stable, and scalable AWS environment while supporting ongoing product development and business growth.

KEY RESPONSIBILITIES

- Manage AWS production environments across Australia and the United States using Terraform.

- Improve monitoring, logging, alerting, and observability to identify and resolve issues early.

- Maintain CI/CD pipelines and improve deployment processes, including automated health checks and rollbacks.

- Ensure PostgreSQL and Amazon RDS databases remain reliable, secure, and optimized for performance.

- Manage database backups, restoration procedures, query optimization, and scaling.

- Develop, document, and regularly test disaster recovery procedures.

- Monitor infrastructure capacity and prepare systems for increasing customer demand.

- Collaborate with software development teams to improve application reliability and performance.

- Manage infrastructure maintenance, server updates, security patches, and vulnerability fixes.

- Investigate production incidents, identify root causes, and implement long-term solutions.

- Improve infrastructure automation and operational processes without disrupting ongoing development.

REQUIRED SKILLS AND EXPERIENCE

- Minimum 5 years of experience in Platform Engineering, Site Reliability Engineering, DevOps, or Cloud Infrastructure.

- At least 3 years of hands-on experience managing AWS production environments for SaaS applications.

- Strong experience with Terraform and Infrastructure as Code (IaC).

- Hands-on knowledge of CI/CD pipelines, automated deployments, and rollback procedures.

- Experience with monitoring, alerting, logging, and observability tools.

- Strong PostgreSQL and Amazon RDS experience, including performance tuning, backups, restores, and scalability.

- Experience developing and testing disaster recovery plans, including RTO and RPO.

- Good understanding of AWS security, networking, access management, and infrastructure management.

- Experience with capacity planning, performance optimization, and production incident resolution.

- Ability to independently manage production infrastructure and take full technical ownership.

- Strong problem-solving and communication skills.

PREFERRED QUALIFICATIONS

- Familiarity with SOC 2, ISO 27001, or similar compliance standards.

- Experience working with managed-service providers or external support teams.

- Background supporting enterprise SaaS applications, particularly in industrial or supply chain environments.

- Previous experience as the primary or sole Platform/SRE Engineer.

- Experience managing cloud infrastructure across multiple regions.

IDEAL CANDIDATE

We are looking for a proactive, hands-on engineer who can independently manage and improve a production SaaS environment.

The ideal candidate will have strong AWS, Terraform, PostgreSQL/RDS, and CI/CD experience, along with proven skills in monitoring, disaster recovery, infrastructure security, and platform scalability.

This position is suitable for someone who takes ownership, works independently, and focuses on practical improvements that strengthen reliability and support business growth.

Similar roles