Aviato Consulting is seeking an experienced Site Reliability Engineer to join our growing team. This isn't just another SRE role; it's an opportunity to own critical infrastructure, drive technical strategy, and shape the reliability culture for major Australian and EU clients, all within a supportive, G-inspired environment built on transparency and collaboration.
What's In It For You?
Learn from the Best: Report directly to and receive mentorship from our Head of SRE, an experienced ex-Google Manager.
High-Impact Projects: Take ownership of complex GCP environments for diverse, significant clients across Australia and the EU.
Drive Innovation, Not Just Tickets: Architect solutions and implement cutting-edge practices like SLOs, error budgets, and predictive AI anomaly detection to proactively improve systems.
A Culture That Works: Founded by ex-Googlers, we foster a transparent, collaborative, and low-bureaucracy environment.
What You'll Do (Your Impact):
Own & Architect Reliability: Design, implement, and manage highly available, scalable architectures on Google Cloud Platform (GCP).
Master Kubernetes & AI Infrastructure: Architect and manage production-grade Kubernetes clusters (GKE), and support high-performance infrastructure for AI/ML workloads (e.g., GPU/TPU pools).
Drive Automation & IaC: Lead robust automation strategies using Terraform, Ansible, and scripting (Python, Go, Bash) for CI/CD pipelines.
Elevate Observability with AIOps: Architect monitoring, logging, and alerting using Grafana, Dynatrace, and Sentry, integrating AI-driven insights for automated root-cause analysis.
Lead Incident Response: Spearhead incident management, conduct blameless post-mortems, and leverage GenAI to automate runbook generation.
Champion SRE Principles: Actively promote SLOs, SLIs, and error budgets, and mentor team members.
What You'll Bring (Your Expertise):
Proven SRE Experience: 5+ years of hands-on experience in a Site Reliability Engineering or Cloud Engineering role focusing on production systems.
Deep GCP & Kubernetes Knowledge: Demonstrable expertise in core GCP services and managing Kubernetes clusters in production (GKE highly desirable).
Infrastructure as Code Mastery: Significant experience using Terraform in complex environments.
Automation & Scripting Prowess: Strong proficiency in Python or Go for automating operational tasks.
Next-Gen Observability Expertise: Experience with modern monitoring tools, preferably with exposure to AI-driven alerting and logs clustering.
Problem-Solving Acumen: Strong analytical skills with experience leading incident response for critical systems.
(Desirable): Experience with AI infrastructure (vector databases, LLM inference), or API Management platforms like Apigee.
Technologies We Use (You'll Master):
Cloud: Google Cloud Platform (GCP)
Containerisation & Orchestration: Kubernetes (GKE), Docker
Infrastructure & Automation: Terraform, Ansible
Monitoring & AIOps: Grafana, Dynatrace, Sentry, Google Cloud Operations Suite
CI/CD: Jenkins, GitHub Actions, Bamboo (or similar)
Scripting & AI Tooling: Python, Go, Bash, GitHub Copilot, Gemini/Vertex AI
Collaboration: JIRA, Confluence, Slack
Work timings: 5 am to 2pm IST
Ready to Elevate Your SRE Career?
If you're a passionate Senior SRE ready to tackle complex challenges on GCP, work with leading clients, and benefit from exceptional mentorship in a fantastic culture, Aviato is the place for you. Apply now and help us build the future of reliable cloud infrastructure!