Full-time Site Reliability Engineer (SRE)

Remote, IndiaCloud


Aviato Consulting is seeking an experienced Site Reliability Engineer to join our growing team. This isn't just another SRE role; it's an opportunity to own critical infrastructure, drive technical strategy, and shape the reliability culture for major Australian and EU clients, all within a supportive, G-inspired environment built on transparency and collaboration. 


What's In It For You?

  • Learn from the Best: Report directly to and receive mentorship from our Head of SRE, an experienced ex-Google Manager.

  • High-Impact Projects: Take ownership of complex GCP environments for diverse, significant clients across Australia and the EU.

  • Drive Innovation, Not Just Tickets: Architect solutions and implement cutting-edge practices like SLOs, error budgets, and predictive AI anomaly detection to proactively improve systems.

  • A Culture That Works: Founded by ex-Googlers, we foster a transparent, collaborative, and low-bureaucracy environment.

What You'll Do (Your Impact):

  • Own & Architect Reliability: Design, implement, and manage highly available, scalable architectures on Google Cloud Platform (GCP).

  • Master Kubernetes & AI Infrastructure: Architect and manage production-grade Kubernetes clusters (GKE), and support high-performance infrastructure for AI/ML workloads (e.g., GPU/TPU pools).

  • Drive Automation & IaC: Lead robust automation strategies using Terraform, Ansible, and scripting (Python, Go, Bash) for CI/CD pipelines.

  • Elevate Observability with AIOps: Architect monitoring, logging, and alerting using Grafana, Dynatrace, and Sentry, integrating AI-driven insights for automated root-cause analysis.

  • Lead Incident Response: Spearhead incident management, conduct blameless post-mortems, and leverage GenAI to automate runbook generation.

  • Champion SRE Principles: Actively promote SLOs, SLIs, and error budgets, and mentor team members.

What You'll Bring (Your Expertise):

  • Proven SRE Experience: 5+ years of hands-on experience in a Site Reliability Engineering or Cloud Engineering role focusing on production systems.

  • Deep GCP & Kubernetes Knowledge: Demonstrable expertise in core GCP services and managing Kubernetes clusters in production (GKE highly desirable).

  • Infrastructure as Code Mastery: Significant experience using Terraform in complex environments.

  • Automation & Scripting Prowess: Strong proficiency in Python or Go for automating operational tasks.

  • Next-Gen Observability Expertise: Experience with modern monitoring tools, preferably with exposure to AI-driven alerting and logs clustering.

  • Problem-Solving Acumen: Strong analytical skills with experience leading incident response for critical systems.

  • (Desirable): Experience with AI infrastructure (vector databases, LLM inference), or API Management platforms like Apigee.

Technologies We Use (You'll Master):

  • Cloud: Google Cloud Platform (GCP)

  • Containerisation & Orchestration: Kubernetes (GKE), Docker

  • Infrastructure & Automation: Terraform, Ansible

  • Monitoring & AIOps: Grafana, Dynatrace, Sentry, Google Cloud Operations Suite

  • CI/CD: Jenkins, GitHub Actions, Bamboo (or similar)

  • Scripting & AI Tooling: Python, Go, Bash, GitHub Copilot, Gemini/Vertex AI

  • Collaboration: JIRA, Confluence, Slack

Work timings: 5 am to 2pm IST

Ready to Elevate Your SRE Career?

If you're a passionate Senior SRE ready to tackle complex challenges on GCP, work with leading clients, and benefit from exceptional mentorship in a fantastic culture, Aviato is the place for you. Apply now and help us build the future of reliable cloud infrastructure!

Apply