Manager | Site Reliability Engineering | Bengaluru | Engineering | Platform Development & Integratio

Deloitte · Bengaluru, IN · India

Manager | Site Reliability Engineering | Bengaluru | Engineering | Platform Development & Integratio • Job requisition ID : 108924  • Location: Bengaluru • Entity: Deloitte Touche Tohmatsu India LLP  The Team Deloitte’s Technology & Transformation practice can help you uncover and unlock the value buried deep inside vast amounts of data....

Manager | Site Reliability Engineering | Bengaluru | Engineering | Platform Development & Integratio
Job requisition ID : 108924 
Location: Bengaluru
Entity: Deloitte Touche Tohmatsu India LLP 

The Team


Deloitte’s Technology & Transformation practice can help you uncover and unlock the value buried deep inside vast amounts of data. Our global network provides strategic guidance and implementation services to help companies manage data from disparate sources and convert it into accurate, actionable information that can support fact-driven decision-making and generate an insight-driven advantage. Our practice addresses the continuum of opportunities in business intelligence & visualization, data management, performance management and next-generation analytics and technologies, including big data, cloud, cognitive and machine learning.   


Work Profile

Experienced DevOps Engineer with expertise in designing, scaling, and maintaining high-availability, production-grade cloud infrastructure on AWS. Skilled in managing Amazon EKS clusters, automating infrastructure using Terraform, implementing CI/CD pipelines, and ensuring platform reliability through monitoring and incident management.

Roles & Responsibilities

Design, deploy, and manage scalable cloud infrastructure on AWS. Administer and optimize Amazon EKS clusters and containerized workloads. Automate infrastructure provisioning using Terraform (IaC). Build and maintain CI/CD and GitOps pipelines for reliable software delivery. Implement observability solutions using Prometheus and Grafana to monitor system health and performance. Lead incident response, troubleshoot production issues, and conduct blameless postmortems to improve system resilience.

Required Skills

 

Success Metrics

 

Location and Way of Working: 

 

   

Encuentra más ofertas como esta

Explora más ofertas activas de esta empresa o crea una cuenta en Insider Jobs para buscar, guardar y seguir oportunidades en todo el job board.