Vista previa de la oferta
Senior Consultant | ITSM | Bengaluru | Engineering | Platform Development & Integration
senior · Consulting
Senior Consultant | ITSM | Bengaluru | Engineering | Platform Development & Integration
• Job requisition ID : 109992
• Location: Bengaluru
• Entity: Deloitte Touche Tohmatsu India LLP
Senior Consultant | Engineering | Platform Development & Integration | ITSM
Location: Bangalore
The team
Our Enterprise Technology & Performance team helps organizations build, operate, and optimize resilient cloud-native platforms. We are looking for an experienced Lead Production Incident Manager (IM) to lead enterprise production operations, incident management, and cloud infrastructure reliability across large-scale AWS environments. The ideal candidate should possess strong expertise in production support, Site Reliability Engineering (SRE), cloud technologies, and ITIL-based service management, preferably with experience in Banking & Financial Services, specifically Cards & Payments and Mobile Applications. Enterprise technology has to do much more than keep the wheels turning; it is the engine that drives functional excellence and the enabler of innovation and long-term growth. Learn more about: Customer
Role Summary
We are seeking an experienced Lead Production Incident Manager (IM) to lead enterprise production operations, major incident management, and cloud infrastructure reliability across large-scale AWS environments. The ideal candidate will be responsible for ensuring secure, scalable, highly available, and cost-effective cloud operations while driving operational excellence, service reliability, and continuous improvement across mission-critical enterprise applications. This role requires strong expertise in ITIL-based service management, Site Reliability Engineering (SRE), AWS cloud technologies, production support, automation, and stakeholder management. Experience in the Banking & Financial Services domain, particularly Cards & Payments, Mobile Applications, and Cloud-Native Solutions, will be highly advantageous.
-
Lead enterprise-wide Incident, Problem, and Change Management activities aligned with ITIL best practices.
-
Own the end-to-end lifecycle of production incidents (P1–P4), ensuring timely identification, escalation, communication, resolution, and closure.
-
Drive service recovery, Root Cause Analysis (RCA), Post Incident Reviews (PIR), and Corrective & Preventive Actions (CAPA) to improve service reliability and reduce MTTR.
-
Lead 24x7 production support operations, managing L2/L3 application and infrastructure support across Linux-based environments.
-
Oversee production deployments, release management, infrastructure operations, and Site Reliability Engineering (SRE) initiatives.
-
Manage Disaster Recovery (DR), High Availability (HA) architecture, and Active-Active/Active-Passive failover strategies.
-
Drive automation, cloud infrastructure optimization, capacity planning, and operational excellence using AWS, Kubernetes, Docker, and CI/CD pipelines.
-
Monitor and optimize production environments using observability tools including Grafana, Kibana, ELK, Splunk, CloudWatch, Prometheus, Nagios, Zenduty, and Site24x7.
-
Support REST API-based applications, perform production troubleshooting using SQL, and collaborate with engineering teams to ensure highly resilient and secure production environments.
-
Lead and mentor SRE and Production Support teams while ensuring SLA, SLO, KPI, and customer satisfaction targets are consistently achieved.
Key Skills Required
- Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related discipline.
- 10–15+ years of overall IT experience with at least 8+ years specializing in Enterprise Production Support, Incident Management, Cloud Infrastructure, and Site Reliability Engineering (SRE).
- Strong experience in Banking & Financial Services, preferably supporting Cards & Payments, Mobile Applications, and Cloud-Native solutions.
- Proven expertise in Incident, Problem & Change Management (ITIL), Major Incident Management, Service Recovery, Root Cause Analysis (RCA), and Production Operations.
- Experience leading 24x7 production support teams and managing L2/L3 application and infrastructure support across Linux environments.
- Strong knowledge of AWS services including EC2, S3, RDS, Lambda, VPC, IAM, DynamoDB, CloudWatch, and secure cloud architecture.
- Experience designing highly available, scalable, secure, and disaster recovery-enabled cloud solutions.
- Hands-on experience with Docker, Kubernetes, Jenkins, CI/CD pipelines, Infrastructure as Code (Terraform, AWS CloudFormation, Ansible), and automation using Python or Bash.
- Strong understanding of Linux administration, networking, load balancing, SQL/NoSQL databases, REST APIs, and cloud security best practices.
- Experience with monitoring and observability tools including Grafana, Kibana, ELK Stack, Splunk, CloudWatch, Prometheus, Nagios, Zenduty, and Site24x7.
- Proven ability to drive SLA/SLO/KPI compliance, reduce MTTR, improve service reliability, and implement automation and continuous improvement initiatives.
- Strong leadership, stakeholder management, client communication, analytical, and problem-solving skills with the ability to coordinate cross-functional teams during critical incidents.
- ITIL Foundation Certification is preferred.
- AWS Certified Solutions Architect – Associate or Professional certification is highly preferred.