Senior Team Lead | Engineering, AI & Data - Engineering | Site Reliability Engineering

Deloitte · Bengaluru, IN · India

Senior Team Lead | Engineering, AI & Data - Engineering | Site Reliability Engineering senior · Technology / Software Development

Vista previa de la oferta

Senior Team Lead | Engineering, AI & Data - Engineering | Site Reliability Engineering

senior · Technology / Software Development

Senior Team Lead | Engineering, AI & Data - Engineering | Site Reliability Engineering
Job requisition ID : 108935 
Location: Bengaluru
Entity: Deloitte Touche Tohmatsu India LLP 

Senior Team Lead | Engineering, AI & Data - Engineering | Site Reliability Engineering

 

Location: Bangalore

 

The team

Engineering helps Reimagine and re-engineer mission-critical operations and processes; Leverage engineering-led design, deep industry knowledge, and AI and data-driven insights to transform the technology platforms at the heart of business.  

Working alongside team, we empower and drive mission-critical solutions whether we need to modernize existing systems or implement new technology products and platforms. Through innovation, we improve financial performance, accelerate new digital businesses and fuel growth. Learn more about Engineering, AI and Data

 

Your work profile

We are looking for a highly skilled Site Reliability Engineer (SRE) to manage and scale mission-critical, production-grade distributed systems running on Google Cloud Platform (GCP). The ideal candidate will focus on reliability, automation, observability, and operational excellence while minimizing toil and improving system availability. Maintaining and improving 4 Nines of uptime to 5 Nines with engineering efforts.

This role requires deep technical expertise in cloud-native technologies, Kubernetes, infrastructure automation, Linux administration and TCP/IP fundamentals, programming in one language and strong troubleshooting capabilities for distributed systems. The candidate needs to participate in the overall lifecycle management of mission critical banking services with a 24x7 operations mode in an rotational on-call basis. The job requires the candidate to have strong troubleshooting skills in a distributed environment spread across multiple cloud environments. The bare minimum ask would be to maintain high level of agility, learnability and adaptability in different scenarios. An engineer with a zeal to learn fast and having a bias for action would be the best fit for the role.

 

Reliability & Operations

 

Cloud & Infrastructure

 

Kubernetes & Containers

•   Deploy and manage containerized workloads using Kubernetes (GKE).

•   Troubleshoot issues related to: Pods, nodes, networking, storage, services on an ongoing basis

•   Manage deployments using Helm, YAML, and rollout strategies (Canary/Blue-Green).

 

Automation & CI/CD

 

Observability & Monitoring

 

System & Application Troubleshooting

 

Key skills required

 

 

Technical Skills

Cloud & Platform

 

Infrastructure as Code

•   Strong hands-on experience with Terraform

•   Ability to write and debug Terraform code from scratch Containers & Orchestration

•   Deep expertise in: Kubernetes (GKE), Docker

•   Strong troubleshooting experience in Kubernetes environments

 

CI/CD & Automation

•   Hands-on experience with: Jenkins (pipeline-based CI/CD), GitHub

•   Strong scripting skills: Python (preferred), Shell scripting

• Experience with automation frameworks and tooling

 

Observability

 

Programming & Debugging

•   Working knowledge of: Java and/or Golang applications

•   Strong debugging skills across application and infrastructure layers

 

Linux & Networking

•   Strong Linux fundamentals

•   Deep understanding of TCP/IP networking

•   Ability to debug network issues in distributed systems

 

Reliability Engineering Skills

•   Solid understanding of: SLI, SLO, SLA, Error Budgets

•   Demonstrable and Proven Experience improving: MTTR, MTTA

•   Experience handling incident management lifecycle

 

Soft Skills

•       Strong analytical and troubleshooting mindset

•       Excellent communication and stakeholder management

•       Ability to work in high-pressure production environments

•       Ownership-driven and proactive approach

 

Ideal Candidate Profile in summary would be like :

Encuentra más ofertas como esta

Explora más ofertas activas de esta empresa o crea una cuenta en Insider Jobs para buscar, guardar y seguir oportunidades en todo el job board.