Vista previa generada por IA
Vice President of Production Services Application Support – Chennai, India
Lead mission-critical production support, incident management, and automation across complex environments.
Vice President, Production Services Application Support
Location: Chennai, Tamil Nadu, India
In this role, you’ll make an impact in the following ways:
- In this role, you will play a critical part in protecting production stability, accelerating service recovery, and driving continuous improvement across mission-critical platforms.
- Own end-to-end support for mission-critical applications, ensuring high availability, stability, resiliency, and timely incident resolution.
- Lead complex technical triage across application, database, middleware, messaging, batch, and infrastructure layers to restore service quickly and effectively.
- Support payments applications, including transaction flow validation, issue analysis, health checks, and operational readiness.
- Drive incident and problem management, root cause analysis, service restoration, and permanent fix follow-up for high-impact production issues.
- Partner across application development, DBA, infrastructure, network, middleware, operations, and business teams to resolve issues, reduce risk, and improve reliability.
- Monitor application health and identify risks using Splunk, Grafana/Optics, AppDynamics, Moogsoft, and other observability tools.
- Support batch scheduling, recovery, and operational validations using Control-M and approved support procedures.
- Execute and validate deployment, release, and change activities through GitLab / CI/CD pipelines aligned to change standards.
- Champion automation using Ansible and UNIX/Linux scripting to reduce manual effort, improve consistency, and strengthen operational controls.
- Support MQ and Kafka platforms to ensure stable integration and operational continuity.
- Maintain support documentation, runbooks, knowledge articles, escalation guides, and recovery procedures.
- Participate in production readiness reviews, resiliency testing, disaster recovery exercises, and failover validations.
- Use ServiceNow for incident, problem, change, service request, and follow-up tracking.
- Collaborate with global teams and deliver clear, confident updates during incidents, bridge calls, and leadership communications.
- Leverage AI tools such as Microsoft Copilot to elevate documentation, analysis, reporting, knowledge management, and team productivity.
To be successful in this role, we’re seeking the following:
- Bachelor’s or higher degree in computer science, engineering, or a related discipline, or equivalent work experience.
- 10+ years of proven experience in production support, application support, or technology operations.
- Demonstrated success supporting enterprise-scale, mission-critical applications in complex production environments.
- Knowledge of payments domain flows, transaction processing, and production issue analysis.
- Hands‑on Oracle SQL experience for querying, troubleshooting, data validation, and issue investigation.
- Strong UNIX/Linux skills, including log analysis, file system checks, process monitoring, and command-line troubleshooting.
- Hands‑on experience with Ansible for automation and operational task execution.
- Knowledge of AppEngine and application runtime support.
- Experience with GitLab and CI/CD pipelines for deployment, release, and validation activities.
- Experience with Control-M or similar enterprise batch scheduling tools.
- Experience with Splunk, Grafana/Optics, AppDynamics, Moogsoft, and related observability platforms.
- Experience using ServiceNow for ITSM, including incident, problem, change, and service request management.
- Proficiency in messaging technologies such as MQ and Kafka.
- Exposure to AI-enabled productivity tools such as Microsoft Copilot.
- Strong understanding of incident, problem, and change management, production readiness, and operational risk controls.
- Ability to analyze logs, alerts, batch failures, messaging issues, transaction breaks, infrastructure events, and application errors to identify root cause and remediation.
- Excellent communication and stakeholder management skills, with the ability to translate technical issues into clear updates for technology, operations, business, and leadership stakeholders.
- Ability to stay composed under pressure, take ownership during critical incidents, and drive issues to closure with urgency and accountability.