Senior Platform Engineer

Standard Chartered · Chennai, IND · India

Senior Platform Engineer senior · Technology / Software Development

Vista previa de la oferta

Senior Platform Engineer

senior · Technology / Software Development

Job Description Apply now Requisition Number: 59454 Job Location: Chennai, IND Global Grade: Band 6 Work Type: Office Working Employment Type: Permanent Posting Start Date: 04/08/2026 Posting End Date: 18/08/2026 Job Description:

Job Summary

To own the end-to-end operational health, reliability, and performance of our AI/ML platform running across Azure, AWS, and On-Premises GPU environments. The ideal candidate will have deep expertise in maintaining GPU-accelerated infrastructure, AI/ML application stacks, and hybrid-cloud deployments while ensuring maximum uptime, scalability, and security.

You will play a critical role in supporting production AI workloads, troubleshooting complex platform issues, and continuously optimizing the infrastructure that powers our AI applications (LLMs, ML models, inference services, RAG pipelines, etc.).

Key Responsibilities

Platform Operations & Maintenance
• Manage day-to-day operations, monitoring, and maintenance of the AI Platform across Azure, AWS, and on-premises GPU clusters.
• Ensure high availability, reliability, and performance of AI/ML applications and services (training, inference, fine-tuning, RAG pipelines).
• Perform proactive health checks, capacity planning, patching, and lifecycle management of infrastructure components.
• Handle incident management, root cause analysis (RCA), and problem resolution within defined SLAs.
GPU Infrastructure Management
• Maintain and troubleshoot GPU nodes (NVIDIA A100, H100, V100, L40S, etc.), including driver installations, CUDA  updates.
• Optimize GPU utilization, workload scheduling, and resource allocation across multi-tenant environments.
• Manage GPU clusters using Kubernetes (with GPU operators) or similar orchestrators.
Cloud & Hybrid Environment Management
• Administer AI/ML workloads on Azure ML, AWS SageMaker, EKS/AKS, EC2/Azure VMs, and on-prem Kubernetes/OpenShift clusters.
• Manage container-based deployments using Docker, Kubernetes, Helm, and GPU-enabled runtimes.
• Configure and maintain networking, storage (NFS, Blob, S3), and IAM across hybrid environments.

AI Application Support
• Support production AI applications including LLM serving, MLflow, Kubeflow, Ray, and similar platforms.
• Deploy, upgrade, and maintain model registries, vector databases and orchestration tools.
• Collaborate with Data Scientists and ML Engineers to onboard models, optimize inference, and troubleshoot performance bottlenecks.
Automation & CI/CD
• Automate platform provisioning, configuration, and maintenance using Terraform, Ansible, ARM/Bicep, CloudFormation.
• Build and maintain CI/CD pipelines for AI/ML workloads (GitHub Actions, Azure DevOps, Jenkins, GitLab).
• Implement Infrastructure as Code (IaC) practices for reproducible environments.
Monitoring, Logging & Observability
• Implement and manage observability stacks: Prometheus, Grafana, ELK,  Azure Monitor, CloudWatch.
• Set up GPU-specific monitoring and alerting for AI workloads.
• Track model performance, drift, latency, and throughput.
Security & Compliance
• Enforce security best practices (identity, network, secrets management, encryption).
• Ensure compliance with organizational and regulatory standards 
• Manage RBAC, key vaults, and secure model artifact storage.

Strategy
• To Maintain a highly available, secure, scalable, and cost-optimized AI Platform across Azure, AWS, and On-Premises GPU environments that enables seamless development, deployment, and operation of AI/ML and GenAI workloads with enterprise-grade reliability.

Business

• AI Factory architecture and engineering leadership
• T&A and TTO management
• Chief Data Office and central data governance
• Software Engineering Platform Teams
• Business technology teams across Trade, Cash, Financial Markets, Risk & Compliance, and Retail Operational, Technology & Cyber Risk; Group Internal Audit Compliance, Regulators, and the independent AI governance function 

Regulatory & Business Conduct

• Display exemplary conduct and live by the Group’s Values and Code of Conduct. 
• Take personal responsibility for embedding the highest standards of ethics, including regulatory and business conduct, across Standard Chartered Bank. This includes understanding and ensuring compliance with, in letter and spirit, all applicable laws, regulations, guidelines and the Group Code of Conduct.
• Effectively and collaboratively identify, escalate, mitigate and resolve risk, conduct and compliance matters.
Key stakeholders

• AI Factory architecture and engineering leadership
• T&A and TTO management
• Chief Data Office and central data governance
• Software Engineering Platform Teams
• Business technology teams across Trade, Cash, Financial Markets, Risk & Compliance, and Retail Operational, Technology & Cyber Risk; Group Internal Audit Compliance, Regulators, and the independent AI governance function 

Skills and Experience

• Azure Administrator
• Databricks
• Linux Administrator
• Kubernetes
• GPU
• Python 

Qualifications

• Bachelor's/Master's degree in Computer Science, Engineering, or related field.
• 8+ years of experience in Platform/Infrastructure/DevOps/SRE roles, with at least 3+ years in AI/ML platform engineering.
• Strong hands-on experience with Azure and AWS (compute, networking, storage, IAM, GPU instances).
• Solid experience managing on-premises GPU infrastructure (NVIDIA DGX, HPE, Dell, Supermicro GPU servers).
• Expertise in Kubernetes, GPU Operator, container runtimes, and Helm.
• Proficiency in Linux system administration, shell scripting, and Python.
• Experience with IaC tools (Terraform, Ansible) and CI/CD pipelines.
• Familiarity with AI/ML frameworks: PyTorch, TensorFlow, Hugging Face, and inference servers (Triton, vLLM, TGI).
• Strong troubleshooting skills across networking, storage, GPU, and application layers.
• Experience with monitoring/observability tools (Prometheus, Grafana, DCGM).
• Certifications: Azure Solutions Architect, AWS Solutions Architect/DevOps, CKA/CKAD, NVIDIA DLI.

About Standard Chartered

We're an international bank, nimble enough to act, big enough for impact. For more than 170 years, we've worked to make a positive difference for our clients, communities, and each other. We question the status quo, love a challenge and enjoy finding new opportunities to grow and do better than before. If you're looking for a career with purpose and you want to work for a bank making a difference, we want to hear from you. You can count on us to celebrate your unique talents and we can't wait to see the talents you can bring us.

Our purpose, to drive commerce and prosperity through our unique diversity, together with our brand promise, to be here for good are achieved by how we each live our valued behaviours. When you work with us, you'll see how we value difference and advocate inclusion.

Together we:

What we offer

In line with our Fair Pay Charter, we offer a competitive salary and benefits to support your mental, physical, financial and social wellbeing.

Apply now Information at a Glance

Similar Jobs

Senior SRE EngineerGuangzhou, CHN TechnologyRequisition Number58869ChinaAVP, Chapter lead, Engineering - LARCBangalore, IND TechnologyRequisition Number58509IndiaAVP, Engineering Lead, WRB TechBangalore, IND TechnologyRequisition Number58903IndiaSenior SRE EngineerGuangzhou, CHN TechnologyRequisition Number58870China Job SkillsKey SkillsDevOpsLinuxInfrastructure as Code (IaC)ScalabilityOpenShiftKubernetesTerraformSystem AvailabilityMicrosoft AzureDockerAnsibleAmazon Web ServicesCloud ComputingLinux System AdministrationRelevant SkillsUptimeData LoggingApplication LayersCloudwatchInformation TechnologySoftware EngineeringGitlabAzure Machine LearningShell ScriptNetwork PerformanceSchedulingComputer ClustersIncident ManagementSelf MotivationCapital MarketsKey ManagementJenkinsMaintenanceResource AllocationRetail CommerceInnovationContinuous IntegrationGeographyContinuous TrainingPytorchBusiness TechnologiesArtificial IntelligenceBicepIdentity and Access ManagementGithubRisk AnalysisProduct Family EngineeringHybridCloudLeadershipPython (Programming Language)Installation TechnologyAutomationSafety PrinciplesNvidia CUDAGovernanceCapacity PlanningRole-Based Access ControlGrafanaIT Risk ManagementCloudformationMachine Learning OperationsAssertivenessLifecycle ManagementData GovernanceVector DatabasesRoot Cause AnalysisAI PlatformsArchitectureAmazon S3ElearningSecurity ManagingDatabricksAudit ManagementMental HealthMachine LearningLarge Language ModelsKnowledge of FinanceAmazon Elastic Compute CloudInfrastructure ManagementHealth AssessmentCryptographyHardware InfrastructureRegulatory ComplianceReliabilityInfrastructure Automation FrameworksPerseverancePrometheusTensorflowHuggingFaceProblem SolvingMore

Skills Matching

Upload Your CVUse your CV to see how well your skills match up with this jobGet Started

Encuentra más ofertas como esta

Explora más ofertas activas de esta empresa o crea una cuenta en Insider Jobs para buscar, guardar y seguir oportunidades en todo el job board.