Volver a ofertas

Senior SRE Engineer

Standard Chartered · Guangzhou, CHN · China

Job Description Apply now Requisition Number: 58869 Job Location: Guangzhou, CHN Global Grade: Band 6 Work Type: Hybrid Working Employment Type: Permanent Posting Start Date: 27/07/2026 Posting End Date: 01/11/2026 Job Description: Job Summary As a Site Reliability Engineer (SRE) within Markets Technology, you will be responsible for...

Job Description Apply now Requisition Number: 58869 Job Location: Guangzhou, CHN Global Grade: Band 6 Work Type: Hybrid Working Employment Type: Permanent Posting Start Date: 27/07/2026 Posting End Date: 01/11/2026 Job Description:

Job Summary

As a Site Reliability Engineer (SRE) within Markets Technology, you will be responsible for improving the reliability, stability, performance, scalability, and operational resilience of the bank’s Markets applications and platforms. Working closely with development, infrastructure, cloud, and production support teams, you will drive the adoption of SRE practices to reduce operational risk, improve service availability, and ensure critical applications and platforms always operate reliably.

The role combines traditional SRE responsibilities with the opportunity to help shape the next generation of operational tooling through AI-enabled automation. In addition to supporting the reliability of Markets systems, you will contribute to the development of AI-powered operational assistants and Microsoft Copilot-integrated agents that improve incident management, troubleshooting, observability, knowledge management, and operational efficiency.

Key Responsibilities

Site Reliability Engineering
• Improve the reliability, availability, scalability, and performance of business-critical Markets applications and platforms.
• Establish and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets to drive reliability-focused engineering decisions.
• Lead the diagnosis and resolution of production incidents, reducing Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR).
• Drive root cause analysis and permanent remediation of recurring production issues.
• Proactively identify reliability risks and implement measures to improve resilience and fault tolerance.
• Design and implement monitoring, alerting, observability, and capacity management solutions across Markets technology platforms.
• Partner with development teams to improve application operability, observability, recoverability, and production readiness.
• Support disaster recovery, resiliency testing, failover exercises, and service continuity planning.
• Drive operational excellence through automation, self-healing capabilities, and reduction of manual operational effort.
• Provide technical leadership during high-severity incidents and critical production events.
Observability & Automation
• Develop and enhance observability solutions covering metrics, logs, traces, synthetic monitoring, and service health monitoring.
• Build automated operational workflows, runbooks, and remediation capabilities.
• Develop dashboards and analytics that provide actionable insights into application health, performance, capacity, and reliability trends.
• Support the ongoing evolution of the bank’s monitoring and observability capabilities.

AI & Copilot Engineering
• Contribute to the design and development of AI-powered operational tooling and intelligent automation solutions.
• Develop Microsoft Copilot-integrated agents that assist engineers in troubleshooting, incident response, change risk assessment, knowledge retrieval, and operational decision-making.
• Identify opportunities to apply Generative AI and AIOps capabilities to reduce operational workload and improve engineering productivity.
• Collaborate with platform engineering, architecture, and business stakeholders to develop practical AI use cases within production operations.
• Evaluate and implement AI-driven observability, anomaly detection, event correlation, and predictive operational capabilities.
Strategy
• Awareness and understanding of the T&O 30 business strategy and model appropriate to the role. Support and the enablement of the Central Monitoring & Observability strategy, goals and objectives by developing prioritized features aligned to the Catalyst and Tech Simplification programmes.
Business
• The Markets Site Reliability Engineering (SRE) team is a global function focused on ensuring the reliability, performance, resilience, and operational stability of the bank’s critical Markets applications and platforms. Working closely with development, infrastructure, and cloud engineering teams, the SRE team drives observability, automation, incident reduction, and operational excellence across the Markets technology estate.
• The ideal candidate will have strong experience supporting large-scale production environments and deep expertise in observability technologies such as Elastic, Grafana, Open Telemetry, or Geneos, together with supporting technologies including Kafka, cloud platforms, and automation frameworks. They will apply SRE principles to improve service reliability, accelerate incident resolution, reduce operational toil, and enhance platform resilience.
• In addition, the candidate will contribute to the development of AI-powered operational tooling and Microsoft Copilot-integrated agents, helping automate troubleshooting, knowledge retrieval, incident analysis, and operational workflows. This role requires a blend of reliability engineering, software development, and automation skills, with a focus on improving both platform stability and engineering productivity across the Markets domain.

Processes
• You will play a crucial role in ensuring the stability, reliability, and performance of our Markets applications and platforms, thereby enabling our organization to deliver exceptional services to our internal stakeholders by adhering to the Enterprise SDLC (eSDLC) framework and guidelines.
People & Talent
• Actively engaging in stakeholders’ conversations, providing timely, clear and actionable feedback to deliver solution within timeline. 
Risk Management
• The ability to interpret the Group’s technical and security (ICS) control requirements and information to identify potential risks and key issues based on this information and put in place appropriate controls and measures to mitigate or minimize risk to the central monitoring & observability platform delivery.
Governance
• Awareness and understanding of the eSDLC framework, in which the T&O software delivery operates, and the requirements and expectations relevant to the role. 
• Responsible for adhering to the effectiveness of the central monitoring and observability platform deliver governance, based on oversight and controls of the eSDLC framework.
Regulatory & Business Conduct
• Display exemplary conduct and live by the Group’s Values and Code of Conduct. 
• Take personal responsibility for embedding the highest standards of ethics, including regulatory and business conduct, across Standard Chartered Bank. This includes understanding and ensuring compliance with, in letter and spirit, all applicable laws, regulations, guidelines and the Group Code of Conduct.
• Effectively and collaboratively identify, escalate, mitigate and resolve risk, conduct and compliance matters.

Key stakeholders
• Global Head, Markets Production Management 
• Lead, Markets Site Reliability Engineering 
• Markets PSS Leads 
• Markets PSS Managers
• Head, Observability 
Other Responsibilities
• Embed Here for good and Group’s brand and values in the Observability Platform Team; Perform other responsibilities assigned under Group, Country, Business or Functional policies and procedures; Multiple functions (double hats).

Skills and Experience

Strong experience with several of the following technologies:
• Prometheus and Alert Manager
• Grafana
• Open Telemetry (Metrics, Logs, Tracing)
• Elastic Stack or equivalent observability platforms
• Application Performance Monitoring (APM) tools
• Synthetic monitoring solutions
• Kafka / Confluent Kafka
• ServiceNow ITOM and Event Management
• Azure, AWS, AKS, EKS, or Kubernetes-based platforms
• Terraform, Ansible, Chef, Puppet, or similar automation tools
• Experience with Shell scripting, Java, Python or Ruby
• Experience with Web Technologies (Apache, HTML, JavaScript, HTTP, XML)
Programming and scripting experience in one or more of:
• Python
• Java
• Go
• Shell scripting
• PowerShell
Reliability Engineering Competencies
• Advanced troubleshooting skills across applications, middleware, cloud infrastructure, databases, networking, and distributed systems.
• Production incident management and major incident support experience.
• Root cause analysis and problem management expertise.
• Capacity planning, performance engineering, and resiliency testing experience.
• Strong understanding of operational risk and production stability principles.

AI & Automation Experience (Preferred) Experience in one or more of the following areas would be highly advantageous:
• Microsoft Copilot Studio and Copilot extensibility.
• AI agent development and orchestration.
• Generative AI platforms and Large Language Models (LLMs).
• AIOps platforms and intelligent event management.
• Retrieval-Augmented Generation (RAG) architectures.
• Operational knowledge management and search platforms.
• AI-enabled observability and incident management solution

• Reliability Engineering
• Observability Platforms 
• Software Engineering
• Cloud Computing 
• Automation Engineering 
• Incident Management

Qualifications

• Bachelor's Degree in Computer Science, Engineering, Information Systems, or equivalent practical experience.
• Minimum 5+ years of IT experience, including at least 2 years in SRE, Production Engineering, Platform Engineering, DevOps, or a similar reliability-focused role.
• Experience supporting critical production applications within financial services, capital markets, trading platforms, or similarly demanding environments.
• Strong understanding of distributed systems, cloud-native architectures, and large-scale production operations.
• Experience with software engineering principles, source control, CI/CD pipelines, and deployment automation.
• Education: Degree
• Training: Agile Delivery, SRE
• Certifications: Any Professional certifications in SRE, monitoring and observability platforms such as ElasticSearch, Grafana, or ITRS Geneos:
• Certified Kubernetes Administrator (CKA)
• Kubernetes and Cloud Native Associate (KCNA)
• Certified Administrator for Apache Kafka
• Red Hat Certified Specialist in Event-Driven Development with Kafka
• AWS Certified SysOps Administrator – Associate
• Language: English

About Standard Chartered

We're an international bank, nimble enough to act, big enough for impact. For more than 170 years, we've worked to make a positive difference for our clients, communities, and each other. We question the status quo, love a challenge and enjoy finding new opportunities to grow and do better than before. If you're looking for a career with purpose and you want to work for a bank making a difference, we want to hear from you. You can count on us to celebrate your unique talents and we can't wait to see the talents you can bring us.

Our purpose, to drive commerce and prosperity through our unique diversity, together with our brand promise, to be here for good are achieved by how we each live our valued behaviours. When you work with us, you'll see how we value difference and advocate inclusion.

Together we:

What we offer

In line with our Fair Pay Charter, we offer a competitive salary and benefits to support your mental, physical, financial and social wellbeing.

Apply now Information at a Glance

Similar Jobs

Lead, SRE EngineerGuangzhou, CHN TechnologyRequisition Number58871ChinaAVP, Engineering Lead, WRB TechBangalore, IND TechnologyRequisition Number58903IndiaAVP, Chapter lead, Engineering - LARCBangalore, IND TechnologyRequisition Number58509IndiaAssoc Director, Engineering Lead, WRB TechSingapore, SGP TechnologyRequisition Number55589Singapore Job SkillsKey SkillsDevOpsWindows PowerShellAmazon Web ServicesScalabilityCloud ComputingInformation TechnologySoftware EngineeringCloud EngineeringKubernetesMicrosoft AzureContinuous IntegrationRelevant SkillsProduction ManagementCapacity ManagementApache HTTP ServerRisk AssessmentBusiness EfficiencyPredictive Data AnalysisShell ScriptExtensible Markup Language (XML)Incident ManagementMedical SurveillanceJava (Programming Language)ElasticsearchSelf MotivationCapital MarketsBrand ManagementTelemetryRed Hat Enterprise LinuxOperational ExcellenceHTMLPerformance EngineeringInnovationTooling Assembly and DismantlingAccident AnalysisManufacturing EngineeringGeographyDashboardsContinuous TrainingArtificial IntelligenceMicrosoft Power AutomateTesting SkillsCatalyst (Software)Computer ProgrammingRisk AnalysisMttrNetworking SkillsMonitoring of SystemsIncident ResponseProduct Family EngineeringPython (Programming Language)WorkflowsService Level ManagementKnowledge ManagementBusiness StrategiesAnsibleAutomationGovernanceCapacity PlanningApache KafkaBusiness Continuity and Disaster RecoveryGrafanaEnglishOperational Risk ManagementDatabasesAssertivenessDisaster RecoveryTechnical ManagementElectronic Trading PlatformTraceabilityApplication Performance ManagementRoot Cause AnalysisAzure AKSArchitectureEvent ManagementElearningMental HealthLarge Language ModelsSite Reliability Engineering PracticesKnowledge of FinanceRubyProblem IdentificationDeployment AutomationInformation SystemsJavaScript (Programming Language)Infrastructure ManagementProduction SupportHypertext Transfer Protocols (HTTP)Generative AIFault ToleranceDistributed SystemsReliabilityBudgeting SkillsPerseverancePrometheusMiddlewareProblem SolvingSystems Development Life CycleMetricsAdministrative Decision-MakingReliability EngineeringMore

Skills Matching

Upload Your CVUse your CV to see how well your skills match up with this jobGet Started

Encuentra más ofertas como esta

Explora más ofertas activas de esta empresa o crea una cuenta en Insider Jobs para buscar, guardar y seguir oportunidades en todo el job board.