Volver a ofertas

Director, Head of Technology Resilience and Production Operations

MUFG · London · United Kingdom

Lead the technology resilience and production operations at MUFG in London, shaping strategy and ensuring business continuity. Manage a global team, oversee incident, change, and service continuity, driving automation and regulatory compliance.

Vista previa generada por IA

Lead the technology resilience and production operations at MUFG in London, shaping strategy and ensuring business continuity.

Manage a global team, oversee incident, change, and service continuity, driving automation and regulatory compliance.

MAIN PURPOSE OF THE ROLE

The Head of Technology Resilience and Production Operations, Director is accountable for defining and executing the Technology Resilience and Production Operations strategy to strengthen MUFG’s technological service stability, mitigating risks to business continuity whilst handling the complexities of the business landscape

The role holder and will work collaboratively with colleagues in EMEA and across the globe to create an Enterprise Technology Resilience and Production Operations Centre of Excellence platform function to further enhance resilience and existing production operations

The Head of Technology Resilience and Production Operations will strengthen the resiliency and control landscape through the automation of embedded controls and with an agile approach to delivering resilience enhancements, improving service stability and security, increasing velocity and speed with the ability to commoditise on a global scale

The role holder will heighten customer satisfaction through the protection of revenue, safeguarding MUFG’s reputation and vision to become ‘the worlds most trusted Bank’ and ensuring continued regulatory confidence

The role holder will work with the Head of Digital Engineering Services and Solutions Department Head to deliver transformation through a reliable, robust, sustainable, scalable and efficient operating model by leveraging best practices

The role holder will be accountable for launching and promoting Technology Resilience and Production Operation offerings to MUFG business lines including communications, education and service readiness. They will own and lead the investment planning of Technology Resilience related projects and programmes measuring the effectiveness of these services delivered

This is a Leadership position and an integral part of the Digital Engineering Solutions and Services Leadership team maintain compliance and regulatory obligations.

NUMBER OF DIRECT REPORTS

TBC - Team Size circa 25

KEY RESPONSIBILITIES

Planning & Strategy

Own and execute a comprehensive Technology Resilience and Production Operations strategy to scale and commoditise the platform where necessary whilst meeting regulatory obligation

Accountable for leading the strategic direction of Technology Resilience, service design and delivery performance to protect the MUFG business and enable the business and IT to respond to disruption whilst continuing with critical business operations

Accountable for leading the strategic direction for safeguarding the organisation’s infrastructure and applications by proactively identifying, assessing, and remediating security vulnerabilities

Accountable for defining, leading and governing the Change and Release Management strategy across the enterprise translating business strategy into technological action by ensuring digital transformation efforts are stable

Accountable of leading the strategic direction of Incident, Problem and Event Management to maintain operational resilience

Provide leadership and insight on future tooling, technologies, promoting adoption of these within regions to support the ongoing modernisation of the Technology Resilience and Production Operations enterprise function

Own and manage cost and feasibility while planning and delivering Technology Resilience and Production Operations platform services

Provide leadership and subject matter expertise on Technology Resilience and Product Operation technologies and adoption of these across multiple regions and business lines to support the ongoing modernisation of the function, staying ahead of technological advancements through knowledge and experience through internal and external networks

Identify opportunities and implement strategies for efficiency enhancement, functionalities and service quality improvement, whilst adopting technology advancements where possible

Own and manage the Technology Resilience and Production Operations platform budget, licensing and vendor relationships to ensure cost efficiency, optimisation and Return on Investment

Identify opportunities and implement strategies for efficiency enhancement, functionalities and service quality improvement, whilst adopting technology advancements where possible

Work with Head of Digital Engineering Services and Solutions to deliver a reliable, robust, sustainable and efficient operating model by leveraging best practices

Ensure relationships with EMEA Technology functions and business clients are actively managed to fully deliver to business requirements

Centre of Excellence, Platform Management and Governance

Accountable for building and implementing an Internationalised Technology Resilience and Production Operations Centre of Excellence Platform ecosystem that is robust, secure and resilient

Accountable for Manage and overseeing successful platform upgrades pertaining to the Technology Resilience and Production Operations function

Accountable for deliver best practice operations and strategic development, shaping the department’s security posture while adhering to MUFG policies and procedures

Accountable for resolving platform escalations with the autonomy and experience to make final decision on prioritisation, resourcing and design configurations for tooling, products and systems pertaining tot the Technology Resilience and Production Operations function

Accountable for delivering high quality management information and dashboards that provide a comprehensive view of the Technology Resilience and Production Operations estate. Deliver senior stakeholder communications by translating complex technical details into clear, accessible insights

Service Continuity, Design and Performance

Accountable for designing, implementing and management of the Technology Resilience framework to ensure systems can adapt, recover and withstand disruption for service continuity and business performance aligning with the Operational Resilience and Business Continuity divisions outside of Technology

Own and execute the strategic planning, governance and testing of IT systems to ensure business resilience

Accountable for designing, implementing and managing Technology service design, acceptance and continuous performance of Technology Production operations

Accountable for ensuring new technology or changed services establish document standards for availability, capacity, security and continuity to ensure they operate at peak performance once live.

Accountable for tracking the operational readiness for in-scope projects, to ensure new services and changes are fully “ready” to be supported maintained and used by the business and Technology without causing disruptions

Accountable for defining and enforcing service acceptance criteria and conducting readiness assessments to ensure new olive systems meet security, reliability and supportable standards

Accountable for defining Key Performance Indicators and metrics (e.g system uptime, response times) with the Business and Operational Resilience to then tracking Technology service performance and health against Service Level Agreements

Accountable for providing regular data for senior management that includes Technology and KPI, SLA and risk management MI to Senior Executives in the form of executive dashboard and automated alerts to drive investment and governance decisions to meet business objectives

Accountable for service analysis of performance data to identify bottlenecks and lead initiatives that enhance services efficiency and user satisfaction

Accountable for business impact analysis to identify vulnerabilities, critical assets and definition recovery time/point objectives

Accountable for identifying Technology business continuity risk, creating Technology Division Business Continuity Plans in alignment with Product and Platform owners to identify necessary risk mitigation, including managing concentration risk and ensuring SaaS recovery assurance

Accountable for orchestrating and leading complex, end to end Disaster Recovery exercises, including cycler attack simulations, ransomware scenarios and infrastructure failures to maintain Technology service continuity and resilience in the event of a disaster or critical event

Incident, Problem and Event Management

Own, execute and driving the Incident, Problem and Event Management associated procedures to resolution using strong facilitation, planning and time management

Accountable for responding to “event” (changes in Technology state), collaborating with the Observability function Lead and monitoring functions to mitigate service-impacting incidents

Accountable for providing major Incident cover, coordinating resolution efforts to ensure timely incident management

Accountable for the assessment and prioritisation of multiple incidents based on the customer, business, regulatory, reputational and financial impacts, knowing when to escalat et without sacrificing SLA commitments

Accountable for communicating the incident status, resolution and impacts to internal and external stakeholders clearly and concisely, including gathering relevant information to communicate to regulators

Accountable for Postmortem analysis with accountable parties to identify root cause and deliver eradication actions with the correct ownership

Drive a culture that reduces repeat incidents, shared learning, identifying thematic root causes, impacts and actions to drive improved decision making. Owns and executes enhance capabilities in root cause identification, implementing proactive Problem Management practices

Accountable for delivering high quality management information and dashboards that provide a comprehensive view of the Technology Resilience and Production Operations estate. Deliver senior stakeholder communications by translating complex technical details into clear, accessible insights

Change Management

Accountable for the design, implementation and management framework for the introduction of technology changes and software releases into the live environment.to meet business goals whilst maintaining operational resilience

Accountable for the execution of the enterprise Change Advisory Board to successfully manage risk, business impact and the authorisation of Technology change to ensure all changes meet regulatory and audit requirements

Accountable for ensuring all changes are authorised, documented and appropriately assessed for risk

Accountable for maintaining forward schedule of change and deployments manging conflicts and conflicts between multiple technology stacks and ensuring high availability of systems (uptime)

Accountable for release orchestration, coordinating within complex deployment pipelines integrating traditional change with agile methodology to maintain speed and stability

Accountable for evaluating the risks to the integrity of the Technology service environment, including security, availability and performance impact, ensuring comprehensive rollback plans are in place for all deployments

Own and execute automation across the delivery pipeline to improve efficiency and reduce manual risk

G overnance

Own and oversee the health, performance and operational integrity of Technology Resilience and Production Operations Platforms and products

Ensure accurate reporting of all KPIs/KRIs for the Technology Resilience and Production Operations Platforms and products, systems and tooling

Accountable for defining, refining and communicating frameworks, policies and procedures, ensuring alignment and collaboration with third parties for consistent delivery

Identify key risks within the Technology Resilience and Production Operations team, assessing and mitigating those risks in accordance with the appropriate risk management framework

Ensure the Technology Resilience and Production Operations function operates in a controlled way in accordance with standard procedures by adhering to related security and compliance procedures

Interpret relevant regulatory, requirements and security, design and technical best practices and translate these to business aligned cloud programme requirements. In conjunction with EMEA Technology Risk, Security and Control, ensure that all regulatory requirements are fully complied with, including DORA assessments and appropriate defences and controls are in place to deal with all cyber risks

Project/ Change

Accountable for tracking the operational readiness for in-scope projects, to ensure new services and changes are fully “ready” to be supported maintained and used by the business and Technology without causing disruptions

Accountable for defining and enforcing service acceptance criteria and conducting readiness assessments to ensure new olive systems meet security, reliability and supportable standards

Accountable for complex, large scale Technology Resilience and Production Operations Implementations, ensuring quality, security and compliance

Responsible for the definition, documentation and safe technical execution of projects, working closely with the allocated Programme / Project Manager.  Actively participating in all phases of the project

Analyse accuracy of technical demands, work in progress, take action to ensure targets are met within safety and quality procedures, including hand-over to business and operational supports teams where appropriate

Provide effective investigation and technical resolutions for any issues that may occur as part of the project lifecycle

Investigate potential and actual service problems that may arise from implementation and recommend solutions. Follow formal procedures to plan and test proposed solutions

Manage the technical delivery of projects within the agreed technical scope, cost and timescale across Bank and Securities

Provide specialized technical guidance, establish technical delivery goals and set reasonable delivery timeframes

Interpret relevant regulatory, requirements and security, design and technical best practices and translate these to business aligned Digital Engineering Services and Solutions programme requirements

Leadership, Culture and People Management

Accountable for leadership across the Technology Resilience and Production Operations function, ensuring that the team have the appropriate capabilities to be successful in their roles.

Accountable for development of team with a strong focus on building future capabilities

Lead and champion MUFG’s inclusive, diverse, and values-led culture while fostering a growth mindset to embrace new technologies, industry advancements, and innovative use cases

Actively manage performance, develop talent, identify key positions and persons and create sustainable success plans

Ensure appropriate training is in place to fulfil current and future skill requirements

Lead and promote a dynamic, delivery driven culture that works alongside business units to provide responsive resolutions and value driven solutions

SKILLS AND EXPERIENCE

Essential

Strong understanding of IT resilience frameworks, Cloud Resilience, Operational Resilience with strong experience managing critical service mapping and impact tolerances

Proven Experience setting and executing resilience strategy aligned to business priorities with enterprise wide thinking linking resilience to business continuity, customer impact and financial risk

Proven experience successfully leasing responsiveness to audits regulatory reviews and resilience testing scenarios

Exceptional ability to influence senior stakeholders across CIO, Risk, Compliance, Audit and Business executives

Extensive and proven experience of leading associated teams across regions to manage Incidents, Problems and Events that have compromised Technology services

Extensive experience leading Priority 1 system outages within global and regulated environments across financial services

Experience conducting root cause investigations and applying problem management methodologies to understand why issues occur and to prevent recurrence

Proficient in Trend Analysis and statistical modelling to predict system failures

Proven extensive experience leading Event, Incident, Problem and Change management teams/ departments with demonstrable proactive functional enhancement to Incident, Problem, event and change management frameworks and practices

Proven extensive experience working within infrastructure environments and cloud platforms (AWS, Azure, Oracle), with a high-level understanding of platforms, operating systems, and technologies and software lifecycles

Strong analytical and problem-solving skills to analyse data, identify patterns and develop effective solutions to mitigate risk

Extensive and proven experience of managing Change/Release functions within a regulated financial services environment with high success rates for changes (CAB approval accuracy), increase in release velocity and reductions in Change related incidents

Advanced proficiency in ServiceNow for Incident and problem tracking, analysis and reporting functions

Highly Desirable

Proven experience leading strategic security initiatives and process automation in large-scale environments

EDUCATION / QUALIFICATIONS/ TECHNICAL COMPETENCIES

Highly Desirable

ITIL 4

Proficient in ITSM tools (Extensive experience of ServiceNow, Jira Service Management)

Knowledge of:

Ivanti LANDesk, Qualys, Splunk

Windows Server/Desktop, RHEL/OEL Linux

PowerShell and Python scripting

Familiarity with:

CyberArk PAM, ServiceNow SecOps Vulnerability Response / Application Vulnerability

Agile and DevOps Methodologies /Certifications

Knowledge of ISO22301 / ISO/IEC 20000 standards

CBCP certified Business Continuity

MBCI Member of the Business Continuity Institute

Desirable

GenAI & Now Assist (Generative AI)

Familiarity with Site Reliability Engineering Principles

PERSONAL REQUIREMENTS

Strong decision-making skills with a structured and logical approach to work with a data driven and KPI focuses mindset

Enterprise wide and strong strategic thinking to link resilience to business risk, customer impact and financial outcomes

Excellent communication skills to bridge the gap between Technology and the business, influence across functions to build credibility across both technical and non-technical audience

Excellent interpersonal and leadership skills to confidently challenge constructively whilst maintain alignment

A curious and motivated approach with the ability to explore new solutions to embed a culture of proactive risk identification and mitigation

Ability to remain calm, decisive and structures with the ability to lead in a pressurised environment, making timely risk-based decisions under pressure delivering clan, authoritive updates during crisis

Legal Notice

We are open to considering flexible working requests in line with organisational requirements.

MUFG is committed to embracing diversity and building an inclusive culture where all employees are valued, respected and their opinions count. We support the principles of equality, diversity and inclusion in recruitment and employment, and oppose all forms of discrimination on the grounds of age, sex, gender, sexual orientation, disability, pregnancy and maternity, race, gender reassignment, religion or belief and marriage or civil partnership.

We make our recruitment decisions in a non‑discriminatory manner in accordance with our commitment to identifying the right skills for the right role and our obligations under the law.

Encuentra más ofertas como esta

Explora más ofertas activas de esta empresa o crea una cuenta en Insider Jobs para buscar, guardar y seguir oportunidades en todo el job board.