Vista previa generada por IA
Lead the technology resilience and production operations at MUFG in London, shaping strategy and ensuring business continuity.
Manage a global team, oversee incident, change, and service continuity, driving automation and regulatory compliance.
MAIN PURPOSE OF THE ROLE
The Head of Technology Resilience and Production Operations, Director is accountable for defining and executing the Technology Resilience and Production Operations strategy to strengthen MUFG’s technological service stability, mitigating risks to business continuity whilst handling the complexities of the business landscape
The role holder and will work collaboratively with colleagues in EMEA and across the globe to create an Enterprise Technology Resilience and Production Operations Centre of Excellence platform function to further enhance resilience and existing production operations
The Head of Technology Resilience and Production Operations will strengthen the resiliency and control landscape through the automation of embedded controls and with an agile approach to delivering resilience enhancements, improving service stability and security, increasing velocity and speed with the ability to commoditise on a global scale
The role holder will heighten customer satisfaction through the protection of revenue, safeguarding MUFG’s reputation and vision to become ‘the worlds most trusted Bank’ and ensuring continued regulatory confidence
The role holder will work with the Head of Digital Engineering Services and Solutions Department Head to deliver transformation through a reliable, robust, sustainable, scalable and efficient operating model by leveraging best practices
The role holder will be accountable for launching and promoting Technology Resilience and Production Operation offerings to MUFG business lines including communications, education and service readiness. They will own and lead the investment planning of Technology Resilience related projects and programmes measuring the effectiveness of these services delivered
This is a Leadership position and an integral part of the Digital Engineering Solutions and Services Leadership team maintain compliance and regulatory obligations.
NUMBER OF DIRECT REPORTS
TBC - Team Size circa 25
KEY RESPONSIBILITIES
Planning & Strategy
Own and execute a comprehensive Technology Resilience and Production Operations strategy to scale and commoditise the platform where necessary whilst meeting regulatory obligation
Accountable for leading the strategic direction of Technology Resilience, service design and delivery performance to protect the MUFG business and enable the business and IT to respond to disruption whilst continuing with critical business operations
Accountable for leading the strategic direction for safeguarding the organisation’s infrastructure and applications by proactively identifying, assessing, and remediating security vulnerabilities
Accountable for defining, leading and governing the Change and Release Management strategy across the enterprise translating business strategy into technological action by ensuring digital transformation efforts are stable
Accountable of leading the strategic direction of Incident, Problem and Event Management to maintain operational resilience
Provide leadership and insight on future tooling, technologies, promoting adoption of these within regions to support the ongoing modernisation of the Technology Resilience and Production Operations enterprise function
Own and manage cost and feasibility while planning and delivering Technology Resilience and Production Operations platform services
Provide leadership and subject matter expertise on Technology Resilience and Product Operation technologies and adoption of these across multiple regions and business lines to support the ongoing modernisation of the function, staying ahead of technological advancements through knowledge and experience through internal and external networks
Identify opportunities and implement strategies for efficiency enhancement, functionalities and service quality improvement, whilst adopting technology advancements where possible
Own and manage the Technology Resilience and Production Operations platform budget, licensing and vendor relationships to ensure cost efficiency, optimisation and Return on Investment
Identify opportunities and implement strategies for efficiency enhancement, functionalities and service quality improvement, whilst adopting technology advancements where possible
Work with Head of Digital Engineering Services and Solutions to deliver a reliable, robust, sustainable and efficient operating model by leveraging best practices
Ensure relationships with EMEA Technology functions and business clients are actively managed to fully deliver to business requirements
Centre of Excellence, Platform Management and Governance
Accountable for building and implementing an Internationalised Technology Resilience and Production Operations Centre of Excellence Platform ecosystem that is robust, secure and resilient
Accountable for Manage and overseeing successful platform upgrades pertaining to the Technology Resilience and Production Operations function
Accountable for deliver best practice operations and strategic development, shaping the department’s security posture while adhering to MUFG policies and procedures
Accountable for resolving platform escalations with the autonomy and experience to make final decision on prioritisation, resourcing and design configurations for tooling, products and systems pertaining tot the Technology Resilience and Production Operations function
Accountable for delivering high quality management information and dashboards that provide a comprehensive view of the Technology Resilience and Production Operations estate. Deliver senior stakeholder communications by translating complex technical details into clear, accessible insights
Service Continuity, Design and Performance
Accountable for designing, implementing and management of the Technology Resilience framework to ensure systems can adapt, recover and withstand disruption for service continuity and business performance aligning with the Operational Resilience and Business Continuity divisions outside of Technology
Own and execute the strategic planning, governance and testing of IT systems to ensure business resilience
Accountable for designing, implementing and managing Technology service design, acceptance and continuous performance of Technology Production operations
Accountable for ensuring new technology or changed services establish document standards for availability, capacity, security and continuity to ensure they operate at peak performance once live.
Accountable for tracking the operational readiness for in-scope projects, to ensure new services and changes are fully “ready” to be supported maintained and used by the business and Technology without causing disruptions
Accountable for defining and enforcing service acceptance criteria and conducting readiness assessments to ensure new olive systems meet security, reliability and supportable standards
Accountable for defining Key Performance Indicators and metrics (e.g system uptime, response times) with the Business and Operational Resilience to then tracking Technology service performance and health against Service Level Agreements
Accountable for providing regular data for senior management that includes Technology and KPI, SLA and risk management MI to Senior Executives in the form of executive dashboard and automated alerts to drive investment and governance decisions to meet business objectives
Accountable for service analysis of performance data to identify bottlenecks and lead initiatives that enhance services efficiency and user satisfaction
Accountable for business impact analysis to identify vulnerabilities, critical assets and definition recovery time/point objectives
Accountable for identifying Technology business continuity risk, creating Technology Division Business Continuity Plans in alignment with Product and Platform owners to identify necessary risk mitigation, including managing concentration risk and ensuring SaaS recovery assurance
Accountable for orchestrating and leading complex, end to end Disaster Recovery exercises, including cycler attack simulations, ransomware scenarios and infrastructure failures to maintain Technology service continuity and resilience in the event of a disaster or critical event
Incident, Problem and Event Management
Own, execute and driving the Incident, Problem and Event Management associated procedures to resolution using strong facilitation, planning and time management
Accountable for responding to “event” (changes in Technology state), collaborating with the Observability function Lead and monitoring functions to mitigate service-impacting incidents
Accountable for providing major Incident cover, coordinating resolution efforts to ensure timely incident management
Accountable for the assessment and prioritisation of multiple incidents based on the customer, business, regulatory, reputational and financial impacts, knowing when to escalat et without sacrificing SLA commitments
Accountable for communicating the incident status, resolution and impacts to internal and external stakeholders clearly and concisely, including gathering relevant information to communicate to regulators
Accountable for Postmortem analysis with accountable parties to identify root cause and deliver eradication actions with the correct ownership
Drive a culture that reduces repeat incidents, shared learning, identifying thematic root causes, impacts and actions to drive improved decision making. Owns and executes enhance capabilities in root cause identification, implementing proactive Problem Management practices
Accountable for delivering high quality management information and dashboards that provide a comprehensive view of the Technology Resilience and Production Operations estate. Deliver senior stakeholder communications by translating complex technical details into clear, accessible insights
Change Management
Accountable for the design, implementation and management framework for the introduction of technology changes and software releases into the live environment.to meet business goals whilst maintaining operational resilience
Accountable for the execution of the enterprise Change Advisory Board to successfully manage risk, business impact and the authorisation of Technology change to ensure all changes meet regulatory and audit requirements
Accountable for ensuring all changes are authorised, documented and appropriately assessed for risk
Accountable for maintaining forward schedule of change and deployments manging conflicts and conflicts between multiple technology stacks and ensuring high availability of systems (uptime)
Accountable for release orchestration, coordinating within complex deployment pipelines integrating traditional change with agile methodology to maintain speed and stability
Accountable for evaluating the risks to the integrity of the Technology service environment, including security, availability and performance impact, ensuring comprehensive rollback plans are in place for all deployments
Own and execute automation across the delivery pipeline to improve efficiency and reduce manual risk
G overnance
Own and oversee the health, performance and operational integrity of Technology Resilience and Production Operations Platforms and products
Ensure accurate reporting of all KPIs/KRIs for the Technology Resilience and Production Operations Platforms and products, systems and tooling
Accountable for defining, refining and communicating frameworks, policies and procedures, ensuring alignment and collaboration with third parties for consistent delivery
Identify key risks within the Technology Resilience and Production Operations team, assessing and mitigating those risks in accordance with the appropriate risk management framework
Ensure the Technology Resilience and Production Operations function operates in a controlled way in accordance with standard procedures by adhering to related security and compliance procedures
Interpret relevant regulatory, requirements and security, design and technical best practices and translate these to business aligned cloud programme requirements. In conjunction with EMEA Technology Risk, Security and Control, ensure that all regulatory requirements are fully complied with, including DORA assessments and appropriate defences and controls are in place to deal with all cyber risks
Project/ Change
Accountable for tracking the operational readiness for in-scope projects, to ensure new services and changes are fully “ready” to be supported maintained and used by the business and Technology without causing disruptions
Accountable for defining and enforcing service acceptance criteria and conducting readiness assessments to ensure new olive systems meet security, reliability and supportable standards
Accountable for complex, large scale Technology Resilience and Production Operations Implementations, ensuring quality, security and compliance
Responsible for the definition, documentation and safe technical execution of projects, working closely with the allocated Programme / Project Manager. Actively participating in all phases of the project
Analyse accuracy of technical demands, work in progress, take action to ensure targets are met within safety and quality procedures, including hand-over to business and operational supports teams where appropriate
Provide effective investigation and technical resolutions for any issues that may occur as part of the project lifecycle
Investigate potential and actual service problems that may arise from implementation and recommend solutions. Follow formal procedures to plan and test proposed solutions
Manage the technical delivery of projects within the agreed technical scope, cost and timescale across Bank and Securities
Provide specialized technical guidance, establish technical delivery goals and set reasonable delivery timeframes
Interpret relevant regulatory, requirements and security, design and technical best practices and translate these to business aligned Digital Engineering Services and Solutions programme requirements
Leadership, Culture and People Management
Accountable for leadership across the Technology Resilience and Production Operations function, ensuring that the team have the appropriate capabilities to be successful in their roles.
Accountable for development of team with a strong focus on building future capabilities
Lead and champion MUFG’s inclusive, diverse, and values-led culture while fostering a growth mindset to embrace new technologies, industry advancements, and innovative use cases
Actively manage performance, develop talent, identify key positions and persons and create sustainable success plans
Ensure appropriate training is in place to fulfil current and future skill requirements
Lead and promote a dynamic, delivery driven culture that works alongside business units to provide responsive resolutions and value driven solutions
SKILLS AND EXPERIENCE
Essential
Strong understanding of IT resilience frameworks, Cloud Resilience, Operational Resilience with strong experience managing critical service mapping and impact tolerances
Proven Experience setting and executing resilience strategy aligned to business priorities with enterprise wide thinking linking resilience to business continuity, customer impact and financial risk
Proven experience successfully leasing responsiveness to audits regulatory reviews and resilience testing scenarios
Exceptional ability to influence senior stakeholders across CIO, Risk, Compliance, Audit and Business executives
Extensive and proven experience of leading associated teams across regions to manage Incidents, Problems and Events that have compromised Technology services
Extensive experience leading Priority 1 system outages within global and regulated environments across financial services
Experience conducting root cause investigations and applying problem management methodologies to understand why issues occur and to prevent recurrence
Proficient in Trend Analysis and statistical modelling to predict system failures
Proven extensive experience leading Event, Incident, Problem and Change management teams/ departments with demonstrable proactive functional enhancement to Incident, Problem, event and change management frameworks and practices
Proven extensive experience working within infrastructure environments and cloud platforms (AWS, Azure, Oracle), with a high-level understanding of platforms, operating systems, and technologies and software lifecycles
Strong analytical and problem-solving skills to analyse data, identify patterns and develop effective solutions to mitigate risk
Extensive and proven experience of managing Change/Release functions within a regulated financial services environment with high success rates for changes (CAB approval accuracy), increase in release velocity and reductions in Change related incidents
Advanced proficiency in ServiceNow for Incident and problem tracking, analysis and reporting functions
Highly Desirable
Proven experience leading strategic security initiatives and process automation in large-scale environments
EDUCATION / QUALIFICATIONS/ TECHNICAL COMPETENCIES
Highly Desirable
ITIL 4
Proficient in ITSM tools (Extensive experience of ServiceNow, Jira Service Management)
Knowledge of:
Ivanti LANDesk, Qualys, Splunk
Windows Server/Desktop, RHEL/OEL Linux
PowerShell and Python scripting
Familiarity with:
CyberArk PAM, ServiceNow SecOps Vulnerability Response / Application Vulnerability
Agile and DevOps Methodologies /Certifications
Knowledge of ISO22301 / ISO/IEC 20000 standards
CBCP certified Business Continuity
MBCI Member of the Business Continuity Institute
Desirable
GenAI & Now Assist (Generative AI)
Familiarity with Site Reliability Engineering Principles
PERSONAL REQUIREMENTS
Strong decision-making skills with a structured and logical approach to work with a data driven and KPI focuses mindset
Enterprise wide and strong strategic thinking to link resilience to business risk, customer impact and financial outcomes
Excellent communication skills to bridge the gap between Technology and the business, influence across functions to build credibility across both technical and non-technical audience
Excellent interpersonal and leadership skills to confidently challenge constructively whilst maintain alignment
A curious and motivated approach with the ability to explore new solutions to embed a culture of proactive risk identification and mitigation
Ability to remain calm, decisive and structures with the ability to lead in a pressurised environment, making timely risk-based decisions under pressure delivering clan, authoritive updates during crisis
Legal Notice
We are open to considering flexible working requests in line with organisational requirements.
MUFG is committed to embracing diversity and building an inclusive culture where all employees are valued, respected and their opinions count. We support the principles of equality, diversity and inclusion in recruitment and employment, and oppose all forms of discrimination on the grounds of age, sex, gender, sexual orientation, disability, pregnancy and maternity, race, gender reassignment, religion or belief and marriage or civil partnership.
We make our recruitment decisions in a non‑discriminatory manner in accordance with our commitment to identifying the right skills for the right role and our obligations under the law.