- Location
- HK
- Workplace
- Remote, Hybrid
- Type
- Full-time
- Department
- IT
- Experience
- 8+ years
- Education
- Bachelor
- Closing date
- Today
- Source
- Vincere
Description
Job Title: Head of Group IT Infrastructure & Operations
Reports to: Group IT Director Direct Reports: 4+ Team Size: 10+ (internal + managed services) Location: Hong Kong
1. Role Overview
As the Group Head of IT Infrastructure & Operations, you are the accountable for the stability, performance, and evolution of the group's global technology foundation—from the data centre to the end-user's device.
This encompasses the full breadth of the I&O function: service delivery and operational reliability (ensuring the availability, performance, and resilience of all hybrid-cloud, network, compute, and storage platforms); infrastructure engineering and modernisation (driving architecture, automation, and technical debt reduction to enable business agility); employee experience and support (owning the service desk as the strategic front door to IT, not a cost centre); and governance and resilience (managing technology lifecycle, vendor performance, capacity planning, and business continuity frameworks across the group).
Your success is measured by service reliability, operational efficiency, team capability, and business stakeholder satisfaction.
2. Why This Role
This is a working executive role – you are not a detached visionary, nor are you a ticket-taker. You are the architect of operational discipline and the driver of strategic change, leading a talented team to deliver a world- class, reliable, and future-ready technology platform—and an exceptional support experience—for the entire group.
3. Key Responsibilities
A. Strategy, Roadmap & Transformation
• Define and own the multi-year infrastructure and operations roadmap, balancing aggressive modernization (cloud migration, automation, Dev/AIOps, and AI readiness) with the imperative of maintaining a secure, stable core
• Lead major transformation programmes from business case to operational handover, including data centre consolidation, global network refreshes, and the scaling of high-throughput AI compute and storage foundations.
• Evaluate emerging technologies and make risk-aware recommendations, delivering tangible business value while preparing core network and routing architectures for low-latency AI performance.
• Implement advanced infrastructure cost-management and FinOps frameworks, introducing granular tracking to accurately monitor, forecast, and optimize the unique computing expenses of AI workloads
B. Operational Governance & Service Excellence
• Establish and govern a group-wide operating model – defining clear roles, responsibilities, and standardised processes (ITIL-aligned) across all regions and entities.
• Own the service lifecycle: set SLAs/OLAs, monitor performance through executive dashboards, and drive a culture of continuous service improvement.
• Server & Compute Lifecycle Management – Govern the end-to-end lifecycle of physical and virtual server infrastructure, covering hardware standardisation, firmware/patching cadence, warranty refresh cycles, performance tuning, capacity headroom planning, and decommissioning – ensuring operational hygiene, cost efficiency (TCO), and alignment with business growth demands.
• Network Infrastructure Governance – Direct global network operations across SD-WAN, MPLS, LAN, Wi-Fi, firewalls, load balancers, and WAN optimisation – ensuring resilient, low-latency connectivity across data centres, regional offices, cloud platforms, and remote workforces. Oversee circuit redundancy, bandwidth utilisation, QoS policies, routing protocols, and network hardware lifecycle (switches, routers, firewalls) to maintain zero-downtime connectivity and support real-time business applications.
• Service Desk Oversight – Set the strategic direction for the regional service desk, defining tiered support models (T0–T3), channel strategies (chat, portal, phone, walk-up), and support hour coverage.
• Govern service desk performance against key metrics—First Contact Resolution (FCR), Average Handle Time (AHT), and Customer Satisfaction (CSAT)—and drive initiatives to continuously improve the employee experience.
• Champion self-service and knowledge management to deflect repetitive tickets and ensure seamless escalation pathways from the service desk into infrastructure engineering teams.
• Act as the primary escalation point for major incidents (P1/P2) – mobilising cross-functional resources, driving business communication, and ensuring rapid, decisive closure.
• Institutionalise a rigorous Problem Management discipline – ensuring systemic issues are addressed at their root, not merely fire-fought.
• Drive capacity planning and lifecycle management – ensuring infrastructure refresh cycles (servers, network hardware, storage, endpoints) are funded, scheduled, and executed without service disruption.
C. Financial & Vendor Stewardship
• Own the group's infrastructure budget (CAPEX/OPEX) – including forecasting, cost optimisation, and value reporting to leadership and finance.
• Implement specialized cost-management tools and FinOps frameworks to monitor, forecast, and optimize cloud and AI-related computing expenses.
• Lead strategic vendor relationships and master contract negotiations with key partners (cloud providers, SIs, networking and hardware vendors, and outsourced service desk providers if applicable).
E. Security, Risk & Business Continuity
• Partner with the cyber security team to ensure infrastructure security controls (patching, IAM, network segmentation) are operationally effective and compliant with regulatory requirements.
• Own the IT Disaster Recovery strategy and support the group’s Business Continuity Planning – ensuring recovery capabilities are tested, documented, and aligned with business RTO/RPO requirements.
F. M&A Integration
• Lead infrastructure due diligence and integration for acquisitions – ensuring new entities are onboarded into the group's standard operating model (including service desk onboarding and user adoption) with minimal business disruption.
4. Qualifications & Experience
• Education: Bachelor's degree in IT, Computer Science, or Engineering
• Experience:
o 15+ years in IT infrastructure and operations, with at least 8 years in a managerial capacity overseeing global/regional, multi-site environments (2,000+ users).
o Proven track record of leading large-scale cloud migrations (Azure/Alicloud) in complex, regulated environments.
o Demonstrated experience managing multi-million-dollar budgets and leading strategic vendor negotiations.
o Solid technical grounding across networking, virtualisation, storage, cloud platforms, and end- user computing – sufficient to challenge and guide your technical teams.
o Experience overseeing or transforming global/regional service desk operations (in-house or outsourced)
• Certifications (desired): ITIL Expert / Master, Azure/AWS/Alicloud Solutions Architect, PMP, CISSP.
5. Core Competencies
• Balanced Leadership: Strategically minded but operationally grounded – you speak fluently to both the management and the engineering floor.
• Decisive Under Pressure: Calm and authoritative during major outages, able to cut through noise and align stakeholders on the path forward.
• Commercial Pragmatism: You prioritise investments based on business impact, not technical novelty.
• Influential Communication: Equally comfortable presenting a cost-benefit analysis to the management
• Employee-Centric Mindset: You view the service desk as the voice of the user and champion initiatives that make technology effortless for the workforce.
6. Success Metrics (First 3-6 Months)
• Define and secure approval for the group's infrastructure transformation roadmap, with clear milestones and funding.
• Transform the service desk – implement a modern self-service portal, achieve ≥70% FCR, and establish measurements and capture customer satisfactions.
• Establish a formalised infrastructure governance forum (operating model) that unify group and business lines infrastructure into a consistent standard.
• Maintain operational stability across all infrastructure domains, meeting or exceeding availability targets while driving continuous reduction in incident frequency and severity.