Hiring.Camp

Senior Manager, Reliability (GERA)

Dayone

·

Yesterday

Location
CFA Campus (SG), Singapore
Type
Full-time
Seniority
Senior
Experience
8+ years
Source
Workday

Description

Join DayOne – Shaping the Future of Data Infrastructure

DayOne is a global leader in the development and operation of high-performance data centers. As one of the fastest-growing companies in the industry, we’ve built a robust presence across Asia and Europe — and we’re just getting started.

As we expand into new international markets, we’re looking for talented, driven individuals to join us on this exciting journey. This is more than a job — it’s an opportunity to be a key contributor to our dynamic team and help shape the future of global data infrastructure.

If you're passionate about innovation, technology, and growth, we invite you to be part of DayOne’s next chapter.

This position represents an in-country operational leadership role within the DayOne Operations team. The incumbent will support the Reliability Director of GERA by providing technical leadership, operational governance, and execution oversight for reliability initiatives across the designated data centre portfolio.

The Manager / Senior Manager will be entrusted with ensuring the safe, reliable, efficient, and scalable operation of mission-critical data centres.

Core responsibilities encompass:

  • taking care equipment reliability and performance Conducting comprehensive operational risk assessments utilizing the Failure Mode and Effects Analysis (FMEA) methodology.
  • Developing and implementing containment measures, permanent corrective actions, and preventive strategies to mitigate recurrence.
  • Coordinating effectively with internal and external stakeholders to achieve optimal operational outcomes.
  • Driving continuous improvement initiatives to strengthen reliability, resilience, and performance across the portfolio.

Role and Responsibilities

Reliability Governance: Oversee and uphold the reliability governance framework across the assigned data centre portfolio. Establish and maintain a comprehensive systemic risk register with clearly defined risk categories and corresponding mitigation measures. Conduct periodic evaluations of the effectiveness of these measures to drive continuous improvement.

Asset management and critical parts management: Consolidate and maintain system specification information and critical parts following Risk Prioritization Number (RPN) methodology.

Problem Management: Lead the problem management process by conducting thorough technical and holistic root cause analyses, implementing Corrective and Preventive Actions (CAPA), and overseeing reviews of major incidents. Establish and maintain the Day One Data Centre Known Error Database (KEDB) to ensure systematic documentation and resolution of recurring issues.

Energy and Water Performance Management: Familiar with ISO 50001 Energy Management Scheme, monitor and analyze performance indicators for the data centre portfolio assigned, and apply PDCA methodology for continuous improvement.

Maintenance Governance: Advance system reliability and availability through the implementation of preventive, predictive, and condition-based maintenance strategies. Evaluate equipment health, performance data, alarms, and failure trends to ensure that all maintenance objectives are appropriately executed and aligned with operational standards.

Operational Engineering Support: Assist with any technical issues arise from operations including customer request, regulatory and legal requirements, design intent fulfillment etc.

Equipment Lifecycle and Obsolescence: Develop and execute strategic plans for major equipment maintenance, overhauls, and cyclic replacements. Manage end-of-life (EOL) and end-of-service-life (EOSL) transitions, and oversee asset refresh programs to ensure sustained reliability, efficiency, and alignment with operational standards.

Vendor and Contractor Management: Coordinate and engage with internal stakeholders, original equipment manufacturers (OEMs), service providers, consultants, and contractors to ensure the consistent delivery of high-quality maintenance services in alignment with organizational standards and operational requirements.

Regional Capability Building: Provide coaching and mentorship to site teams, disseminate best practices, and strengthen regional competencies in the operation of mission-critical environments.

Candidate Requirements

  • Bachelor’s degree in mechanical/ electrical engineering, Building Services Engineering or related discipline.
  • Candidates with SCEM, GMAP, CDCS or other related professional certificates are preferred.
  • Minimum 8-12 years of experience in in managing critical infrastructure and equipment, preferably in data centers environments.
  • Demonstrated experience as a Technical Subject Matter Expert (SME) responsible for the complete electrical and mechanical infrastructure lifecycle.
  • Strong understanding of local and international standards on equipment maintenance scope, frequency and operating parameters.
  • Have in depth knowledge of local regulations in operations and maintenance management.
  • Proven experience leading complex incident response, performing advanced fault diagnosis, protection system analysis, forensic Root Cause Analysis (RCA), and driving permanent corrective and preventive actions across equipment lifecycle and operations.
  • Strong communication and stakeholder management skills.
  • Willingness to travel across Southeast Asia to support site operations, audits, technical certification, commissioning, incident response, and technical reviews.

DayOne is proud to be an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

If you're ready to grow with one of the fastest-moving companies in the data center industry, apply now and be part of our global journey.

Similar Jobs

30

Senior Manager- Reliability

First Solar · Perrysburg, OH, United States, US

2 days ago

Senior Manager, Site Reliability Engineering (SRE)

Nium · Bangalore +1 · Hybrid

Yesterday

(USA) Senior Manager I, Reliability

Walmart · (USA) IN MCCORDSVILLE 08972 ECOMM FULFILLMENT SERVICES, United States of America

Yesterday

Senior Technical Product Manager - AI Agents, Evals & Reliability

Bjakcareer · China · Remote

3 days ago

Senior Technical Product Manager - AI Agents, Evals & Reliability

Bjakcareer · Indonesia · Remote

3 days ago

Senior Technical Product Manager - AI Agents, Evals & Reliability

Bjakcareer · United Kingdom · Hybrid

3 days ago

Senior Technical Product Manager - AI Agents, Evals & Reliability

Bjakcareer · Seoul, Korea · Hybrid

3 days ago

Senior Technical Product Manager - AI Agents, Evals & Reliability

Bjakcareer · Singapore · Remote

3 days ago

Sr Manager, Site Reliability Engineering

Fis · IND PUNE FL7, India

3 days ago

Senior Technical Product Manager - AI Agents, Evals & Reliability

Bjakcareer · United States · Remote

3 days ago

Senior Manager, Reliability & Platform Engineering

McMaster-Carr · Chicago, IL (Elmhurst) +1 · Hybrid

1 week ago

Senior Manager, AI Reliability Engineering -Kroger Technology & Digital (P2498)

84.51° · Cincinnati, OH +2

1 week ago

Senior Manager, Site Reliability Engineering

Oracle · Ireland, IE

1 week ago

Senior Manager, Service Reliability and Customer Operations

InterSystems · Dublin

1 week ago

Senior Manager, Site Reliability Engineering – Paylo Platform

PDI Technologies · Alpharetta, GA +3 · Hybrid

1 week ago

Senior Manager, Technology Development Engineering (NAND Device physics, Memory Reliability , NAND Failure Analysis)

SanDisk · Bengaluru, KA, India

1 week ago

Senior Maintenance Manager , Reliability Maintenance and Engineering

Amazon

1 week ago

Senior Manager, Site Reliability & Operational Resilience

Zelis Careers · US NJ Morristown, United States of America +4

2 weeks ago

Senior Manager, Maintenance & Reliability Engineering

Avantor · USA-NJ Phillipsburg, United States of America

2 weeks ago

Senior Manager, Site Reliability Engineering

Oracle · Nashville, TN, United States, US

2 weeks ago

Sr. Manager Technical Infra Program Management , Capacity Delivery Reliability

Amazon

2 weeks ago

Senior Manager, Site Reliability & Infrastructure Engineering

Aviva the most attractive choice · Canada - Markham ON 10 Aviva Way · Hybrid

3 weeks ago

Senior Manager, Data & AI Platforms Reliability

AmeriLife has served the needs · Remote, FL, United States of America +47 · Remote

3 weeks ago

Senior Lab Manager - Reliability

Vertiv · Pune, India

3 weeks ago

Senior Product Manager, Reliability & Call Quality

Tin Can · Seattle · Hybrid

3 weeks ago

Senior Manager System Reliability Engineering

Gevernova · Hyderabad TS IN 26, India · Hybrid

3 weeks ago

Senior Manager System Reliability Engineering

gevernova · Hyderabad TS IN 26, India · Hybrid

3 weeks ago

Senior Manager, Hardware Reliability & Test

Muonspace · San Jose, CA +1 · Hybrid, Onsite

3 weeks ago

Senior Manager, Site Reliability Engineering

Finastra · Mississauga - Avebury, Canada

4 weeks ago

Senior Engineering Manager - Enterprise Trust & Reliability

Multiverse · London · Hybrid

4 weeks ago