Hiring.Camp

Senior Manager, Reliability

Dayone

·

Yesterday

Location
NTP Campus (JB), Malaysia
Type
Full-time
Seniority
Senior
Experience
8+ years
Source
Workday

Description

Join DayOne – Shaping the Future of Data Infrastructure

DayOne is a global leader in the development and operation of high-performance data centers. As one of the fastest-growing companies in the industry, we’ve built a robust presence across Asia and Europe — and we’re just getting started.

As we expand into new international markets, we’re looking for talented, driven individuals to join us on this exciting journey. This is more than a job — it’s an opportunity to be a key contributor to our dynamic team and help shape the future of global data infrastructure.

If you're passionate about innovation, technology, and growth, we invite you to be part of DayOne’s next chapter.

This position represents an in-country operational leadership role within the DayOne Operations team. The incumbent will support the Reliability Director of GERA by providing technical leadership, operational governance, and execution oversight for reliability initiatives across the designated data centre portfolio.

The Manager / Senior Manager will be entrusted with ensuring the safe, reliable, efficient, and scalable operation of mission-critical data centres.

Core responsibilities encompass:

  • taking care equipment reliability and performance Conducting comprehensive operational risk assessments utilizing the Failure Mode and Effects Analysis (FMEA) methodology.
  • Developing and implementing containment measures, permanent corrective actions, and preventive strategies to mitigate recurrence.
  • Coordinating effectively with internal and external stakeholders to achieve optimal operational outcomes.
  • Driving continuous improvement initiatives to strengthen reliability, resilience, and performance across the portfolio.

Role and Responsibilities

Reliability Governance: Oversee and uphold the reliability governance framework across the assigned data centre portfolio. Establish and maintain a comprehensive systemic risk register with clearly defined risk categories and corresponding mitigation measures. Conduct periodic evaluations of the effectiveness of these measures to drive continuous improvement.

Asset management and critical parts management: Consolidate and maintain system specification information and critical parts following Risk Prioritization Number (RPN) methodology.

Problem Management: Lead the problem management process by conducting thorough technical and holistic root cause analyses, implementing Corrective and Preventive Actions (CAPA), and overseeing reviews of major incidents. Establish and maintain the Day One Data Centre Known Error Database (KEDB) to ensure systematic documentation and resolution of recurring issues.

Energy and Water Performance Management: Familiar with ISO 50001 Energy Management Scheme, monitor and analyze performance indicators for the data centre portfolio assigned, and apply PDCA methodology for continuous improvement.

Maintenance Governance: Advance system reliability and availability through the implementation of preventive, predictive, and condition-based maintenance strategies. Evaluate equipment health, performance data, alarms, and failure trends to ensure that all maintenance objectives are appropriately executed and aligned with operational standards.

Operational Engineering Support: Assist with any technical issues arise from operations including customer request, regulatory and legal requirements, design intent fulfillment etc.

Equipment Lifecycle and Obsolescence: Develop and execute strategic plans for major equipment maintenance, overhauls, and cyclic replacements. Manage end-of-life (EOL) and end-of-service-life (EOSL) transitions, and oversee asset refresh programs to ensure sustained reliability, efficiency, and alignment with operational standards.

Vendor and Contractor Management: Coordinate and engage with internal stakeholders, original equipment manufacturers (OEMs), service providers, consultants, and contractors to ensure the consistent delivery of high-quality maintenance services in alignment with organizational standards and operational requirements.

Regional Capability Building: Provide coaching and mentorship to site teams, disseminate best practices, and strengthen regional competencies in the operation of mission-critical environments.

Candidate Requirements

  • Bachelor’s degree in mechanical/ electrical engineering, Building Services Engineering or related discipline.
  • Candidates with SCEM, GMAP, CDCS or other related professional certificates are preferred.
  • Minimum 8-12 years of experience in in managing critical infrastructure and equipment, preferably in data centers environments.
  • Demonstrated experience as a Technical Subject Matter Expert (SME) responsible for the complete electrical and mechanical infrastructure lifecycle.
  • Strong understanding of local and international standards on equipment maintenance scope, frequency and operating parameters.
  • Have in depth knowledge of local regulations in operations and maintenance management.
  • Proven experience leading complex incident response, performing advanced fault diagnosis, protection system analysis, forensic Root Cause Analysis (RCA), and driving permanent corrective and preventive actions across equipment lifecycle and operations.
  • Strong communication and stakeholder management skills.
  • Willingness to travel across Southeast Asia to support site operations, audits, technical certification, commissioning, incident response, and technical reviews.

DayOne is proud to be an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

If you're ready to grow with one of the fastest-moving companies in the data center industry, apply now and be part of our global journey.

Similar Jobs

30

Senior Maintenance Manager , Reliability Maintenance and Engineering

Amazon

Today

Senior Manager, Site Reliability & Operational Resilience

Zelis Careers · US NJ Morristown, United States of America +4

Yesterday

Senior Manager, Maintenance & Reliability Engineering

Avantor · USA-NJ Phillipsburg, United States of America

Yesterday

Sr. Manager Technical Infra Program Management , Capacity Delivery Reliability

Amazon

2 days ago

Sr. Manager, Reliability & Engineering Programs

Coke Florida Careers · Tampa, FL, USA

4 days ago

Senior Manager, Product Development Engineering (Memory Reliability , NAND,Failure Analysis, Root Cause Analysis, Device Physics)

SanDisk · Bengaluru, KA, India

1 week ago

Senior Manager, Site Reliability & Infrastructure Engineering

Aviva the most attractive choice · Canada - Markham ON 10 Aviva Way · Hybrid

1 week ago

Senior Manager, Data & AI Platforms Reliability

AmeriLife has served the needs · Remote, FL, United States of America +47 · Remote

1 week ago

Senior Lab Manager - Reliability

Vertiv · Pune, India

1 week ago

Senior Product Manager, Reliability & Call Quality

Tin Can · Seattle · Hybrid

1 week ago

Senior Manager System Reliability Engineering

Gevernova · Hyderabad TS IN 26, India · Hybrid

1 week ago

Senior Manager System Reliability Engineering

gevernova · Hyderabad TS IN 26, India · Hybrid

1 week ago

Senior Manager, Hardware Reliability & Test

Muonspace · San Jose, CA +1 · Hybrid, Onsite

1 week ago

Senior Manager, Site Reliability Engineering

Finastra · Mississauga - Avebury, Canada

2 weeks ago

Senior Engineering Manager - Enterprise Trust & Reliability

Multiverse · London · Hybrid

2 weeks ago

Senior Manager, Site Reliability Engineering

Oracle · United States, US

3 weeks ago

Senior Maintenance and Reliability Manager

Abbott · United States > Casa Grande : Plant, United States of America

3 weeks ago

Senior Maintenance and Reliability Manager

Abbott · United States > Casa Grande : Plant, United States of America

3 weeks ago

Senior Manager, Site Reliability Engineering

Oracle · Reston, VA, United States, US

3 weeks ago

Senior Program Manager, RME Sort Center Mile Partner, Reliability and Maintenance Engineering, INOPs

Amazon

3 weeks ago

Sr. Reliability Technical Program Manager, Infrastructure Reliability & Quality

Amazon

3 weeks ago

Senior Engineering Manager, Site Reliability

Horizon3Ai · US, Remote · Remote

4 weeks ago

Senior Manager, Reliability Engineering- (Nashville, TN - onsite)

Oracle · Nashville, TN, United States, US

1 month ago

Senior Manager, Foundry Customer Quality & Reliability

Samsung Semiconductor · San Jose, California, United States · Onsite

1 month ago

Senior Manager, Product Reliability

Nvidia · Santa Clara, CA,US, US

1 month ago

Senior Manager, Product Reliability

Nvidia · US, CA, Santa Clara, United States of America

1 month ago

Sr Engineering Manager 2 - Applied Reliability Engineering

Gevernova · Atlanta, United States of America +2

1 month ago

Senior Manager of Site Reliability Engineering - Securitized Products, Production Management - NA

JPMorgan Chase · NY, United States, US

1 month ago

Senior Manager of Site Reliability Engineering - Securitized Products, Production Management - NA

JP Morgan Chase · NY, United States, US

1 month ago

Senior Regional Maintenance & Reliability Manager

Smithfieldfoods · Remote, NC, United States of America · Remote

1 month ago