- Location
- NTP Campus (JB), Malaysia
- Type
- Full-time
- Seniority
- Senior
- Experience
- 8+ years
- Source
- Workday
Description
Join DayOne – Shaping the Future of Data Infrastructure
DayOne is a global leader in the development and operation of high-performance data centers. As one of the fastest-growing companies in the industry, we’ve built a robust presence across Asia and Europe — and we’re just getting started.
As we expand into new international markets, we’re looking for talented, driven individuals to join us on this exciting journey. This is more than a job — it’s an opportunity to be a key contributor to our dynamic team and help shape the future of global data infrastructure.
If you're passionate about innovation, technology, and growth, we invite you to be part of DayOne’s next chapter.
This position represents an in-country operational leadership role within the DayOne Operations team. The incumbent will support the Reliability Director of GERA by providing technical leadership, operational governance, and execution oversight for reliability initiatives across the designated data centre portfolio.
The Manager / Senior Manager will be entrusted with ensuring the safe, reliable, efficient, and scalable operation of mission-critical data centres.
Core responsibilities encompass:
- taking care equipment reliability and performance Conducting comprehensive operational risk assessments utilizing the Failure Mode and Effects Analysis (FMEA) methodology.
- Developing and implementing containment measures, permanent corrective actions, and preventive strategies to mitigate recurrence.
- Coordinating effectively with internal and external stakeholders to achieve optimal operational outcomes.
- Driving continuous improvement initiatives to strengthen reliability, resilience, and performance across the portfolio.
Role and Responsibilities
Reliability Governance: Oversee and uphold the reliability governance framework across the assigned data centre portfolio. Establish and maintain a comprehensive systemic risk register with clearly defined risk categories and corresponding mitigation measures. Conduct periodic evaluations of the effectiveness of these measures to drive continuous improvement.
Asset management and critical parts management: Consolidate and maintain system specification information and critical parts following Risk Prioritization Number (RPN) methodology.
Problem Management: Lead the problem management process by conducting thorough technical and holistic root cause analyses, implementing Corrective and Preventive Actions (CAPA), and overseeing reviews of major incidents. Establish and maintain the Day One Data Centre Known Error Database (KEDB) to ensure systematic documentation and resolution of recurring issues.
Energy and Water Performance Management: Familiar with ISO 50001 Energy Management Scheme, monitor and analyze performance indicators for the data centre portfolio assigned, and apply PDCA methodology for continuous improvement.
Maintenance Governance: Advance system reliability and availability through the implementation of preventive, predictive, and condition-based maintenance strategies. Evaluate equipment health, performance data, alarms, and failure trends to ensure that all maintenance objectives are appropriately executed and aligned with operational standards.
Operational Engineering Support: Assist with any technical issues arise from operations including customer request, regulatory and legal requirements, design intent fulfillment etc.
Equipment Lifecycle and Obsolescence: Develop and execute strategic plans for major equipment maintenance, overhauls, and cyclic replacements. Manage end-of-life (EOL) and end-of-service-life (EOSL) transitions, and oversee asset refresh programs to ensure sustained reliability, efficiency, and alignment with operational standards.
Vendor and Contractor Management: Coordinate and engage with internal stakeholders, original equipment manufacturers (OEMs), service providers, consultants, and contractors to ensure the consistent delivery of high-quality maintenance services in alignment with organizational standards and operational requirements.
Regional Capability Building: Provide coaching and mentorship to site teams, disseminate best practices, and strengthen regional competencies in the operation of mission-critical environments.
Candidate Requirements
- Bachelor’s degree in mechanical/ electrical engineering, Building Services Engineering or related discipline.
- Candidates with SCEM, GMAP, CDCS or other related professional certificates are preferred.
- Minimum 8-12 years of experience in in managing critical infrastructure and equipment, preferably in data centers environments.
- Demonstrated experience as a Technical Subject Matter Expert (SME) responsible for the complete electrical and mechanical infrastructure lifecycle.
- Strong understanding of local and international standards on equipment maintenance scope, frequency and operating parameters.
- Have in depth knowledge of local regulations in operations and maintenance management.
- Proven experience leading complex incident response, performing advanced fault diagnosis, protection system analysis, forensic Root Cause Analysis (RCA), and driving permanent corrective and preventive actions across equipment lifecycle and operations.
- Strong communication and stakeholder management skills.
- Willingness to travel across Southeast Asia to support site operations, audits, technical certification, commissioning, incident response, and technical reviews.
DayOne is proud to be an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
If you're ready to grow with one of the fastest-moving companies in the data center industry, apply now and be part of our global journey.