- Location
- Singapore
- Type
- Full-time
- Department
- Engineering
- Closing date
- Today
- Source
- CareersPage
Description
About the Role
We are looking for a hands-on Storage Engineer to operate and maintain Ceph-based storage services supporting our on-premises Kubernetes and OpenShift platforms.
The role focuses on the day-to-day operation, reliability, and continuous improvement of production storage services, including monitoring, maintenance, upgrades, capacity management, hardware replacement, performance troubleshooting, and failure recovery.
This is primarily an operations and engineering role, rather than a position focused on developing a new distributed storage solution. Candidates with experience operating or engineering distributed storage platforms are encouraged to apply.
You will work closely with Kubernetes platform, network, infrastructure, and application teams to ensure storage services remain reliable, scalable, and supportable.
Key Responsibilities
- Operate and maintain Ceph and OpenShift Data Foundation (ODF) environments.
- Monitor cluster health, capacity, latency, throughput, placement groups, and storage device status.
- Manage OSD and node replacement, recovery, rebalancing, backfill, scrubbing, and routine maintenance activities.
- Diagnose and resolve degraded placement groups, slow operations, quorum issues, device failures, and network-related storage issues.
- Plan and perform Ceph and ODF upgrades, patching, capacity expansion, and configuration changes.
- Support block and shared-file storage using ODF, Ceph CSI, RBD, and CephFS.
- Troubleshoot persistent volume provisioning, attachment, mounting, expansion, and storage performance issues.
- Maintain monitoring dashboards, alerts, operational runbooks, capacity plans, and recovery procedures.
- Test and validate recovery procedures for disk, node, service, and network failures.
- Coordinate server, disk, firmware, and network maintenance activities with infrastructure teams.
- Participate in production incident response, troubleshooting, and root-cause analysis.
- Automate routine storage operations and maintenance activities where appropriate.
Requirements
- Hands-on experience operating and supporting Ceph environments.
- Experience with storage monitoring, capacity planning, upgrades, expansion, maintenance, and component replacement.
- Experience benchmarking storage workloads and analysing performance bottlenecks.
- Strong troubleshooting skills across storage software, Linux, networking, and physical hardware.
- Good understanding of HDD, SSD, NVMe, HBA, firmware, and storage networking.
- Experience with automation or scripting using Ansible, Python, Shell, or similar tools.
- Knowledge of backup, snapshots, replication, recovery, and disaster recovery operations.
- Experience supporting production environments and responding to storage-related incidents.
- Experience with Kubernetes, OpenShift, or OpenShift Data Foundation will be highly advantageous.