- Location
- Beijing, Beijing,CN, CN · Shanghai, Shanghai,CN, CN
- Type
- Full-time
- Department
- Engineering
- Seniority
- Senior
- Source
- Eightfold
Description
Profile, analyze, and optimize GPU‑accelerated code to improve training and inference performance for large‑scale recommender systems. Design, implement, and maintain high‑performance C++/CUDA components within our core recommendation framework. Develop and execute tests (unit, integration, and performance) to ensure numerical correctness, stability, and regression prevention in GPU workloads. Collaborate closely with CUDA and ML engineers to interpret profiling results, refine designs, and implement optimization strategies. Design and optimize high‑throughput data flows between GPUs, RDMA‑capable NICs, and NVMe SSDs using technologies such as GPUDirect RDMA and GPUDirect Storage. Experience with RDMA (verbs, UCX, or CUDA‑aware MPI) and high‑speed data movement between compute and storage.