- Location
- GB · ES · PL · DE · CH · IT
- Workplace
- Remote
- Type
- Full-time
- Department
- Engineering
- Seniority
- Senior
- Closing date
- Today
- Source
- Eightfold
Description
Collaborating with NVIDIA's training framework developers and product teams to stay ahead of the latest features and help partners to adopt them effectively. Assisting with deployment, debugging, and improving the efficiency of AI workloads on extensive NVIDIA platforms. Benchmarking new framework features, analyzing performance, and sharing actionable insights with both customers and internal teams. Working directly with external customers to solve cluster performance and stability issues, identify bottlenecks, and implement effective solutions. Build expertise and guide customers in scaling workloads efficiently and reliably on the latest generation of NVIDIA GPUs. Contributing to Europe's Sovereign AI initiative by helping customers implement advanced resiliency features within AI training pipelines. Experienced in working with large compute clusters with an understanding of their internal scheduling and resource management mechanisms (e.g. SLURM or Cloud based clusters). Proficient knowledge of training pipelines and frameworks, encompassing their internal operations and performance attributes. Experience in debugging training pipelines running on thousands of GPUs in production environment. Ability to debug stability issues across the entire stack: parallel application, training frameworks, runtime libraries, schedulers, and hardware.