- Location
- Lahore
- Type
- Full-time
- Department
- Engineering
- Experience
- 5+ years
- Education
- Bachelor
- Closing date
- Today
- Source
- CareersPage
Description
Requirements:
- 5+ years of experience in systems, infrastructure, or SRE engineering, operating production systems at scale.
- Deep Linux troubleshooting skills across the OS, networking, storage, and performance, with hands-on experience working as root on production systems. Experience with Ubuntu is highly relevant, as it is used almost exclusively.
- Hands-on experience operating GPU servers in production, including troubleshooting driver, device, and hardware-level issues, rather than only the workloads running on top of them.
- Practical network troubleshooting experience, including diagnosing physical-layer faults.
- Strong automation mindset with programming skills in Python or a comparable language.
- Experience with configuration management, node provisioning, and infrastructure-as-code (IaC) using Ansible, Terraform, or similar tools.
- Experience building observability and alerting solutions using Grafana and Prometheus.
- Bachelor's degree in Computer Science or equivalent experience.
- Experience operating GPU clusters or AI infrastructure at production scale.
- Production experience with Kubernetes or Slurm; experience with both is a bonus.
- Background in HPC or research computing.
- Familiarity with the NVIDIA GPU stack, InfiniBand/RDMA, and NCCL.
- Experience with CLI-based AI coding agents such as Claude Code, rather than browser-based assistants alone.
- Contributions to open-source projects within the cloud-native, HPC, or AI infrastructure ecosystem.
Responsibilities:
- Own the reliability, availability, and performance of production Linux GPU clusters, covering the operating system, drivers, GPUs, high-speed networking, and storage.
- Lead deep, end-to-end troubleshooting of complex distributed systems, GPU nodes, networking, and storage issues.
- Troubleshoot and resolve GPU rail and NCCL performance issues across multi-GPU and multi-node collective communication paths.
- Diagnose network faults end to end, including configuration, routing, and physical-layer issues such as cabling, transceivers, and link errors.
- Configure and maintain workload managers that schedule customer jobs, including Kubernetes, Slurm, or both, along with the identity, storage, and networking services they depend on.
- Build automation and tooling to eliminate operational toil, using a modern language such as Python or Go to design solutions, review implementations, and redirect approaches when needed.
- Use AI-assisted engineering tools such as Claude to accelerate automation, runbook development, and incident analysis.
- Automate provisioning, image deployment, configuration, and remediation using Ansible and infrastructure-as-code.
- Design and operate observability using Grafana, Prometheus, and Loki, while tuning alerts for meaningful signals and building self-healing capabilities that reduce the need for human intervention.
- Lead incident response, on-call activities, blameless postmortems, and reliability improvements that maintain customer SLAs.
- Partner with Platform and Systems Engineering teams on capacity planning, rollouts, and continuous improvement.
Working Hours:
8 PM - 4 AM