Hiring.Camp

Site Reliability Engineer (Onsite, Lahore, PKR Salary)

Hr Pod Hiring Talent Globally

·

Today

Location
Lahore
Type
Full-time
Department
Engineering
Experience
5+ years
Education
Bachelor
Closing date
Today
Source
CareersPage

Description

Requirements:

  • 5+ years of experience in systems, infrastructure, or SRE engineering, operating production systems at scale.
  • Deep Linux troubleshooting skills across the OS, networking, storage, and performance, with hands-on experience working as root on production systems. Experience with Ubuntu is highly relevant, as it is used almost exclusively.
  • Hands-on experience operating GPU servers in production, including troubleshooting driver, device, and hardware-level issues, rather than only the workloads running on top of them.
  • Practical network troubleshooting experience, including diagnosing physical-layer faults.
  • Strong automation mindset with programming skills in Python or a comparable language.
  • Experience with configuration management, node provisioning, and infrastructure-as-code (IaC) using Ansible, Terraform, or similar tools.
  • Experience building observability and alerting solutions using Grafana and Prometheus.
  • Bachelor's degree in Computer Science or equivalent experience.
  • Experience operating GPU clusters or AI infrastructure at production scale.
  • Production experience with Kubernetes or Slurm; experience with both is a bonus.
  • Background in HPC or research computing.
  • Familiarity with the NVIDIA GPU stack, InfiniBand/RDMA, and NCCL.
  • Experience with CLI-based AI coding agents such as Claude Code, rather than browser-based assistants alone.
  • Contributions to open-source projects within the cloud-native, HPC, or AI infrastructure ecosystem.

Responsibilities:

  • Own the reliability, availability, and performance of production Linux GPU clusters, covering the operating system, drivers, GPUs, high-speed networking, and storage.
  • Lead deep, end-to-end troubleshooting of complex distributed systems, GPU nodes, networking, and storage issues.
  • Troubleshoot and resolve GPU rail and NCCL performance issues across multi-GPU and multi-node collective communication paths.
  • Diagnose network faults end to end, including configuration, routing, and physical-layer issues such as cabling, transceivers, and link errors.
  • Configure and maintain workload managers that schedule customer jobs, including Kubernetes, Slurm, or both, along with the identity, storage, and networking services they depend on.
  • Build automation and tooling to eliminate operational toil, using a modern language such as Python or Go to design solutions, review implementations, and redirect approaches when needed.
  • Use AI-assisted engineering tools such as Claude to accelerate automation, runbook development, and incident analysis.
  • Automate provisioning, image deployment, configuration, and remediation using Ansible and infrastructure-as-code.
  • Design and operate observability using Grafana, Prometheus, and Loki, while tuning alerts for meaningful signals and building self-healing capabilities that reduce the need for human intervention.
  • Lead incident response, on-call activities, blameless postmortems, and reliability improvements that maintain customer SLAs.
  • Partner with Platform and Systems Engineering teams on capacity planning, rollouts, and continuous improvement.

Working Hours:

8 PM - 4 AM

Skills

PythonKubernetesTerraformAnsibleLinuxSRE

Similar Jobs

30

DevOps Engineer

BeVera Solutions LLC · Atlanta, GA

Today

Devops Engineer

Miratech · Bengaluru, KA, India · Remote

Today

Release Platform Engineer

bet365 · Manchester, England, United Kingdom · Hybrid

Today

Release Platform Engineer

bet365 · Stoke-on-Trent, England, United Kingdom · Hybrid

Today

Site Reliability Engineer

Fortive · IN +1 · Remote

Today

Site Reliability Engineer

Fluke · Remote, India · Remote

Today

DevOps Engineer

Agdata · Pune, Maharashtra

Today

Site Reliability Engineer

Andurilindustries · Waltham, Massachusetts, United States

Today

Devops Engineer

Playpower Labs · Fully Remote - India Only · Remote

Today

Staff Platform Engineer

Ftmo · Prague office · Onsite

Today

DevOps Engineer

G2I · USA

Today

Technology Platform Engineer

Accenture · Bengaluru, BDC7C, India

Today

Technology Platform Engineer

Accenture · Chennai, CDC2F, India

Today

Senior Platform Engineer

Amadeus · Lisbon - Carnaxide, Portugal · Hybrid

Today

Cloud Platform Engineer

Accenture · Pune, PDC5A, India

Today

Cloud Platform Engineer

Accenture · Bhubaneswar, BBDC1A, India

Today

Data Platform Engineer

Accenture · Bengaluru, BDC9A, India

Today

Cloud Platform Engineer

Accenture · Bengaluru, BDC14A, India

Today

Data Platform Engineer

Accenture · Hyderabad, HDC3B, India

Today

Data Platform Engineer

Accenture · Hyderabad, HDC4A, India

Today

Data Platform Engineer

Accenture · Pune, PDC3B, India

Today

Technology Platform Engineer

Accenture · Bengaluru, BDC11A, India

Today

Technology Platform Engineer

Accenture · Bengaluru, BDC11A, India

Today

Technology Platform Engineer

Accenture · Bengaluru, BDC11A, India

Today

Technology Platform Engineer

Accenture · Bengaluru, BDC11A, India

Today

Technology Platform Engineer

Accenture · Bengaluru, BDC11A, India

Today

DevOps Engineer

Accenture · Navi Mumbai, MDC5C, India

Today

DevOps Engineer

Accenture · Bengaluru, BDC7B, India

Today

Data Platform Engineer

Accenture · Hyderabad, HDC4A, India

Today

Data Platform Engineer

Accenture · Pune, PDC3B, India

Today
Site Reliability Engineer (Onsite, Lahore, PKR Salary) at Hr Pod Hiring Talent Globally | Hiring.Camp