- Location
- Lahore
- Type
- Full-time
- Department
- Engineering
- Experience
- 5+ years
- Education
- Bachelor
- Closing date
- Today
- Source
- CareersPage
Description
Requirements:
- 5+ years of experience in Linux systems administration or infrastructure engineering at scale.
- Strong knowledge of Linux internals with hands-on troubleshooting skills across the OS, kernel, storage, and system performance.
- Solid understanding of TCP/IP networking fundamentals with hands-on network troubleshooting experience.
- Hands-on experience operating NFS and other shared storage systems in production environments.
- Proven experience with configuration management using Ansible.
- Experience with monitoring tools such as Grafana and ticketing tools such as ServiceNow and JIRA.
- Proficiency in Bash and Python scripting for automation.
- Bachelor's degree in Computer Science, Engineering, or an equivalent field/experience.
- Experience operating GPU servers, HPC, or AI infrastructure.
- Familiarity with Kubernetes, Slurm, or other cluster schedulers.
- Exposure to storage technologies such as Lustre, GPFS/Spectrum Scale, or Ceph.
- Knowledge of InfiniBand or RDMA and high-performance networking.
- Familiarity with or working knowledge of AI coding tools such as Claude, OpenAI, or other similar tools to accelerate automation and troubleshooting.
- Relevant certifications such as RHCSA/RHCE, CCNA, or cloud provider certifications.
Responsibilities:
- Administer, patch, and harden a large fleet of Linux servers, including RHEL, Rocky, and Ubuntu, across bare-metal and cloud environments.
- Troubleshoot complex OS, kernel, performance, and hardware issues to identify and resolve root causes.
- Design, configure, and manage TCP/IP networking, including routing, VLANs, DNS, DHCP, firewalls, and network bonding.
- Deploy and manage NFS storage and other shared file systems for high-throughput workloads.
- Automate provisioning, configuration, and remediation using Ansible and infrastructure-as-code practices.
- Build and maintain monitoring, alerting, and dashboards using Grafana and related observability tools.
- Manage operational workflows through ticketing systems such as ServiceNow (SNOW) and JIRA.
- Own capacity planning, OS lifecycle management, standardized system builds, and image management.
- Collaborate with Platform Engineering and SRE teams on reliability, security, and the rollout of new services.
- Document standards, runbooks, and procedures while continuously reducing operational toil.
Working Hours:
8 PM - 4 AM