- Location
- México D.F., CDMX,MX, MX
- Type
- Full-time
- Department
- Engineering
- Seniority
- Senior
- Education
- PhD
- Source
- Eightfold
Description
##
Company:
Qualcomm Intl Inc., Mexico Branch Office
## Job Area:
Information Technology Group, Information Technology Group > IT Engineering
General Summary:
Position Summary
Qualcomm's Grid Solution team is seeking a Senior HPC Scheduler Operations Engineer to operate, scale, and continuously optimize large-scale job scheduling infrastructure supporting Electronic Design Automation (EDA) and compute-intensive engineering workloads. This role is embedded within the Engineering IT Hardware Infrastructure — EDA Compute organization and requires deep expertise in IBM Spectrum LSF as the primary scheduler, with a working knowledge of Slurm as an emerging priority. The ideal candidate brings strong Linux systems administration skills, a reliability engineering mindset, and the ability to translate complex scheduler behavior into actionable insights for engineering and management stakeholders.
Key Responsibilities
Scheduler Operations & Optimization
- Manage, scale, and optimize IBM Spectrum LSF job scheduling systems for EDA and HPC compute-intensive workloads across multi-site environments.
- Administer scheduler configuration including queue policies, resource limits, fairshare, job arrays, and preemption strategies.
- Analyze scheduler and infrastructure performance data to identify bottlenecks and improve utilization, throughput, and job turnaround time.
- Perform capacity planning and workload characterization to support SKU-level packing models and tiered job scheduling strategies.
- Implement and maintain scheduler tuning parameters to minimize dispatch latency and maximize grid efficiency under high-saturation conditions.
Reliability & Incident Management
- Troubleshoot and resolve service-impacting issues across scheduler, OS, and workload layers with minimal time-to-resolution.
- Define and track SLOs for service performance and reliability; partner with customer teams to set and communicate realistic expectations.
- Implement automation and process improvements to reduce manual toil and prevent recurring incidents.
- Develop and enforce operational standards, runbooks, and best practices ensuring consistency across all sites.
Observability, Metrics & Automation
- Build and maintain observability systems including metrics pipelines, dashboards, and alerting frameworks for scheduler and compute infrastructure health.
- Leverage accounting and telemetry data (e.g., LSF stream logs, bjobs, bacct) to quantify workload coverage, grid efficiency, and tier performance.
- Develop automation tooling (Python, Shell, APIs) to streamline operations, enforce policy guardrails, and surface actionable insights.
- Produce operational reports and management summaries that combine automated metrics with human operational narratives.
Stakeholder Collaboration
- Collaborate directly with CAD and engineering teams to clarify workload requirements, translate technical tradeoffs, and drive issues to closure.
- Communicate scheduler performance metrics, reliability posture, and capacity status clearly to engineering and management audiences.
- Partner with peer infrastructure teams to coordinate cross-functional changes and align on shared operational standards.
Required Qualifications
Education
Bachelor's degree in Computer Science, Electrical Engineering, or related field, or equivalent practical experience.
Experience
5+ years operating and supporting large-scale Linux-based compute infrastructure in HPC or silicon design environments.
Scheduler Expertise
Strong hands-on experience with IBM Spectrum LSF, including queue configuration, policy tuning, fairshare, job arrays, and multi-cluster operations.
Linux Systems
Proficiency in SLES administration including OS-level troubleshooting, workload profiling and user environment management.
Problem Solving
Ability to independently analyze complex system behavior under load; experience with root cause analysis across scheduler, OS, and hardware layers.
Scripting & Automation
Proficiency in Python and/or Shell scripting for operational automation, data processing, and tooling development.
Communication
Demonstrated ability to clearly articulate technical tradeoffs, reliability metrics, and operational status to both engineering and management audiences.
Preferred Qualifications
- Working knowledge of Slurm (Simple Linux Utility for Resource Management); familiarity with Slurm configuration, partition management, and job scheduling concepts is a strong plus as the team evaluates multi-scheduler environments.
- Deep knowledge of scheduler internals, configuration tuning, and advanced troubleshooting for LSF and/or Slurm.
- Familiarity with container technologies (Docker, Singularity/Apptainer, Podman) in HPC or EDA compute environments; experience deploying or managing containerized workloads on HPC schedulers is advantageous as the team evolves toward container-native compute delivery.
- Background influencing adoption of new infrastructure standards and operational practices across distributed engineering teams and multi-site organizations.
- Exposure to hardware-level telemetry and performance monitoring (e.g., PMU counters, IPC, memory bandwidth) for workload analysis.
Minimum Qualifications:
- 3+ years of IT-related work experience with a Bachelor's degree.
OR
5+ years of IT-related work experience without a Bachelor’s degree.
*Completed advanced degrees in a relevant field may be substituted for up to two years (Master’s = one year, Doctorate = two years) of work experience.
Applicants: Qualcomm is an equal opportunity employer. If you are an individual with a disability and need an accommodation during the application/hiring process, rest assured that Qualcomm is committed to providing an accessible process. You may e-mail [email protected] or call Qualcomm's toll-free number found here. Upon request, Qualcomm will provide reasonable accommodations to support individuals with disabilities to be able participate in the hiring process. Qualcomm is also committed to making our workplace accessible for individuals with disabilities. (Keep in mind that this email address is used to provide reasonable accommodations for individuals with disabilities. We will not respond here to requests for updates on applications or resume inquiries).
Qualcomm expects its employees to abide by all applicable policies and procedures, including but not limited to security and other requirements regarding protection of Company confidential information and other confidential and/or proprietary information, to the extent those requirements are permissible under applicable law.
To all Staffing and Recruiting Agencies: Our Careers Site is only for individuals seeking a job at Qualcomm. Staffing and recruiting agencies and individuals being represented by an agency are not authorized to use this site or to submit profiles, applications or resumes, and any such submissions will be considered unsolicited. Qualcomm does not accept unsolicited resumes or applications from agencies. Please do not forward resumes to our jobs alias, Qualcomm employees or any other company location. Qualcomm is not responsible for any fees related to unsolicited resumes/applications.
If you would like more information about this role, please contact Qualcomm Careers.