Hiring.Camp

Software Engineer, LiteRT / On-Device AI

Qualcomm

·

Yesterday

Location
Taipei City,TW, TW · Hsinchu, Hsinchu City,TW, TW
Type
Full-time
Department
Engineering
Education
PhD
Source
Eightfold

Description

##

Company:

Qualcomm Semiconductor Limited

## Job Area:

Engineering Group, Engineering Group > Machine Learning Engineering

General Summary:

About the Team and What You Will Build

Our team maintains the software that runs AI and LLM models on Qualcomm's NPU, working hand in hand with Google and other industry partners. You will contribute to the development of:

  • LiteRT — Google's on-device ML inference runtime (the successor to TensorFlow Lite), and its Qualcomm backend that makes models run efficiently on Qualcomm's NPU.
  • LiteRT-LM — the LLM-focused layer on top of LiteRT for running large language and generative models on-device, including the model bring-up and optimization work that makes them fast enough for real products.
  • QNN SDK — Qualcomm's AI Engine software stack that helps you build model files to run on various device processors across multiple operating systems.
  • QNN TFLite Delegate — the delegate extends the QNN SDK to support running TFLite models on the Qualcomm NPU.

You will work across this stack, from the model graph down to the runtime and the hardware, with a particular focus on bringing up and optimizing modern LLMs.

Key Responsibilities

You will develop in the core runtime and backend libraries that execute AI and LLM workloads on the NPU — an AI model inferencing stack scalable across Qualcomm platforms — reasoning about how a model maps onto the hardware end to end, from the model graph down to the runtime, memory, and scheduling behavior on the device.

  • Enable, bring up, and optimize AI and LLM models running on the Qualcomm NPU, from graph-level transformations down to individual operations.
  • Design, implement, profile, and analyze the runtime, backend, and operation libraries that execute these models.
  • Diagnose accuracy and performance issues across the stack (model, runtime, backend, and hardware) and drive them to root cause.
  • Deliver high-quality code and collaborate with open-source software communities.
  • Work with key technical specialists across Qualcomm, our partners, and our customers to improve the libraries for commercial use cases and industrial benchmarks.
  • Build and maintain the tooling and test coverage that keep the runtime stable and its numerics correct across platforms.

Qualifications (New College Graduate / Junior)

  • Master's degree or above in Computer Science, Electrical Engineering, or a related field.
  • Proficiency in programming languages such as C, C++, or Python.
  • Strong knowledge of object-oriented programming, data structures, algorithms, operating systems, and computer architecture.
  • Understanding of machine learning and deep learning fundamentals, including the structure of modern neural networks and the transformer / LLM architecture (attention, KV cache, tokenization, prefill vs. decode).
  • Self-motivated and capable of working independently with minimal oversight.
  • Ability to communicate technical concepts effectively and to work collaboratively within cross-functional teams.

Minimum Qualifications (Experienced / Senior)

  • Master's degree or above in Computer Science, Electrical Engineering, or a related field, and relevant software engineering experience.
  • Strong proficiency in modern C++ and Python for building and optimizing performance-critical software.
  • Solid grounding in computer architecture, and an understanding of NPU / DSP / GPU architectures and parallel programming concepts.
  • Solid operating-system fundamentals, including how the OS manages and schedules compute and memory resources, and how that affects on-device inference (threading, memory allocation, buffer / cache management, contention, and power / performance trade-offs).
  • Familiarity with embedded systems and the constraints of edge devices.
  • Experience with low-level programming for efficient hardware utilization, and with analyzing and optimizing performance bottlenecks.
  • Proficiency with version control and development tools such as Git, Gerrit, and Jira.

Preferred Qualifications (Experienced / Senior)

  • Master's or Ph.D. in Computer Science, Electrical Engineering, or a related field.
  • Solid understanding of LLM and generative-model architecture, and hands-on experience bringing such models up or optimizing them for inference.
  • Familiarity with quantization concepts and their effect on model accuracy and performance (for example, integer / mixed-precision quantization and its tradeoffs).
  • Familiarity with ML frameworks and runtimes such as TensorFlow, PyTorch, and LiteRT, and an understanding of mainstream ML model formats and their runtime environments.
  • Experience optimizing code specifically for NPU / DSP / GPU architectures, and in fine-tuning performance-critical applications on edge devices.
  • Understanding of the interaction between the model graph, the runtime, and OS-level resource scheduling, and how to reason about the whole pipeline when debugging accuracy or latency.
  • Understanding of hardware-software co-design principles for edge devices.
  • Proficiency with tools for debugging and profiling on-device code.
  • Strong communication skills to collaborate with model, framework, and hardware engineers and to convey complex technical concepts clearly.

Expertise in at least one of the following areas

  • LLM / generative-model enablement and optimization: bringing modern language or multimodal models onto an accelerator and improving their prefill / decode performance and accuracy.
  • Runtime and backend development for an ML inference stack: designing the layers that translate a model graph into hardware-executable operations.
  • Performance profiling and optimization for parallel or heterogeneous compute: strong practical experience finding and removing bottlenecks.
  • Operating-system-level resource management and scheduling as it applies to on-device inference: memory, threading, and compute-resource contention on constrained hardware.
  • Compiler technology: familiarity with ML compilers such as TVM / XLA / Glow, or experience with LLVM / GCC backend development, optimization analysis, and implementation, is highly advantageous.
  • Deep learning fundamentals: knowledge of neural-network and transformer fundamentals, experience training or fine-tuning models, and familiarity with TensorFlow / PyTorch.

Other Qualifications

  • Enthusiasm for machine learning and on-device AI, especially the LLM and generative-model space.
  • Hands-on experience with the design or implementation of deep-learning networks via modern frameworks such as TensorFlow and PyTorch.
  • Ability to quickly learn new technologies and to resolve customer-reported technical problems during product development cycles.
  • Excellent analytical, problem-solving, and communication skills, and a willingness to work directly with partners and customers.

Minimum Qualifications:

  • Bachelor's degree in Computer Science, Engineering, Information Systems, or related field and 2+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.

OR

Master's degree in Computer Science, Engineering, Information Systems, or related field and 1+ year of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.

OR

PhD in Computer Science, Engineering, Information Systems, or related field.

Applicants: Qualcomm is an equal opportunity employer. If you are an individual with a disability and need an accommodation during the application/hiring process, rest assured that Qualcomm is committed to providing an accessible process. You may e-mail [email protected] or call Qualcomm's toll-free number found here. Upon request, Qualcomm will provide reasonable accommodations to support individuals with disabilities to be able participate in the hiring process. Qualcomm is also committed to making our workplace accessible for individuals with disabilities. (Keep in mind that this email address is used to provide reasonable accommodations for individuals with disabilities. We will not respond here to requests for updates on applications or resume inquiries).

Qualcomm expects its employees to abide by all applicable policies and procedures, including but not limited to security and other requirements regarding protection of Company confidential information and other confidential and/or proprietary information, to the extent those requirements are permissible under applicable law.

To all Staffing and Recruiting Agencies: Our Careers Site is only for individuals seeking a job at Qualcomm. Staffing and recruiting agencies and individuals being represented by an agency are not authorized to use this site or to submit profiles, applications or resumes, and any such submissions will be considered unsolicited. Qualcomm does not accept unsolicited resumes or applications from agencies. Please do not forward resumes to our jobs alias, Qualcomm employees or any other company location. Qualcomm is not responsible for any fees related to unsolicited resumes/applications.

If you would like more information about this role, please contact Qualcomm Careers.

Skills

PythonMachine LearningDeep LearningTensorFlowPyTorchGitJira