- Location
- Markham, ON,CA, CA
- Type
- Full-time
- Department
- Engineering
- Seniority
- Senior
- Experience
- 1+ years
- Education
- PhD
- Source
- Eightfold
Description
Explore and prototype novel or emerging ML model architectures optimized for on-device, low-power inference, including vision, audio, and multimodal workloads. Design, evaluate, and refine quantization, mixed-precision, sparsity, and compression techniques, with careful analysis of accuracy-performance-power trade-offs. Develop and optimize computational graphs, including operator fusion, scheduling strategies, and memory-aware execution. Conduct rigorous performance and accuracy investigations using profiling tools, hardware counters, and targeted experiments. Collaborate closely with compiler, runtime, and hardware teams to convert exploratory prototypes into production-viable execution paths. Influence future accelerator features, compiler capabilities, and deployment strategies through technical insights and experimental results. Strong track record in machine learning research or advanced applied ML development, with demonstrated focus on inference efficiency. Deep understanding of ML model architecture, operator behavior, and inference-time performance characteristics. Hands-on experience with quantization and reduced-precision inference (e.g., INT8/INT4, FP8/FP4, mixed precision, PTQ/QAT). Proven ability to prototype, analyze, and iterate on ideas under strict compute, memory, and power constraints. Proficiency in Python and C/C++, with comfort working across modeling, systems, and low-level execution layers. Strong background in computer architecture and hardware-aware optimization, particularly for AI accelerators. Ability to reason about computational graphs, tensor layouts, and memory movement at a detailed level. Master's degree in Computer Science, Engineering, Information Systems, or a related field, and 1+ year of experience in Hardware Engineering, Software Engineering, Systems Engineering, or a related area; OR PhD in Computer Science, Engineering, Information Systems, or a related field Experience targeting or co-designing for custom accelerators, NPUs, DSPs, or GPUs. Familiarity with compiler-assisted ML optimization, graph transformations, or operator scheduling. Experience with multimodal or sensor-driven models. Evidence of technical leadership, such as driving complex investigations, publishing, patenting, or shaping internal technical direction. Comfort operating in ambiguous, research-heavy problem spaces with minimal upfront specification. Bachelor's degree in Computer Science, Engineering, Information Systems, or related field and 2+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience. OR Master's degree in Computer Science, Engineering, Information Systems, or related field and 1+ year of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience. OR PhD in Computer Science, Engineering, Information Systems, or related field.