Hiring.Camp

Tech Lead, AI Compute Infrastructure

Heygen

·

Oct 16, 2025

Location
Los Angeles, Palo Alto, San Francisco, Toronto, Singapore · Los Angeles, California, United States · Palo Alto, California, United States · San Francisco, California, United States · Toronto, Ontario, Canada
Workplace
Remote
Department
Engineering
Seniority
Lead
Education
PhD
Source
Greenhouse

Description

About HeyGen

At HeyGen, our mission is to make visual storytelling accessible to all. Over the last decade, visual content has become the preferred method of information creation, consumption, and retention. But the ability to create such content, in particular videos, continues to be costly and challenging to scale. Our ambition is to build technology that equips more people with the power to reach, captivate, and inspire audiences.
Learn more at www.heygen.com.  Visit our Mission and Culture doc here

We are seeking a seasoned Technical Leader to build and scale the foundational compute infrastructure that powers our state-of-the-art AI models—from multimodal training data pipelines to high-throughput, low-latency video generation.

Responsibilities

You will be the core engineer responsible for building the robust, efficient, and scalable platform that enables our research and production teams to rapidly iterate on HeyGen's generative video models. Your contributions will directly impact model performance, developer productivity, and the final quality of every AI-generated video.

  • Optimize GPU Utilization: Design and implement mechanisms to aggressively optimize GPU and cluster utilization across thousands of devices for inference, training, data processing and large-scale deployment of our state-of-art video generation models.

  • Develop Large-Scale AI Job Framework: Build highly scalable, reliable frameworks for launching and managing massive, heterogeneous compute jobs, including multi-modal high-volume data ingestion/processing, distributed model training, and continuous evaluation/benchmarking.

  • Enhance Observability: Develop world-class observability, tracing, and visualization tools for our compute cluster to ensure reliability, diagnose performance bottlenecks (e.g., memory, bandwidth, communication).

  • Accelerate Pipelines: Collaborate closely with AI researchers and AI engineers to integrate innovative acceleration techniques (e.g., custom CUDA kernels, distributed training libraries) into production-ready, scalable training and inference pipelines.

  • Infrastructure Management: Champion the adoption and optimization of modern cloud and container technologies (Kubernetes, Ray) for elastic, cost-efficient scaling of our distributed systems.

Minimum Requirements

We are looking for a highly motivated engineer with deep experience operating and optimizing AI infrastructure at scale.

  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.

  • 5+ years of full-time industry experience in large-scale MLOps, AI infrastructure, or HPC systems.

  • Experience with data frameworks and standards like Ray, Apache Spark, LanceDB

  • Strong proficiency in Python and a high-performance language such as C++  for developing core infrastructure components.

  • Deep understanding and hands-on experience with modern orchestration and distributed computing frameworks such as Kubernetes and Ray.

  • Experience with core ML frameworks such as PyTorch, TensorFlow, or JAX.

Preferred Qualifications

  • Master's or PhD in Computer Science or a related technical field.

  • Demonstrated Tech Lead experience, driving projects from conceptual design through to production deployment across cross-functional teams.

  • Prior experience building infrastructure specifically for Generative AI models (e.g., diffusion models, GANs, or large language models) where cost and latency are critical.

  • Proven background in building and operating large-scale data infrastructure (e.g., Ray, Apache Spark) to manage petabytes of multi-modal data (video, audio, text).

  • Expertise in GPU acceleration and deep familiarity with low-level compute programming, including CUDA, NCCL, or similar technologies for efficient inter-GPU communication.

What HeyGen Offers

  • Competitive salary and benefits package.
  • Dynamic and inclusive work environment.
  • Opportunities for professional growth and advancement.
  • Collaborative culture that values innovation and creativity.
  • Access to the latest technologies and tools.

 

HeyGen is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Skills

PythonKubernetesTensorFlowPyTorchSpark

Similar Jobs

30

AI Tech Lead

Inorsa · Austin, TX +1

2 months ago

AI Tech Lead

FeverUp · Madrid

3 months ago

Tech Lead AI

Qualifinds® · Onsite

1+ year ago

AI Tech Lead - Commercial

Smithnephew · US - Field, United States of America

Yesterday

Full Stack AI Tech Lead - Vice President

citibank · Mississauga, ON,CA, CA

5 days ago

Full Stack AI Tech Lead - Vice President

Citi Bank · 5900 HURONTARIO STREET MISSISSAUGA, Canada · Hybrid

5 days ago

Tech Lead AI Product Engineering

Cobre · LATAM +1

6 days ago

AI Tech Lead - Pipeline

Adastragrp · Calgary, AB, CA · Hybrid, Onsite

2 weeks ago

Frontend Tech lead ( AI Agent)

Binance · Asia · Hybrid

2 weeks ago

Tech Lead AI F/H

Talan · Lyon, Auvergne-Rhône-Alpes, France · Hybrid

2 weeks ago

Agentic AI Tech Lead

citibank · Pune, MH,IN, IN

1 month ago

Agentic AI Tech Lead

Citi Bank · PLOT NO-1, S.NO. 77, India · Hybrid

1 month ago

Tech Lead, AI Engineering

BMO is · BMOPLACE, Canada · Remote, Hybrid

1 month ago

Tech Lead AI Conversational Agents

Santander Effect Our work touches · JUAN IGNACIO LUCA DE TENA-PLANTA BAJA, Spain

1 month ago

Tech Lead AI Conversational Agents

Santander · JUAN IGNACIO LUCA DE TENA-PLANTA BAJA, Spain

1 month ago

Applied AI Tech Lead

Amigo · New York City · Onsite

1 month ago

Agentic AI Tech Lead

citibank · Mississauga, ON,CA, CA

1 month ago

Tech Lead AI/ML (F/H/X)

Referral Publicisgroupe · Paris, FR

1 month ago

Staff Software Engineer (Tech Lead) - AI Data Platform (d/f/m, Berlin)

Mondaai · Germany · Hybrid

4 months ago

Staff Software Engineer (Tech Lead) - AI Data Platform (d/f/m, Berlin)

Mondaai · Germany · Hybrid

4 months ago

AI Tech Lead/Architect - PALO IT Labs

Paloit · Mexico +1

5 months ago

AGENTIC AI TECH LEAD

Mango · Palau-solità i Plegamans, Catalonia, Spain

5 months ago

AI Tech Lead Manager

Umusic · GBR-4PS, United Kingdom

6 months ago

AI Tech Lead (Founding Engineer)

ambi.careers · Remote, Hybrid

6 months ago

Staff Machine Learning Engineer - AI Tech Lead

Sumologic · United States

6 months ago

Rezolve.AI- Tech Lead

Nexthire · Chennai, IN · Onsite

10 months ago

III Tech Lead AI

Qualifinds® · Onsite

1+ year ago

Tech Lead AI/ML Engineer

Aiserajobs · Athens, Greece +1

1+ year ago

AI Agent Tech Lead

Lavora con Noi - Gruppo Credem · REGGIO EMILIA, EMILIA ROMAGNA, Italy · Hybrid

1 week ago

Tech Lead, Applied AI

Abby Care · San Francisco · Hybrid

2 weeks ago