- Location
- Palo Alto, CA
- Workplace
- Hybrid
- Department
- Product & Technology
- Seniority
- Manager
- Experience
- 8+ years
- Source
- Lever
Description
About the Team
Enterprise AI is a platform that provides end-to-end machine learning tooling experience to support and accelerate machine learning development, including autonomous driving and other related projects. Our platform serves customers as a standardized machine learning platform within Woven by Toyota as the larger Toyota Group companies.
The Enterprise AI ML Training Infrastructure team builds and operates the infrastructure that enables engineers and researchers to train, evaluate, and iterate on machine learning models at scale. Our customers include teams working on AD/ADAS and Woven City, where large-scale machine learning workloads, simulation, and data processing are critical to developing the next generation of mobility technologies.
The team operates at the intersection of machine learning, distributed systems, cloud infrastructure, and developer platforms. Our systems include large-scale GPU compute environments, Kubernetes-based cluster infrastructure, and workflow and pipeline orchestration technologies such as Ray, Airflow, and Temporal.
We are looking for an Engineering Manager to lead a team to help shape the technical direction, engineering culture, and long-term evolution of our ML training infrastructure platform.
WHO ARE WE LOOKING FOR?
As the Engineering Manager, ML Training Infrastructure, you will lead the team responsible for building and operating the infrastructure that powers machine learning training workloads across Woven by Toyota.
You will combine engineering leadership with strong technical judgment. You will work closely with engineering, research, and product stakeholders to understand customer needs, translate them into scalable technical solutions, and ensure that our platform is reliable, efficient, secure, and easy to use.
This role is particularly suited to an engineering leader who enjoys working on complex infrastructure problems and is comfortable operating across Kubernetes, distributed systems, GPU infrastructure, cloud platforms, and machine learning workloads.
You will also play an important role in building a strong engineering culture, developing engineers, establishing effective engineering practices, and partnering with our customers to continuously improve the platform.
What You Will Enable
In this role, you will help build the infrastructure that enables teams across Woven by Toyota to train and operate increasingly large and sophisticated machine learning workloads.
Your work will directly support engineers and researchers working on AD/ADAS, Woven City, and other mobility initiatives by making compute resources easier to access, workloads more reliable, and ML development faster and more scalable.
You will have the opportunity to shape both the technical architecture of a critical ML infrastructure platform and the engineering culture of the team building it.