We are seeking PhD research interns with strong expertise in generative modeling and a demonstrated record of original research. In this role, you will work alongside our research team to develop world models that learn the dynamics of the physical world from large-scale multimodal data — predicting how a scene evolves under an agent's actions, and serving as a learned simulator for training and evaluating driving and robotic policies. You will work with state-of-the-art generative architectures, including diffusion and flow-matching models, video tokenizers, and transformer-based multimodal backbones, with access to vast amounts of real-world multimodal data from our autonomous fleet and robotics platforms. Interns are expected to drive a focused research project end to end, and strong results are supported for publication at top-tier venues.
-
Drive a focused research project on predictive world models, spanning problem formulation, architecture design, training, evaluation, and empirical analysis, in close collaboration with a mentor and the broader research team.
-
Contribute to one or more of the following directions: high-quality multi-view future prediction and generation, supporting both action-conditioned rollouts and formulations that forecast the future without explicit action conditioning; architectures in which a shared backbone both predicts the future and produces trajectories or actions; predictive pre-training to improve Vision-Language-Action (VLA) driving performance.
-
Extend prediction beyond 2D pixels into a shared multimodal latent space that spans 3D scene representations such as Gaussian Splatting, together with occupancy and reward signals, so that a single model can support simulation, evaluation, and policy training.
-
Investigate cross-embodiment generalization through unified observation and action representations and embodiment-conditioning mechanisms, so that a single world model transfers across vehicles, robots, and sensor configurations with only few-shot data.
-
Build evaluation methodology for predictive world models, spanning representation quality, prediction accuracy, generation fidelity, physical plausibility, long-horizon rollout consistency, and closed-loop policy performance.
-
Collaborate with research engineers to move research prototypes into scalable training and inference pipelines, and publish and open-source results where appropriate.
-
Currently pursuing a PhD in Engineering, Computer Science, or a related field, with a focus on Deep Learning, Computer Vision, or Generative Models.
-
First-author publications at top-tier venues such as CVPR, ICCV, ECCV, NeurIPS, ICLR, ICML, CoRL, RSS, or SIGGRAPH. Work under submission may be presented as an arXiv preprint.
-
Strong, up-to-date foundation in generative modeling and experimental methodology, with hands-on experience building, training, fine-tuning, and evaluating models in PyTorch or JAX.
-
Strong Python programming and software design skills, with a solid understanding of data structures, algorithms, code optimization, and large-scale data processing.
-
Available to commit to a minimum of 12 weeks and work on-site at our Santa Clara office.