- Location
- Houston, TX or San Francisco Bay Area · Houston, Texas, United States
- Department
- Algorithm
- Experience
- 3+ years
- Source
- Greenhouse
Description
Company Introduction
At Bot Auto, we are revolutionizing the transportation of goods with our cutting-edge autonomous trucks, enhancing the quality of life for communities around the globe. With the agility of a start-up and the wisdom of seasoned experts, Bot Auto boasts a team that has achieved numerous world-firsts and unparalleled innovations. United by a shared vision, we create miracles and propel the future of transportation. Join us and transform your dreams into reality.
You would collaborate with software engineers, AI researchers, and hardware specialists to develop high-performance solutions that meet the stringent requirements of autonomous driving applications. This is an exciting opportunity to work on next-generation transportation technology and make a meaningful impact on the future of mobility.
Key Responsibilities
- Optimize end-to-end GPU performance for real-time autonomous driving workloads, including sensor processing (e.g., camera, LiDAR) and neural network inference.
- Develop and optimize parallel computing algorithms and GPU-accelerated components using technologies such as CUDA.
- Collaborate with cross-functional teams to design and improve onboard GPU software architectures that meet the computational requirements of perception, planning, and control modules.
- Profile and analyze bottlenecks across GPU computation, memory access, data movement, synchronization, and CPU–GPU interaction.
- Debug and optimize GPU-based software to improve latency, throughput, resource utilization, and runtime stability on embedded platforms.
Qualifications:
Required:
- Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field.
- Strong knowledge of parallel computing principles, GPU architecture, memory hierarchy, and performance optimization techniques.
- Experience profiling GPU applications using tools such as NVIDIA Nsight Systems, Nsight Compute, or equivalent tools.
- Experience deploying or optimizing neural network inference workloads using technologies such as PyTorch, ONNX, and TensorRT.
- Experience with real-time embedded systems and handling large data streams from sensors (camera, LiDAR, radar).
- Strong proficiency in C/C++ and Python.
Preferred:
- 3+ years of experience in GPU programming and optimization (e.g., CUDA, OpenCL, Vulkan).
- Experience with NVIDIA Jetson Thor, NVIDIA DRIVE Thor, or similar embedded GPU platforms.
- Experience with model quantization, including FP8 and NVFP4.
- Experience managing concurrent GPU workloads and resource isolation using technologies such as NVIDIA Multi-Process Service (MPS), Multi-Instance GPU (MIG), or other related technologies.
- Experience with GPU-accelerated sensor data compression, including camera, LiDAR, or other onboard sensor data.