- Salary
- $215k – $285k
- Location
- San Francisco, California
- Type
- Full-time
- Department
- Engineering
- Seniority
- Senior
- Experience
- 5+ years
- Source
- RecruiterFlow
Description
Senior Software Engineer, GPU Sandboxes
Location: San Francisco, CA
Company Stage of Funding: Series A
Office Type: In Person
Salary: $215,000-$285,000 + Equity
Company Description
We're representing a rapidly growing infrastructure company building the compute layer for AI agents. Its platform provides secure, isolated, instantly available sandboxes where AI agents can execute code, interact with computers, and run training workloads.
Following its Series A, the company has already grown 10x and is now expanding its GPU infrastructure to give AI agents on-demand access to dedicated GPU compute with strong isolation and serverless economics. The engineering challenges sit deep in the systems stack, spanning hypervisors, Linux virtualization, GPU drivers, kernels, scheduling, and the orchestration infrastructure required to make dedicated GPUs available behind an API call within seconds.
This is a deeply hands-on systems engineering role for someone who wants to solve difficult infrastructure problems rather than work within a large engineering organization.
What You Will Do
- Build and productionize the runtime powering GPU-backed sandboxes for AI agents and machine learning workloads.
- Develop virtualization infrastructure using KVM/QEMU, VFIO passthrough, and IOMMU isolation.
- Build systems for managing IOMMU groups, PCIe devices, guest drivers, and CUDA stack lifecycles.
- Optimize GPU sandbox cold-start latency across image caching, VM boot paths, driver initialization, and snapshot restoration.
- Build GPU-aware scheduling and bin-packing systems that maximize fleet utilization and improve serverless GPU economics.
- Develop fleet-wide GPU health monitoring, failure detection, and automated remediation infrastructure.
- Implement pause, snapshot, restore, and fork capabilities for GPU-attached workloads.
- Design and harden security-sensitive infrastructure around device isolation, tenant boundaries, and shared-host environments.
- Investigate and resolve complex production issues involving PCIe errors, IOMMU faults, GPU driver crashes, thermal throttling, and hardware failures.
- Build tooling and automation for managing production GPU fleets.
- Own systems from initial architecture and implementation through deployment, observability, on-call, and ongoing reliability.
- Work closely with the GPU Sandboxes Engineering Lead and other senior engineers on the platform's long-term technical architecture.
Ideal Background
- 5+ years of professional systems, infrastructure, virtualization, or platform engineering experience.
- Deep hands-on experience with Linux virtualization internals, including KVM, QEMU, libvirt, VFIO, and IOMMU.
- Strong systems programming experience with Go, Rust, C, or C++.
- Experience debugging Linux systems at or near the kernel, driver, virtualization, or hardware boundary.
- Production experience operating GPU infrastructure, ideally involving NVIDIA data center hardware.
- Strong understanding of virtualization, hardware isolation, PCIe devices, and production Linux environments.
- Experience building or operating infrastructure where performance, reliability, and resource utilization are critical.
- Strong debugging ability across hardware and software boundaries.
- Comfortable owning production services, including observability, incident response, and on-call.
- High ownership mindset and interest in solving difficult low-level technical problems in a small, fast-moving team.
Preferred
- Contributions to virtualization technologies such as QEMU, KVM, Firecracker, Cloud Hypervisor, or similar open-source systems projects.
- Experience with NVIDIA data center GPUs such as H100s or A100s.
- Experience using NVML or DCGM for GPU telemetry, health monitoring, and fleet management.
- GPU driver-level engineering or debugging experience.
- Experience implementing GPU checkpointing, snapshotting, or workload restoration.
- Background building serverless compute, GPU clouds, virtualization platforms, or hyperscaler infrastructure.
- Experience with RDMA or other high-performance networking technologies.
- Familiarity with GPUDirect.
- Experience with NVLink, multi-GPU interconnects, or other multi-GPU infrastructure.
- Experience optimizing VM cold starts, image caching, scheduling, or infrastructure density.
- Familiarity with infrastructure supporting AI agents, reinforcement learning, model training, or inference workloads.
Compensation and Benefits
- Competitive base salary plus equity.
- Full-time position reporting to the Engineering Lead for GPU Sandboxes.
- Location options in San Francisco or Croatia.
- Opportunity to build deeply technical GPU infrastructure spanning hypervisors, kernels, drivers, hardware, and distributed orchestration.
- Join shortly after the company's Series A during a period of rapid growth, with the platform already having grown approximately 10x.
- Work within a small, senior engineering organization where individual engineers have significant architectural and production ownership.
- Opportunity to help solve emerging infrastructure problems around making dedicated GPU compute instantaneous, isolated, and economically viable for AI agents.