- Salary
- $200k – $250k
- Workplace
- Remote, Onsite
- Type
- Full-time
- Department
- Engineering
- Seniority
- Lead
- Experience
- 7+ years
- Education
- Bachelor
- Source
- RecruiterFlow
Description
Principal Infrastructure Engineer / Tech Lead
Location
San Francisco, CA
On-site role requiring 5 days/week at the San Francisco headquarters. Relocation assistance available.
Compensation
$200,000 – $250,000 Base + Competitive Equity
Compensation may be flexible above $250K for exceptional candidates.
Visa
Open to H-1B Transfers / OPT Transfers
No new visa sponsorship.
Company Stage
Series A – Growth-Stage AI / Enterprise Technology Company
Industry
Artificial Intelligence, Machine Learning, Infrastructure, Supply Chain Technology, Enterprise Software, Cloud Infrastructure, MLOps
About the Company
Our client is building AI-powered forecasting technology designed to help large enterprises improve supply chain efficiency and reduce physical waste.
The company develops custom AI models for enterprise customers, requiring sophisticated infrastructure for model training, inference, deployment, and production operations.
The company has raised approximately $28M and is backed by leading venture capital firms. Despite its growth stage, the organization remains highly lean, with a small engineering team and significant technical ownership available to early infrastructure leaders.
As a Principal Infrastructure Engineer / Tech Lead, you'll be one of the foundational infrastructure engineers responsible for designing and scaling the cloud systems that power custom AI model training and inference.
This is an opportunity for a hands-on technical leader who wants to own infrastructure architecture end-to-end while remaining deeply involved in engineering execution, cloud systems, ML infrastructure, reliability, security, and cost optimization.
What You'll Do
- Own and scale the cloud infrastructure powering custom AI model training and inference for enterprise customers
- Serve as a technical lead for infrastructure and platform engineering initiatives
- Design reliable and scalable cloud architecture across production environments
- Build and maintain deployment systems for AI models and production services
- Design and improve CI/CD pipelines for reliable software and model deployments
- Build infrastructure-as-code systems using tools such as Terraform
- Manage GPU allocation, instance sizing, and cloud resource utilization
- Optimize infrastructure and GPU costs across production workloads
- Build monitoring, logging, alerting, and observability infrastructure
- Develop automated testing infrastructure for production systems
- Improve visibility into system health, failures, performance, and resource utilization
- Own critical infrastructure projects from architecture through production
- Build and maintain data and ML operations infrastructure
- Develop ETL pipelines supporting varied customer data sources
- Implement model versioning and model lifecycle management systems
- Support infrastructure powering ML research and experimentation
- Drive security engineering initiatives across the infrastructure stack
- Design systems that maintain strong data isolation between customers
- Improve infrastructure hardening and security practices
- Support SOC2 compliance and security requirements
- Partner closely with ML researchers and engineers to improve infrastructure reliability and scalability
- Balance immediate infrastructure improvements with longer-term platform architecture
- Participate in on-call responsibilities supporting customers during US business hours
- Mentor and partner with existing infrastructure engineers
- Establish infrastructure engineering standards and best practices
- Make architecture and technology decisions in a highly autonomous environment
- Work closely with engineering leadership on technical strategy
- Build infrastructure capable of supporting rapid company and customer growth
- Maintain hands-on involvement in technical implementation and production systems
- Help shape the company's long-term infrastructure and platform strategy
Ideal Candidate Background
Experience Requirements
- 7+ years of experience in cloud infrastructure, DevOps, platform engineering, or related infrastructure roles
- Experience serving as a technical lead for an infrastructure or platform engineering team
- Strong hands-on individual contributor experience in a current or recent role
- Experience owning infrastructure projects from architecture through production
- Experience working at startups or early-stage technology companies
- Experience building and operating production cloud infrastructure
- Experience working with large-scale or highly reliable production systems
- Experience with infrastructure architecture and technical leadership
- Experience collaborating closely with engineering and technical leadership
- Experience supporting ML workflows or research infrastructure preferred
- Experience working in environments requiring high technical ownership
- Strong ability to operate independently in ambiguous startup environments
- Strong technical communication and problem-solving skills
- Experience mentoring other infrastructure or software engineers
Technical Requirements
- Strong cloud infrastructure experience
- Strong AWS experience preferred
- GCP or Azure infrastructure experience considered transferable
- Strong Python coding fluency
- Experience with Infrastructure-as-Code tools such as Terraform
- Strong CI/CD experience
- Experience designing and maintaining production deployment systems
- Experience with cloud networking and infrastructure architecture
- Experience with production monitoring, logging, and observability
- Experience with infrastructure reliability and incident response
- Experience managing cloud infrastructure costs
- Experience with containerized production workloads
- Experience with scalable backend or platform infrastructure
- Strong understanding of distributed systems fundamentals
- Experience supporting high-compute workloads preferred
- Experience scaling GPU workloads preferred
- Experience with MLOps or ML infrastructure preferred
- Experience with ETL and data pipelines preferred
- Experience with model deployment and lifecycle management preferred
- Experience with infrastructure security and data isolation
- Experience with SOC2 or comparable security compliance initiatives preferred
- Strong systems design and architecture skills
- Ability to troubleshoot complex infrastructure problems
- Ability to balance reliability, scalability, performance, and cost
- Ability to remain hands-on while operating as a technical lead
Education
- Bachelor's degree in Computer Science, Engineering, or related technical field preferred
- Strong computer science fundamentals
- Equivalent practical infrastructure engineering experience accepted
Soft Skills
- Exceptional ownership and execution ability
- Strong technical leadership skills
- Highly hands-on engineering mindset
- Comfortable operating as a senior individual contributor and technical lead
- Strong architectural judgment
- Excellent problem-solving ability
- Comfortable operating with high autonomy
- Strong technical communication skills
- Ability to explain infrastructure tradeoffs clearly
- Strong mentoring and collaboration skills
- Low-ego working style
- Comfortable working closely with engineering leadership
- Strong ability to prioritize immediate fixes against long-term technical investments
- Comfortable working in fast-paced startup environments
- Strong production reliability instincts
- Strong security mindset
- Comfortable participating in on-call responsibilities
- Comfortable working with ML and research teams
- Adaptable and willing to work across infrastructure and product engineering
- Strong interest in AI infrastructure and ML systems
- Comfortable working with GPU-heavy infrastructure
- Comfortable working on-site in San Francisco 5 days/week
Compensation & Benefits
- Base Salary: $200,000 – $250,000
- Potential flexibility above $250K for exceptional candidates
- Competitive startup equity targeted around the 90th percentile relative to market for strong candidates
- Opportunity to work as one of the earliest infrastructure hires
- High ownership over foundational cloud and ML infrastructure
- Direct influence over infrastructure architecture and technical strategy
- Opportunity to build systems from the ground up
- Exposure to GPU-heavy AI infrastructure and custom ML model deployment
- Opportunity to solve complex infrastructure challenges at enterprise scale
- Direct collaboration with AI researchers and technical leadership
- Opportunity to mentor and shape the infrastructure engineering function
- High autonomy in a small, technically strong engineering organization
- Relocation assistance to the San Francisco Bay Area
- H-1B and OPT transfer support
Why Join
This is an opportunity to become a foundational infrastructure leader at a highly capital-efficient AI company solving complex problems at the intersection of machine learning, cloud infrastructure, and enterprise technology.
You'll own the infrastructure powering custom AI models for major enterprise customers rather than maintaining mature systems at a large organization.
You'll work closely with a small team of highly technical engineers and researchers while having significant influence over architecture, cloud infrastructure, GPU workloads, security, reliability, and ML operations.
If you enjoy building infrastructure from the ground up, solving difficult cloud and distributed systems problems, working with AI/ML workloads, and operating with exceptional technical ownership, this role offers significant scope and impact.