Hiring.Camp

Systems Performance Modeling Engineer

Tensordyne

·

Today

Location
Sunnyvale, CA · San Jose, California, United States
Department
Engineering
Source
Greenhouse

Description

About Tensordyne:

Artificial intelligence (AI) is transforming our world. It can perform cognitive functions that previously only humans could do, such as perceiving interactions across different modalities and environments - with the ability to quickly learn and then solve complex problems. Tensordyne is an AI system solution company that builds very high-performance, low-power generative AI inference systems. Our mission, through the creation of custom silicon, hardware and software, is to enable multimodal Generative AI inference acceleration at scale, with safe, sustainable, high-performance systems for our hyperscaler and neocloud data center customers. We are at the leading edge of advancing the latest research and product improvements for generative Al inference solutions that will make Al even more advantageous for compelling new generative AI applications. Tensordyne is a well funded, fast-paced startup company with headquarters in both Sunnyvale, CA, and Munich, Germany. We also have many talented team members working remotely across North America and Europe. We take care of our people and their families with comprehensive benefits, competitive compensation, flexible spending options, and recognition programs, because building category-defining technology starts with a healthy, supported team. Come join us as we shape the future of multimodal generative artificial intelligence!

About the Role

We are looking for a Systems Performance Modeling Engineer to build the models and tools that predict how generative AI inference workloads perform on Tensordyne systems, from a single accelerator up through rack, pod, and cluster scale. This is a hands-on engineering role for someone who likes writing simulator code, running experiments, and digging into why a prediction and a measurement don't match.

Working closely with our architects and the silicon, hardware, networking, and software teams, you'll capture workload behavior, extend simulation and analytical models of our silicon, interconnect, and multi-hop fabrics, and validate them against real hardware. You'll be comfortable moving across the stack, from the model graph through collectives to the network fabric, to track down where performance is going.

What You'll Do

  • Implement and extend simulation-based performance models for multimodal generative AI inference at rack, pod, and cluster scale, covering compute, memory, collective communication, and network fabric.
  • Model how serving strategies (tensor, pipeline, and expert parallelism, prefill/decode disaggregation, batching, and KV-cache placement) interact with Tensordyne silicon and fabric topology, and measure the effect on latency, throughput, and cost per token.
  • Build trace-capture and replay tooling that records real execution from our inference runtime and replays it under hypothetical silicon, system, and network configurations.
  • Model collective communication on multi-hop scale-out fabrics, including implementing custom collective algorithms designed for our topology.
  • Run calibration experiments on Tensordyne hardware as systems come up, compare them against model predictions, and fix the sources of error.
  • Run design-space sweeps and write up clear analyses that architects and engineering teams use in ASIC, fabric, and system configuration decisions.
  • Produce performance projections that support product and customer discussions.
  • Keep the modeling codebase fast, tested, and reproducible so other engineers can run it themselves.

What We're Looking For

  • Hands-on experience building performance models, simulators, or analytical tools for ML workloads, distributed systems, or computer architecture.
  • Solid understanding of distributed ML execution, including parallelism strategies, collective communication (All-Reduce, All-Gather, All-to-All, etc.), and how they scale.
  • Working knowledge of system architecture across compute, memory, interconnect, and networking, and the ability to reason about bottlenecks between them.
  • Experience comparing model predictions against real measurements, and debugging where they diverge.
  • Strong programming skills in C++ and Python, with clean, testable, maintainable code.
  • Ability to take a loosely defined performance question, break it into experiments, and deliver results with minimal hand-holding.
  • Clear written and verbal communication, especially when presenting data and trade-offs to other engineers.
  • MS or higher in Computer Science, Computer Engineering, Electrical Engineering, or a related field.

Nice to Have

  • Familiarity with LLM inference serving: batching, KV-cache management, disaggregated prefill/decode, and latency/throughput trade-offs.
  • Experience modeling or benchmarking collective communication libraries (NCCL, RCCL, or similar) on real clusters.
  • Background in data center or HPC networking: topologies, RDMA/RoCE, and congestion behavior.
  • Experience profiling ML workloads on accelerators (GPUs, TPUs, or custom ASICs).
  • Exposure to hardware/software co-design or early-stage architecture evaluation.
  • Publications or open-source contributions in ML systems, architecture, or networking.

 Tensordyne's culture was built on the following values

  • Put people first. We only succeed when our people succeed.
  • Ethics and integrity always; Being open, honest, and respectful of everyone.
  • Think Big. Be ambitious and have audacious goals of global scale.
  • Aim for excellence. Quality and excellence count in everything we do.
  • Own it and get it done. Results matter!
  • Make each person better together, than they would be as an individual.
  • Embrace each others’ differences, and embrace that there will be differences.

Tensordyne is an equal opportunity employer. We believe that a diverse team is better at tackling complex problems and coming up with innovative solutions. All qualified applicants will receive consideration for employment without regard to age, color, gender identity or expression, marital status, national origin, disability, protected veteran status, race, religion, pregnancy, sexual orientation, or any other characteristic protected by applicable laws, regulations and ordinances.

A note to Recruitment Agencies: Please don’t reach out to Tensordyne employees or leaders about our roles -- we’ve got it covered. We don’t accept unsolicited agency resumes and we are not responsible for any fees related to unsolicited resumes. Thank you for your understanding.

Skills

Python

Similar Jobs

6

Principal Systems Performance Engineer – Modeling

Micron·Longmont-MAX- Office, CO

2mo ago

Skillbridge Principal / Senior Modeling Simulation & Performance Analysis Systems Engineer

Northrop Grumman·Baltimore, MD

2mo ago

Skillbridge Principal / Senior Modeling Simulation & Performance Analysis Systems Engineer

Northrop Grumman·MDLI04, US

2mo ago

Principal / Senior Modeling Simulation & Performance Analysis Systems Engineer

Northrop Grumman·Baltimore, MD·Onsite

2mo ago

Principal / Senior Modeling Simulation & Performance Analysis Systems Engineer

Northrop Grumman·MDLI04, US·Onsite

2mo ago

Staff Systems Engineer, Signal-Path Performance & Modeling

Neurophos·Austin, Texas·Onsite

6d ago