- Salary
- $60 – $90/hr
- Workplace
- Remote
- Type
- Contract
- Department
- Engineering
- Education
- PhD
- Closing date
- Today
- Source
- CareersPage
Description
About the role
A frontier engineering-reasoning evaluation run in collaboration with a leading AI research lab. The work measures whether state-of-the-art models can reason from first principles in your engineering domain rather than retrieve facts from training data — and you get visibility into the model's internal reasoning on your own tasks.
What you'll do
Domain experts write simulations and specification sheets; the model attempts to design an artifact — controller gains, circuit parameters, geometry — that satisfies every spec.
You define the simulation, a set of specs with pass/fail thresholds, and the prompt. The model probes your simulation with a limited number of calls, then submits a final design. An agentic grader runs the simulation and scores spec satisfaction.
Who we're looking for
- Deep applied expertise in control systems, analog or RF circuits, power electronics, or mechanical design — PhD or equivalent industry track record.
- Comfortable writing simulations in Python and expressing acceptance criteria as quantitative, testable thresholds.
- Able to design a problem you can't shortcut yourself, but can still judge rigorously.
- Comfortable working in a browser-based studio and GitHub.
Support
Onboarding calls run daily, alongside internal tooling built to help you work faster. Prior model-evaluation experience is not required.