- Salary
- $10k+/yr
- Location
- San Francisco, California, US
- Type
- Internship
- Department
- Operations
- Seniority
- Internship
- Source
- Y Combinator
Description
Mission
Keeper is an AI tax platform for complex returns, with a focus on serving solo business owners and freelancers in the U.S.
Interest in solopreneurship has exploded over the past few years, in no small part aided by the advent of AI. Every one of these new businesses needs help with taxes.
Today, the tools are scattered. Many solos do nothing at all until tax time, when they scramble to file in Turbotax and end up drastically under-claiming deductions. Others cobble together Quickbooks for their expenses, Gusto for 1099 forms, and a CPA for handling tax returns. It’s a lot to keep track of, and it sucks. Keeper handles all of this work in a seamless, AI-native platform.
Keeper serves over 40,000 paying customers and generates over 8 figures of annualized revenue.
The Role
Given our core business, Keeper is uniquely positioned to contribute to the public discourse happening at the intersection of AI and solopreneurship.
To tap into this moment, we're building a public AI benchmark that measures how well frontier models (ChatGPT, Claude, Gemini, etc.) can do real-world self-employment tasks: creating marketing materials, organizing a work portfolio, categorizing a year of expenses, reviewing a client contract, etc.
Think GDPval (OpenAI's benchmark for professional work), but for the tens of millions of Americans who work for themselves. At the end of the project, we'll publish the results, the methodology, and a leaderboard.
As the Research Operations Intern, you’ll co-lead the task compilation portion of this project with the CEO. This is a dynamic role where daily responsibilities will shift dramatically over the course of the internship. At the onset, you’ll work with Keeper’s freelancer customers to design the workflow for the first set of tasks. You’ll study the literature and other benchmarks to assess questions like:
- What does a high-quality benchmark task look like?
- How do we vet qualified Keeper customers to help us with this project?
- What does an ideal (human) completed task look like?
- How do we scale up operational processes to curate 500+ specified tasks?
Examples of projects you’ll own
Discovery interviews:
- Run structured interviews with solopreneurs to get an initial idea for the landscape of possible interesting tasks
Task recruitment:
- Continue interviews to recruit customers to create tasks, with a deliverable that includes full work specifications and a grading rubric
Synthetic input creation:
- Build realistic prompts, decks, bank exports, receipt photos, contracts, and email threads that each task consumes.
Baseline dataset compilation:
- Recruit and coordinate 30-40 real solopreneurs who complete benchmark tasks themselves, so we can measure models against how owners actually do this work.
Why take this role?
This is a role for an ambitious, early-career individual looking to break into AI/tech or launch their own startup someday. By working directly with David, Keeper’s CEO, you’ll get a front-row seat into the workflows of a venture-backed startup operating at the cutting edge of applied AI. At the end of the project, assuming success, your name will be included in the benchmark publication paper.
What we're looking for in a candidate
- You're obsessively detail-oriented. Benchmarks live or die on data quality, and you’ll hold the line on ensuring that shoddy work doesn’t ship into the final product.
- You communicate clearly. One area where frontier AI models have quite a bit of catching up to do is in effective and clear writing. Task prompts, rubrics, and expert guidelines are all writing problems that demand excellent communication skills.
- You're organized & get along with people. In this role, you’ll need to be comfortable hopping on a lot of video calls on a tight schedule, while juggling multiple tasks at once.
- You’re ready for an intense project. We’re going to get a lot done over the course of 10 weeks. To put it bluntly, raw hours worked will count. This role isn’t for the faint of heart.
Location/duration
- We have a mix of remote and in-office employees. We are looking for someone who wants to work one or more days per week in our SF office (Financial District)
- This is a 10-week position. Expected start date is September 28, 2026.
Compensation
$10k / month