- Salary
- $150k – $250k
- Location
- Remote
- Workplace
- Remote
- Department
- Product
- Seniority
- Senior
- Source
- Greenhouse
Description
Flodesk is recognized in the Inc 5000 as one of the world's fastest-growing email marketing companies, built to help entrepreneurs sell online and design emails that people love to get. We're committed to giving small businesses simple and intuitive tools that help them grow, nurture, and monetize their email list.
We’re a remote-first company headquartered in San Francisco with a globally distributed team, including in-person hubs in Da Nang (Vietnam), Barcelona (Spain), and Menlo Park (California). Our team reflects the diversity and creativity of the people we serve. Join our mission to level the playing field for small business owners through good design.
About the role
Flodesk is building toward a future where small business owners do not need to be expert marketers to grow. Today, our AI helps members create and edit emails, but that is only the starting point. We are building a system that can understand a member's business, brand, audience and past performance, then help turn a goal into a complete marketing campaign across email, workflows, forms, sales pages, segmentation, scheduling, analytics and recommendations.
Reporting to the Head of Product, AI Systems and Core Experience, you will shape how that system behaves and how we know it is good. You will own Flodesk's prompts, agent instructions, quality definitions, evaluation criteria and continuous-improvement loop by building the technical tooling, prototypes and evaluation infrastructure that make this work measurable and scalable. This is not a role for writing clever prompts in isolation. It is hands-on systems work at that connects engineering, product and natural language. You will need to understand how context, models, tools, product state and structured outputs come together to create a member experience, then identify the right layer to change when that experience fails.
What you'll do
- Manage member-facing AI behavior across prompt-to-create, agentic editing and future jobs such as segmentation, scheduling, analytics and recommendations.
- Create, test, version and document prompts and agent instructions, with a clear record of what changed, why it changed and how behavior improved
- Define how the system should use context, member data, tools and product state. Build and prototype the agent loops, tools and structured interfaces needed to make those behaviors work
- Own Flodesk's AI evaluation practice, including representative datasets, behavioral scenarios, scoring rubrics, regression suites and human review
- Build and operate eval harnesses, monitoring and quality dashboards so the team can measure quality without relying on one-off manual checks
- Diagnose failures across prompts, context, orchestration, models, tools, data and product code, then fix the right layer or partner with the engineer who owns it.
- Set quality baselines and release gates for prompt, model and agent changes. Turn production failures, traces and member feedback into durable regression cases
- Partner with product, design, marketing and copy experts to encode Flodesk's point of view into the system, while bringing your own judgment about how the product should behave
What you bring
- A track record of building or meaningfully improving production LLM or agentic products, not just prototypes
- Strong software engineering skills in Python, TypeScript or a similar language that let you build reliable internal tools, evaluation systems and prototypes
- The interaction between prompts, context, tools, agent loops, orchestration and product state is clear to you
- LLM evaluation, observability and experimental design are areas where you have direct experience, and you know how to combine automated checks with human judgment
- A strong product sense that helps you translate ambiguous member feedback into testable quality definitions and concrete system changes
- Systems thinking and reusable patterns guide your work, so today's email experience gets better without building in assumptions that make tomorrow's surfaces harder to build
- Across technical and nontechnical disciplines, your communication is clear and you can explain tradeoffs involving quality, latency, reliability and cost
Bonus points if you've:
- Built content-generation, creative or marketing tools
- Built multi-turn agents that use tools and preserve user intent across edits
- Used Langfuse, Braintrust or comparable evaluation and observability platforms
- Worked on products for small businesses, creators or marketers
What we bring
- $150,000 - $250,000 base salary, depending on your location and experience
- Tier 1 city residents: $165,000 - $250,000
- All other US based locations: $150,000 - $235,000
- Fully paid health insurance for individual coverage
- 16 weeks paid parental leave for non-birthing parents; 22 weeks paid maternity leave for birthing parents
- Unlimited flexible time off
- 401(k) match (US employees only)
- $1,000 annual stipend for learning and development