- Location
- Paris
- Type
- Full-time
- Department
- Engineering
- Source
- RecruiterFlow
Description
We are building the AI coworker for legal and professional services — automating emails, document analysis, and complex workflows in high-stakes industries.
We are not looking for someone to “use GPT APIs.” We are looking for someone to own and scale a production AI system that users rely on daily.
What you will own
You will be responsible for one primary outcome: making our AI outputs reliable, fast, and indispensable in real workflows.
Concretely:
- Design and evolve our LLM / agent architecture.
- Own output quality across key use cases (emails, document analysis, etc.).
- Build evaluation systems (datasets, metrics, regression detection).
- Drive fast iteration loops from production data.
- Improve retrieval, reasoning, and tool usage.
- Ensure production reliability (latency, failure modes, fallbacks).
- Work directly with product and founders on what to build and why.
What this role is really about
Most teams fail because they don’t know what “good output” means, lack proper evals, iterate randomly, or overuse agents.
Your job is to fix that. You will turn vague user problems into structured AI systems with measurable performance that improve every week.
What You Need to Be Excellent At
- Shipping Real LLM Systems
- Evaluation-Driven Development
- Debugging Complex Failures
- Speed of Iteration
- Strong Technical Judgment
What we don’t care about
- Number of years of experience.
- Whether you’ve used a specific framework.
- Fancy research credentials.
If you can build, debug, and improve real systems, you’re a fit.
What success looks like (First 90 Days)
- A clear evaluation framework established for core use cases.
- Measurable, quantitative improvement in output quality.
- Faster iteration cycles across the team.
- Significantly reduced hallucinations and failure modes.
- Robust, scalable system architecture decisions.
Tech Stack (Context, Not Strict Requirements)
- Language/Backend: Python (FastAPI)
- Database: Postgres
- Cloud: Google Cloud Platform (GCP)
- Orchestration: LangGraph / LangChain (evolving)
- Observability: PostHog (analytics), Langfuse (LLM tracing)
- LLM Infrastructure: Azure OpenAI / Multi-provider APIs