- Location
- Bogotá, Colombia · Medellin
- Workplace
- Remote
- Type
- Full-time
- Department
- R&D - Technology
- Seniority
- Senior
- Experience
- 30+ years
- Education
- PhD
- Source
- Lever
Description
We are building the agentic platform that powers Caseware's next generation of audit and accounting products, and applied science is how we measure and improve quality across our agents. This is a role for someone who thinks in experiments, has built evaluations for LLM-based systems, and wants to build the tools that other teams and our customers rely on to ship their own agents.
You will design and run experiments that turn ambiguous quality questions into measurable results, and own the evaluation methodology that internal teams and customers depend on. You will build the AI-based capabilities behind our end-to-end agent builder, including the synthetic data generation and eval builders that let teams and customers create, evaluate, and deploy their own agents. You will also work on something genuinely novel: self-learning, self-improving systems that get better both offline and online. A significant part of that is agentic memory: deciding what is worth remembering and validating the criteria that let agents compound useful knowledge on behalf of customers over time, while rigorously upholding the legal and contractual obligations owed to them and their clients. At the staff level, you will also influence technical direction and mentor others.
📍 Location: This is a fully remote position located in Colombia.
Contact
Maira Russo - Senior Talent Acquisition Partner