- Location
- Brazil
- Workplace
- Remote
- Department
- Pyxis
- Experience
- 30+ years
- Source
- Lever
Description
At CI&T, we help large enterprises transform the potential of AI into real business impact with AI Deployment, AI-native execution, and tech-integrated business solutions.
With 30 years of experience in technological transformation, we accelerate innovation with expertise in Agentic SDLC, Application modernization, Data & AI, Martech and Business strategy.
We are 8,000 CI&Ters across more than 25 countries, collaborating to build solutions with real impact. AI is already part of how we work, evolve, and innovate every day.
To move beyond traditional QA by owning the "Evaluation Harness," ensuring the AI agents are safe, compliant, and accurate through rigorous automated testing of non-deterministic outputs.
Key Responsibilities
- Own and manage the Langfuse eval harness, including the creation of "Golden Paths" and adversarial test scenarios.
- Validate regulatory quality gates (SR 26-2, NYDFS Reg 187) and run shadow mode validation to compare AI performance against human benchmarks.
- Test for data quality degradation, population segmentation errors, and human-in-the-loop (HITL) checkpoint reliability.
- Verify that the Designer’s output standards (explanation blocks, signal types) are maintained in all evaluation scenarios.
Required Skills and Qualifications
- Advanced proficiency in Python 3.11+, Pytest, and automated testing frameworks.
- Direct experience with AI Evaluation tools (Langfuse, RAGAS, or similar) and LLM-as-a-judge patterns.
- Knowledge of model risk management and regulatory compliance testing.
- Experience with "Shadow Mode" deployment strategies and statistical validation.
Soft Skills
- A "skeptical" and adversarial mindset to find edge cases.
- Uncompromising attention to detail regarding regulatory evidence.
- Collaborative approach to working with engineers to fix logic failures.