- Location
- Bogota, Colombia (Remote Friendly) · Bogotá, Bogotá, Colombia
- Workplace
- Remote
- Department
- Engineering
- Seniority
- Senior
- Experience
- 15+ years
- Source
- Greenhouse
Description
Robots & Pencils is an applied AI engineering firm building the next frontier of business architecture. We design and ship AI co-workers that integrate into enterprise operations and deliver measurable results for our clients. We're all in on AWS, combining deep UX capability with senior engineering talent to get AI into production fast and keep it there.
We’ve earned the trust of leaders across Consumer Products and Retail, Education, Energy, Financial Services, Healthcare, and Manufacturing and more, and earned a reputation as the nimble alternative to traditional global systems integrators. Founded in 2009, with delivery centers in Canada, the United States, Eastern Europe, and Latin America, we are smaller, faster, and more senior by design. Our teams average 15+ years of experience. We move fast, sweat the details, and build things that actually ship.
Position Overview
We’re looking for a Senior Testing Engineer to lead quality across complex software features, AI/ML systems, and integrated platforms. This role is ideal for an experienced engineer who is knowledgeable across both automation and AI testing, enjoys hands-on work, and is growing into owning testing strategy and architecture.
In this role, you will work as part of a cross-functional team, designing automation frameworks, validating ML models and AI workflows, and driving quality throughout the delivery process. You’ll collaborate closely with developers, ML engineers, and product teams to deliver high-quality features and develop the technical instincts that come from shipping things that matter.
Why This Role Matters
At Robots & Pencils, we design AI systems for a human world. Our name says it all. Robots and pencils means engineering paired with creativity, because every agent we ship has to work for real people in real workflows. That balance is baked into how we operate.
Every role here contributes directly to that mission. Here, you shape how AI systems integrate into enterprise operations, how teams move at real velocity, and how products create measurable impact for clients and the people they serve. We ship production-ready AI in 30 to 45 days. That pace demands people who take ownership, lead with craft, and care deeply about what they put their name on.
What You’ll Do
Craft & Delivery
- Design and develop scalable automation frameworks for UI, API, and integration testing (e.g., Cypress, Playwright, Selenium, PyTest, TestNG, Cucumber)
- Integrate automated tests into CI/CD pipelines and drive CI/CD quality gates (e.g., GitHub Actions, GitLab CI, Jenkins)
- Design AI-specific test strategies and validate ML models across key metrics including accuracy, precision, recall, bias, and drift
- Test data pipelines, feature engineering processes, and validate LLM responses for hallucination risk and output consistency
- Perform test planning, review automation code, and ensure best practices are applied across the test codebase
- Monitor model quality in production environments and support release validation and quality assurance processes
- Bring an AI-forward mindset to your daily work, using tools like Claude, Cursor, and other modern AI assistants to ship higher-quality work at pace
Collaboration & Communication
- Collaborate closely with developers, ML engineers, DevOps, and product teams across the full SDLC
- Communicate quality status, testing findings, and risks clearly to stakeholders
- Participate actively in sprint ceremonies, design reviews, and release planning
Leadership & Influence
- Lead quality efforts end-to-end with growing ownership of testing strategy and automation architecture
- Contribute to QA standards and best practices, and identify opportunities to improve coverage, reliability, and automation
- Begin mentoring junior engineers, sharing knowledge and supporting their growth
What You’ll Bring
- 3–4+ years of experience in QA, data QA, or AI testing, with solid knowledge across both automation and AI testing
- Strong programming skills in Python and at least one other language (e.g., Java, JavaScript)
- Hands-on experience designing and maintaining automation frameworks (e.g., Selenium, Cypress, Playwright, PyTest, TestNG, Cucumber)
- Strong understanding of the ML lifecycle and experience with model evaluation metrics (accuracy, precision, recall, bias, drift)
- Knowledge of AI testing methodologies including LLM validation, hallucination testing, and data pipeline validation
- Experience with CI/CD pipelines, API automation, and version control (e.g., GitHub Actions, Postman, Git)
- Familiarity with MLOps pipelines and model monitoring in production environments
- Strong understanding of Agile methodologies and experience with test management tools (e.g., Jira, TestRail)
- Demonstrable usage of AI-forward tools such as Claude and Cursor
- Experience with mobile automation, performance testing, LLM testing frameworks, prompt engineering, or cloud ML platforms is a plus (e.g., Appium, JMeter, AWS SageMaker, Azure ML)
Helpful Extras and Unique Skills
You’ll Do Well Here if You Are
- A doer. You see something broken and fix it. You'd rather move on clarity than wait for certainty.
- A fast learner who knows you don't know everything. The AI landscape changes weekly. You're senior enough to know better and curious enough to keep learning anyway.
- Direct in a way that makes the work better. You give honest feedback. You'd rather have the hard conversation than blow smoke.
- Obsessed with craft. You know genius is in the details. You ship exceptional, not perfect, and you don't put your name on work you wouldn't stand behind.
- Built for ownership. You honor commitments, admit mistakes fast, and back your teammates when a decision costs something. No handoffs, no finger-pointing.
- All in. You treat clients' businesses like your own. You take the work seriously without taking yourself seriously.
- Resourceful when the budget, timeline, or team is tight. Constraints don't slow you down. They sharpen you.
- Glad to be in the room with people who care as much as you do. Our teams average fifteen-plus years of experience. We hire people who push each other to do better work.