- Location
- Bangalore, IN
- Department
- Engineering
- Seniority
- Lead
- Education
- Master
- Closing date
- Today
- Source
- iCIMS
Description
Overview
In this role, you will work in a fast-paced, agile environment with a diverse team that has a true passion for technology, transformation, and outcomes. You will help build and operate the AWS-based platform services and infrastructure that power our multi-agent AI workflows for the enterprise, as part of the AI CoE
This is a senior, hands-on engineering role. You will implement and run the core services of the AI/Agentic platform, provisioning infrastructure as code, building backend and orchestration services, and hardening them for production, translating the architecture and standards set by the AI Platform Lead into reliable, scalable systems
Here at Waters, we look to our team members to be versatile and enthusiastic about tackling new problems, display strong ownership, and remain focused on business outcomes
Responsibilities
- Build, deploy, and operate AI/Agentic platform services in AWS and the orchestration systems that power multi-agent AI workflows, including agent lifecycle, routing, and coordination
- Provision, manage, and version cloud infrastructure as code so platform environments are reproducible, auditable, and safe to change
- Build and operate the shared platform capabilities that agent teams consume, RAG, memory, and human-in-the-loop services
- Build & maintain CI/CD and MLOps pipelines so AI services stay reliable and scalable in production; apply the CoE standards and contribute improvements back
- Own the reliability of shared platform services: monitoring, logging, and observability for model and agent performance and health; and responding to and resolving production issues
- Apply the platform’s governance and security controls in the services you build, cost/spend guardrails, connector permission scope, and prompt-injection and agent-attack mitigation
- Design reusable, well-documented service abstractions, templates, and SDKs that application and agent teams build on, and partner with those teams to bring new AI capabilities into production and continually improve the experience of building on the platform
- Provide constructive code reviews and mentor junior and mid-level engineers to ensure the engineering standards across the platform team
- Document platform services, interfaces, and patterns so consuming teams can self-serve, and support knowledge-sharing across the AI CoE
- Stay abreast of AI industry trends and proactively identify opportunities for improvement and adoption
Qualifications
- Bachelor or Master degree inCS, AI/ML, Data Science, or equivalent practical experience.
- 5+ years of industry experience in software engineering or ML
- Working understanding of agent frameworks and architectures (e.g., LangChain, LangGraph, or CrewAI), sufficient to design, build, and operate the platform services that support them, including RAG and memory
- Strong programming skills in one or more languages (e.g., Python, Java, Go, Node.js, or TypeScript), with solid experience building backend services and distributed systems in production, including microservices and event-driven architectures, and designing and versioning REST or GraphQL API contracts
- Solid AWS experience (GCP or Azure acceptable) across compute, eventing, API management, and IAM, plus containerization (Docker/Kubernetes) and Infrastructure-as-Code (CDK or Terraform)
- Hands-on experience operating an agent runtime such as AWS Bedrock AgentCore (or an equivalent) is a strong plus
- Hands-on experience building CI/CD pipelines and applying MLOps/LLMOps practices for deploying, scaling, and monitoring LLM-powered services
- Experience with monitoring/logging stacks such as Grafana, Prometheus, or the ELK stack
- Working knowledge of vector databases, memory systems, and human-in-the-loop workflows.
- Strong collaboration skills across platform engineering and product teams, with clear communication and a bias toward shared standards and reusable solutions
- Curious mindset with a strong desire to stay ahead of AI/ML advancements and enterprise best practices