- Location
- India GCC-Puppalaguda Village
- Type
- Full-time
- Department
- Engineering
- Seniority
- Senior
- Experience
- 5+ years
- Source
- Workday
Description
We’re determined to make a difference and are proud to be an insurance company that goes well beyond coverages and policies. Working here means having every opportunity to achieve your goals – and to help others accomplish theirs, too. Join our team as we help shape the future.
Key Responsibilities
Model Monitoring & Observability
- Collaborate with Data Scientists to onboard models into monitoring systems (e.g., MHM database) and maintain configuration files for drift and performance checks.
- Support retraining workflows when drift or degradation is detected.
- Implement monitoring frameworks to track data drift, training-serving skew, and model performance metrics in production environments.
- Utilize observability tools like Arize to set up monitors for accuracy, precision, recall, and other KPIs.
- Configure automated alerts for anomalies in input/output distributions and performance degradation.
Data Quality Assurance
- Perform QC checks on incoming data to validate integrity and completeness before model scoring.
- Develop pipelines for continuous validation of data sources and ensure compliance with quality standards.
Performance Evaluation
- Monitor KPIs such as accuracy, precision, recall, and fairness across demographic slices.
- Conduct root-cause analysis for performance degradation and recommend retraining strategies.
Infrastructure & Automation
- Build and maintain CI/CD pipelines for deploying monitoring solutions and model updates.
- Leverage cloud technologies for scalable monitoring and orchestration.
Documentation & Reporting
- Maintain detailed logs of monitoring activities, thresholds, and alerts.
- Provide periodic reports on model health, including drift metrics and performance trends.
Required Skills & Experience:
• Bachelor’s or Master’s degree in Computer Science, Engineering, or a closely related field; 5+ years of professional experience with a Bachelor’s degree, or 3+ years of experience accepted with an advanced engineering degree combined with applied academic, internship, or hands‑on project experience focused on machine learning, GenAI, or full‑stack software development.
• 3+ years of experience collaborating with engineering leads and data science teams on the design, development, and scaling of machine learning systems, with a focus on reliability, scalability, and security.
• 4+ years of hands‑on experience working with public cloud platforms such as Google Cloud Platform, AWS, and/or Azure to deploy, operate, and scale ML workloads.
• 2+ years of hands‑on experience building Infrastructure as Code (IaC) using tools such as Terraform to support ML environments and pipelines.
• 5+ years of software development experience using programming languages such as Python, Java, or equivalent object‑oriented or scripting languages, including production‑grade ML or data applications.
• 3+ years of experience working with machine learning and AI frameworks or platforms such as TensorFlow, scikit‑learn, Anaconda, SageMaker, Vertex AI, or Agentic AI tooling.
• 3+ years of experience understanding model drift and data drift concepts and contributing to the design and implementation of monitoring, validation, and retraining strategies for ML and AI systems.
• 3+ years of working knowledge of CI/CD and MLOps practices, including automated testing, model validation, automated deployments, and integration pipelines.
• 3+ years of experience working in agile development environments, with SAFe experience required.
• 4+ years of experience collaborating and partnering with business leaders, engineering teams, and data science stakeholders to deliver machine learning solutions aligned to business outcomes.
• Excellent written and verbal communication skills, with the ability to explain technical concepts clearly to diverse audiences.
• Familiarity with agentic productivity and AI developer tools such as Claude Code, Gemini CLI, OpenAI Codex, or similar tools is a plus.
• Experience enabling or supporting enterprise AI services such as Gemini Enterprise, Amazon Q, Microsoft Copilot, or similar platforms is a plus.
• Strong analytical, critical thinking, and problem‑solving abilities, with a demonstrated willingness to challenge the status quo, take ownership, and drive innovation and continuous improvement in AI platform development.
• Ability to work both independently and collaboratively in a fast‑paced, agile environment, demonstrating accountability, initiative, and a bias for action.
Nice to Have
- Understanding of Agile framework is a plus
Experience leveraging Gen AI tools to accelerate insight generation