- Location
- Bangalore, India
- Workplace
- Hybrid
- Type
- Full-time
- Department
- Engineering
- Seniority
- Lead
- Experience
- 8+ years
- Source
- Workday
Description
Role Overview
We are seeking a high-potential, hands-on Lead Data & AI Operations Engineer to own and continuously improve the operational health, governance, controls, reliability, and efficiency of our enterprise Data & AI ecosystem.
This is a high-impact technical leadership role with end-to-end accountability for Data & AI Operations across the company. The successful candidate will establish the operating model, engineering controls, automation, observability, and governance required to run Data & AI platforms as reliable, secure, and cost-efficient enterprise services.
The ideal candidate combines deep Snowflake and Data Engineering expertise with a strong operations and controls mindset. This engineer will also design, build, and deliver technical solutions and platform capabilities required to achieve operational excellence and efficiency goals.
Key Responsibilities
· Supported end-to-end Data & AI Operations and Production Support across enterprise data platforms, data pipelines, analytics, BI, and AI/ML workloads, ensuring availability, reliability, performance, and SLA adherence.
· Provided day-to-day Snowflake production support and administration, including workload monitoring, query performance analysis, troubleshooting, access/RBAC management, capacity monitoring, and platform health checks.
· Supported and enhanced Data Engineering pipelines and workflows, troubleshooting data ingestion, transformation, orchestration, processing, and downstream data delivery issues across production environments.
· Supported AI/ML and GenAI workloads in production, including monitoring application and model-related jobs, data dependencies, API integrations, scheduled processes, failures, and overall operational health.
· Contributed to AI Operations (AIOps) capabilities by using AI/GenAI tools for incident analysis, log summarization, anomaly identification, troubleshooting assistance, knowledge retrieval, and faster root-cause analysis.
· Developed Python scripts, APIs, workflow automation, RPA, and AI-assisted automation to reduce repetitive operational activities, automate health checks and validations, accelerate issue resolution, and improve support productivity.
· Supported the implementation of intelligent monitoring and anomaly detection across data pipelines, Snowflake workloads, and AI services to proactively identify failures, performance degradation, unusual patterns, and operational risks.
· Assisted in developing automated remediation and self-healing operational workflows for common production issues, reducing manual intervention and improving the Resolution SLA.
· Used GenAI-based operational assistants to support troubleshooting, incident summarization, RCA preparation, log analysis, runbook recommendations, and knowledge management activities.
· Monitored production data pipelines, ETL/ELT jobs, orchestration workflows, AI workloads, APIs, and platform services, investigated failures, performed impact analysis, and coordinated timely service restoration.
· Performed data quality checks, reconciliation, validation, and root-cause analysis to identify data discrepancies and ensure accurate, complete, and reliable data delivery to downstream applications and AI/analytics workloads.
· Supported enterprise data platform controls covering data quality, access, security, privacy, metadata, lineage, change management, and production readiness.
· Monitored Snowflake and cloud consumption, performance, and utilization, identified inefficient queries and workloads, and supported optimization initiatives to improve performance and control platform costs.
· Built and maintained observability, monitoring, alerting, operational dashboards, automated health checks, and proactive notifications across Data and AI platforms.
· Managed Incident, Problem, Change, and Release Management activities, including production troubleshooting, service restoration, RCA documentation, change validation, deployment support, and permanent remediation of recurring issues.
· Supported DataOps, MLOps, AIOps, and DevOps practices, including CI/CD pipelines, testing, deployment, release validation, version control, monitoring, documentation, and production support.
· Worked closely with Data Engineering, Analytics, BI, AI/ML, Architecture, Security, Infrastructure, and business teams to troubleshoot production issues, manage dependencies, and implement platform improvements.
· Participated in on-call and production support activities, ensuring critical Data and AI incidents were addressed within agreed SLAs and appropriately communicated to stakeholders.
· Identified recurring operational issues and implemented automation, AI-assisted solutions, process improvements, and permanent fixes to reduce manual effort, prevent repeat incidents, and improve production stability.
· Contributed to continuous improvement by promoting operational discipline, automation-first practices, documentation, reusable runbooks, knowledge sharing, and Data/AI production support best practices.
What We Are Looking For
· 8–12 years of experience across Data Engineering, Data Platforms, Data Ops, Cloud Engineering, or Production Operations, with demonstrated technical leadership.
· Deep hands-on Snowflake expertise, including architecture, administration, SQL, performance tuning, workload management, security/RBAC, monitoring, troubleshooting, and optimization.
· Strong experience designing and building engineering solutions, not just administering or supporting Data Platforms.
· Demonstrated FinOps and cost optimization experience, with measurable outcomes in Snowflake/cloud consumption reduction, workload optimization, cost attribution, and efficiency improvement.
· Strong experience building automation using RPA platforms, Python, APIs, workflow automation, and AI/GenAI tools.
· Strong expertise with dBT and enterprise ETL/ELT technologies such as Fivetran, Informatica, and Azure Data Factory.
· Experience implementing DataOps, CI/CD, observability, data quality, governance, metadata, lineage, and automated platform controls.
· Strong understanding of production operations, incident/problem management, RCA, change management, and platform reliability engineering.
· Experience with enterprise BI platforms such as Power BI and Looker.
· Ability to operate as both a hands-on engineer and technical leader/manager, taking problems from identification through solution architecture, engineering, implementation, and measurable business outcome.
NOTES:
· Prefer candidates already residing in Bangalore.
· Standard Shift Timing is 12noon to 9pm, however this may vary depending on the business requirements.
· 3 Days work from office.
· Weekend on call support is required.