Hiring.Camp

Senior AI Operations Engineer

Allegion

·

Today

Location
Bangalore, India
Type
Full-time
Department
Engineering
Seniority
Senior
Experience
8+ years
Education
Bachelor
Closing date
Today
Source
Workday

Description

Creating Peace of Mind by Pioneering Safety and Security


At Allegion, we help keep the people you know and love safe and secure where they live, work and visit. With more than 40 brands, 14,000+ employees globally and products sold in 130 countries, we specialize in security around the doorway and beyond. 


Additionally, Allegion is proud to be recognized with the 2026 Gallup Exceptional Workplace Award (GEWA) for the third consecutive year, earning distinction in both the employee engagement and strengths categories. This year, Allegion also received Gallup’s With Distinction honor — a designation reserved for a select group of organizations that go above and beyond in building exceptional workplace cultures.

KEY RESPONSIBILITIES 


AI Platform Operations 

AI Platform Administration 

  • Operate and maintain enterprise AI platforms and services across development, test, and production environments. 
  • Support Azure OpenAI, Azure AI Foundry, Azure Machine Learning, and related Azure AI platform services. 
  • Operate the application hosting platform for AI services across Azure App Service, Azure Functions, and Azure Container Apps, including networking, scaling, and runtime configuration. 
  • Ensure platform availability, performance, reliability, scalability, and operational stability. 
  • Perform platform configuration, lifecycle management, upgrades, and operational maintenance. 

AI Service Deployment & Operations 

  • Deploy and operationalize AI and Generative AI workloads in controlled enterprise environments. 
  • Support production AI solutions including copilots, AI assistants, chatbots, agentic applications, retrieval-augmented generation (RAG) services, and machine learning applications. 
  • Manage release activities, prompt and model version control, runtime configurations, deployment validations, and operational handovers. 
  • Deploy and operate machine learning models and endpoints developed by Data Science teams, providing the pipelines, environments, and runtime platform they promote through. 
  • Troubleshoot deployment, integration, performance, and runtime issues across AI services. 

Container & Kubernetes Platform 

  • Support containerised AI workloads using Docker and Kubernetes operational best practices, including image build, hardening, registries, and runtime configuration. 
  • Help design, build, and operate Azure Kubernetes Service environments as the container platform is introduced for AI and platform workloads. 
  • Establish and monitor cluster health, capacity, autoscaling, networking, ingress, security policies, and operational readiness. 
  • Contribute to platform resiliency, backup, recovery, and disaster recovery readiness. 

Model Access & Gateway Operations 

  • Operate the governed model-access layer serving AI applications, agents, and developer tooling — provider routing, failover, model catalogue, and version management across Azure OpenAI, AWS Bedrock, and Google Vertex AI. 
  • Manage model access control, user and application entitlements, quota allocation, rate limiting, and credential provisioning for the gateway. 
  • Monitor and attribute token consumption, throttling and 429 events, latency, and cost by model, application, and consumer. 
  • Support gateway upgrades, provider onboarding, configuration changes, and controlled rollout of new model versions. 

Monitoring, Reliability & Support 

Monitoring & Observability 

  • Implement monitoring, logging, alerting, dashboards, and operational metrics for AI platforms and services. 
  • Monitor AI system utilization, availability, latency, cost consumption, token usage, and performance trends. 
  • Monitor AI-specific signals including model and agent quality, quota and throttling events, data or model drift, and grounding index freshness. 
  • Use Azure Monitor, Application Insights, Log Analytics, and OpenTelemetry instrumentation to identify incidents and improvement areas. 
  • Provide operational insights to engineering, security, governance, and leadership stakeholders. 

Incident & Problem Management 

  • Investigate and resolve platform incidents, production issues, service disruptions, and operational risks. 
  • Perform root cause analysis and document corrective and preventive actions. 
  • Support production support processes, escalation management, and incident communications. 
  • Participate in post-incident reviews and implement service reliability improvements. 

Operational Excellence 

  • Create and maintain runbooks, support procedures, platform documentation, and operational knowledge articles. 
  • Automate recurring support, validation, monitoring, and deployment activities. 
  • Continuously improve platform efficiency, stability, security posture, and cost governance. 
  • Support service-level objectives and operational performance targets for AI platforms. 

Cloud Engineering & Automation 

Infrastructure as Code 

  • Develop and maintain Infrastructure as Code using Terraform, including reusable modules, remote state management, validation, and peer review. 
  • Automate Azure AI platform provisioning, configuration, policy enforcement, and environment standardization across development, test, and production. 
  • Manage controlled promotion of infrastructure changes across environments, with drift detection, change tracking, and infrastructure compliance. 
  • Support configuration consistency, version control, and secure state management for all platform infrastructure. 

CI/CD & Release Engineering 

  • Build and maintain CI/CD pipelines for AI platform services, application workloads, containers, and supporting infrastructure. 
  • Support deployment automation using Azure DevOps, GitHub Actions, or equivalent tools, including approvals, artifact management, and tested rollback. 
  • Implement validation checks, release controls, and operational readiness gates, including automated model and prompt evaluations, content safety, latency, and cost checks before production release. 
  • Bring unmanaged or manually deployed services under governed, automated delivery. 
  • Maintain traceability and reproducibility across code, configuration, models, prompts, evaluation results, and deployed versions. 

Developer Enablement & Self-Service 

  • Build and maintain reusable Terraform modules, pipeline templates, and standardized provisioning patterns consumed by other engineering teams. 
  • Establish golden paths and self-service capabilities that allow AI and application teams to onboard workloads with reduced manual effort and lower operational risk. 
  • Promote reusable platform patterns and controlled deployment models across environments. 
  • Maintain platform documentation, module usage guidance, and onboarding references for consuming teams. 

Security & Identity Management 

  • Implement secure cloud configurations and platform controls for AI services, including private networking, private endpoints, and network security controls. 
  • Support identity and access management using managed identities, workload identities, RBAC, and least-privilege practices. 
  • Manage secrets, certificates, keys, rotation, and secure configuration using Azure Key Vault or approved tools. 
  • Implement and operate privileged access controls, including just-in-time elevation and periodic entitlement review. 
  • Remediate identified platform security and hardening findings, and prevent recurrence through platform defaults, policy, and pipeline checks. 
  • Partner with Information Security teams to address vulnerabilities, access reviews, and security findings.

REQUIRED QUALIFICATIONS 


Education 

  • Bachelor's degree in Computer Science, Information Technology, Engineering, Data Science, or a related STEM discipline, or equivalent practical experience. 

Experience 

  • 8+ years of experience in Cloud Platform Engineering, Infrastructure Engineering, DevOps, Platform Operations, SRE, MLOps, or AI Operations. 
  • Hands-on experience supporting Microsoft Azure-based enterprise environments and production workloads. 
  • Experience operating AI, Machine Learning, or Generative AI platforms in controlled production environments. 
  • Demonstrated experience owning a shared platform, infrastructure module library, or delivery pipelines consumed by other engineering teams — not solely application-level delivery. 
  • Experience supporting Kubernetes-based solutions and containerized workloads, preferably Azure Kubernetes Service. 
  • Exposure to AI governance, information security, compliance, operational risk, and production support practices. 

REQUIRED SKILLS 


Technical Skills 

  • Strong hands-on experience with Microsoft Azure cloud services and production cloud operations. 
  • Hands-on experience operating Azure application hosting services — App Service, Azure Functions, and container-based compute. 
  • Experience with Azure OpenAI, Azure AI Foundry, Azure Machine Learning, and related Azure AI platform services. 
  • Experience with containerized workloads using Docker, and working knowledge of Kubernetes sufficient to operate and build toward a production cluster platform. 
  • Strong understanding of cloud networking, private endpoints, DNS, NSGs, load balancing, and secure platform architecture. 
  • Experience with Infrastructure as Code, preferably Terraform, including reusable modules, remote state, and automated environment provisioning. 
  • Experience with Azure DevOps or GitHub Actions for CI/CD pipeline design, release management, and deployment automation. 
  • Automation and scripting proficiency in Python, and at least one of Bash or PowerShell. 
  • Strong Linux, identity and access management, secret management, API, and Git skills. 
  • Knowledge of monitoring, observability, logging, alerting, incident management, and root cause analysis. 
  • Understanding of the AI/ML operational lifecycle, model and service deployment, runtime support, and platform management. 
  • Working knowledge of generative AI components including prompt and model versioning, agents, RAG, vector stores, evaluations, groundedness, and content safety controls. 
  • Familiarity with AI governance, responsible AI, privacy, security, and compliance best practices. 

Preferred Skills & Experience 

  • Hands-on Azure Kubernetes Service operations: cluster build and upgrade, node pools, ingress, autoscaling, network and security policy, and production troubleshooting. 
  • Experience operating multi-provider model access across AWS Bedrock and Google Vertex AI alongside Azure OpenAI, and gateway technologies such as LiteLLM. 
  • Azure API Management policy authoring, including token validation, rate limiting, and backend protection. 
  • Helm or Kustomize; GitOps tooling such as Argo CD or Flux; Prometheus or Azure Managed Grafana. 
  • Terraform CDK (CDKTF) in TypeScript, or comparable programmatic infrastructure as code. 
  • Microsoft Entra ID governance, including Privileged Identity Management and Conditional Access. 
  • Azure AI Search, vector databases, and event-driven or API-driven application architectures. 
  • Experience building internal developer platforms or self-service engineering capabilities. 

Tools & Platforms Preferred 


  • Azure OpenAI 
  • Azure AI Foundry 
  • Azure Machine Learning 
  • Azure Kubernetes Service 
  • Azure App Service 
  • Azure Functions 
  • Azure Container Apps 
  • Azure API Management 
  • Terraform 
  • Azure DevOps 
  • Docker 
  • Azure Container Registry 
  • Azure Monitor 
  • Application Insights 
  • Log Analytics 
  • Azure Key Vault 

Analytical & Problem-Solving Skills 


  • Strong troubleshooting, analytical thinking, and root cause analysis capabilities. 
  • Ability to resolve complex production issues across cloud, platform, AI service, and integration layers. 
  • Ability to analyze operational metrics, identify trends, and recommend reliability and cost optimization improvements. 
  • Strong ownership mindset with the ability to work independently in an individual contributor capacity. 

Communication & Soft Skills 


  • Strong written and verbal communication skills with technical and non-technical stakeholders. 
  • Ability to collaborate across Data & AI, Enterprise Architecture, Information Security, Governance, Cloud, and business teams. 
  • Strong documentation, follow-up, prioritization, and stakeholder engagement skills. 
  • Self-driven, accountable, and comfortable supporting production environments and governance-related activities. 

LOCATION & WORK EXPECTATIONS 


  • Location: Bangalore. Standard IST working hours. 
  • This is a hands-on individual-contributor role reporting to the Team Lead – AI Operations. 
  • Occasional after-hours support may be required for critical incidents or planned production activities. 

We Celebrate Who We Are! 

Allegion is committed to building and maintaining a diverse and inclusive workplace.  Together, we embrace all differences and similarities among colleagues, as well as the differences and similarities within the relationships that we foster with customers, suppliers and the communities where we live and work. Whatever your background, experience, race, color, national origin, religion, age, gender, gender identity, disability status, sexual orientation, protected veteran status, or any other characteristic protected by law, we will make sure that you have every opportunity to impress us in your application and the opportunity to give your best at work, not because we’re required to, but because it’s the right thing to do.   We are also committed to providing accommodations for persons with disabilities. If for any reason you cannot apply through our career site and require an accommodation or assistance, please contact our Talent Acquisition Team.



© Allegion plc, 2023 | Block D, Iveagh Court, Harcourt Road, Dublin 2, Co. Dublin, Ireland

REGISTERED IN IRELAND WITH LIMITED LIABILITY REGISTERED NUMBER 527370

Allegion is an equal opportunity and affirmative action employer

Privacy Policy

Skills

PythonTypeScriptAWSAzureDockerKubernetesTerraformCI/CDLinuxMachine LearningData ScienceGitGitHubDevOpsSRECompliance

Similar Jobs

30

Senior Strategy Manager - AI & Operations

Gympass·Brazil

Today

AI Operations Senior Manager

Gympass·Brazil

Today

Senior Prog Manager, Ops Int, Operations Integration & AI Transformation , MEA-Tr

Amazon·Remote

1d ago

Senior AI Engineer, Agentforce Operations

Salesforce·California - San Francisco, US +2·Onsite, Remote

2d ago

Sr. AI Strategy and Operations Analyst

Applovin·Toronto

5d ago

Senior Human Data Operations Partner, AI

Prolific·North America +1·Hybrid

1w ago

Sr Associate Support Engineer, Adaptive Planning AI, ML & Platform Operations

Workday·IND.Pune, India

1w ago

Senior Human Data Operations Partner, AI

Prolific·Mexico +1·Hybrid

1w ago

Senior Product Operations Specialist, AI Enablement

Grab·Petaling Jaya, Malaysia·Onsite

2w ago

Senior Product Manager (R&D – AI Platform Operations)

Virtuos·Singapore, SG

2w ago

Senior Human Data Operations Partner, AI

Prolific·US +1·Hybrid

2w ago

Sr. Accountant, AI Operations

Spacex·Palo Alto, CA +1·Hybrid

3w ago

Senior Business Manager - AI Strategy & Operations

Capitalone·McLean, VA

3w ago

Sr Manager, AI Technology Operations

Gapinc·SF - 2 Folsom, US +1·Remote

3w ago

Senior AI Care Operations Manager

Springhealth66·New York·Hybrid

3w ago

Senior AI Specialist (Operations) (x/f/m)

Doctolib·Berlin, Germany

3w ago

Sr. AI Delivery & Operations Engineer

Vizient·Edina, MN +2

4w ago

Sr. AI Delivery & Operations Engineer

Vizient·Edina, MN 55435 +2

4w ago

Sr. Analyst, HR AI & Operations

Fortinet·Sunnyvale, CA·Onsite

1mo ago

Head of TEM Operations & AI Delivery (Senior Manager level) / Responsable des opérations TEM et de la livraison de l'IA

Upland Software·Ottawa, ON

1mo ago

Senior Operations AI Engineer

Pwc·Athens - Kifisias Av. 65, Greece

1mo ago

Senior Analyst, AI & Analytics - Operations

InterContinental Chicago Magnificent Mile·GA, US·Hybrid

1mo ago

Senior Software Engineer - AI Operations Team (Open to hiring at the Lead Software Engineer Level)

Wellmark·Des Moines, IA·Hybrid

1mo ago

Senior AI Operations Engineer

Axiscapital·Alpharetta, GA +1

1mo ago

Sr. Specialist - Platform Operations (AI & Agentic Systems)

Nasdaq·Philadelphia - FMC Tower, US·Hybrid, Onsite

1mo ago

Sr. Specialist - Platform Operations (AI & Agentic Systems)

Nasdaq·CA-Toronto-York St 24, 25·Hybrid, Onsite

1mo ago

Senior AI Engineer, KYC Operations - Vice President

citibank·Mumbai, MH

1mo ago

Senior Cloud & Enterprise Architect (Cloud Services & AI-Augmented Operations)

Robert Bosch·Bengaluru, India

1mo ago

Senior AI Infrastructure & Platform Operations Engineer (remote in the EU)

Mirantis·Sofia, Sofia City Province·Remote

1mo ago

Senior AI Infrastructure & Platform Operations Engineer (remote in the EU)

Mirantis·Barcelona, CT·Remote

1mo ago