Hiring.Camp

Operations Engineer

Accenture

·

Yesterday

Location
Bengaluru, BDC7C, India
Type
Full-time
Department
Engineering
Experience
5+ years
Source
Workday

Description

Project Role : Operations Engineer
Project Role Description : Support the operations and/or manage delivery for production systems and services based on operational requirements and service agreement.
Must have skills : Splunk Enterprise Architecture and Design, Event management with AIOPS , Splunk Enterprise Observability & ITSI
Good to have skills : NA
Minimum 5 year(s) of experience is required
Educational Qualification : 15 years full time education

Summary:
A Tools & Platforms Site Reliability Engineer (SRE) ensures the reliability, availability, performance, and continuous improvement of the infrastructure engineering tooling estate — spanning observability platforms, infrastructure-as-code tooling, CI/CD pipelines, ITSM platforms, internal developer portals, secret management, and AI-augmented operations tooling. The role applies a software engineering discipline to platform operations — building automated remediation, establishing SLIs and SLOs for tooling platforms, reducing toil through systematic automation, and owning reliability outcomes end to end across the four tooling pillars.

At Level 7 / 8, this individual operates at the intersection of platform engineering, SRE practice, and AI operations — not just keeping platforms running but continuously raising their reliability ceiling. A distinctive aspect of this role is ownership of LLMOps reliability — ensuring AI-augmented operations tooling (runbook automation pipelines, agentic ITSM workflows, RAG knowledge bases, and AI alert correlation services) meets production-grade SLOs in a regulated financial services environment.
Observability SRE
– ELK/Splunk
– OpenTelemetry
– SLI/SLO/Error Budget
IaC & Automation SRE
– Terraform
– Ansible/Chef
– GitHub Actions/ArgoCD
– HashiCorp Vault
– Policy-as-Code
ITSM & DevOps SRE
– ServiceNow
– xmatters
– Backstage IDP
– Jira/Confluence
– CMDB Reliability
AI Ops SRE
– LLMOps Reliability
– Agentic ITSM SRE
– AI Alert Pipeline SRE
– RAG Platform SRE
– Model Observability


Roles & Responsibilities:
– Own reliability of observability platforms —Splunk— defining and maintaining SLIs, SLOs, and error budgets for metrics pipelines, alerting systems, and dashboard availability across all infrastructure tiers
– Engineer auto-remediation for common observability failures — scraper restarts, index rollover failures, ingest pipeline blockages — reducing MTTR and eliminating repetitive manual toil
– Implement and govern OpenTelemetry instrumentation standards across the infrastructure estate — ensuring telemetry coverage is comprehensive, consistent, and production-grade
– Drive observability-as-code adoption — dashboards, alert rules, SLO definitions, and recording rules version-controlled and deployed through GitOps pipelines with automated testing
– Perform capacity planning and performance analysis for observability platforms — managing cardinality growth, storage retention, query performance, and ingest throughput at scale
– Lead blameless post-mortems for observability platform failures — producing structured RCA with systemic preventive actions that address root causes rather than symptoms
AI-Augmented Operations SRE
– Own reliability of LLMOps pipelines — monitoring model API health (OpenAI, Anthropic Claude, Google Gemini), prompt execution success rates, token consumption, latency SLOs, and cost anomaly alerting for AI-augmented operations tooling
– Engineer reliability for agentic ITSM workflows — LangChain, LlamaIndex, CrewAI — including agent execution health, tool call success rates, human-in-the-loop handoff reliability, and automated failure recovery
– Build observability for RAG knowledge base platforms — vector database (Pinecone, Weaviate, ChromaDB) availability, retrieval latency SLOs, embedding pipeline health, and index freshness monitoring
– Implement AI alert correlation reliability — ensuring LLM-based alert grouping pipelines maintain accuracy and availability SLOs, with fallback to rule-based alerting during AI platform degradation
– Define and enforce LLMOps governance frameworks — prompt version control, model evaluation pipelines, output quality monitoring, and FSI compliance controls (audit logging, data residency) for AI operations tooling
– Lead blameless post-mortems for AI tooling failures — diagnosing model degradation, hallucination events, pipeline failures, and agent workflow breakdowns with preventive actions that meet FSI audit standards

Professional & Technical Skills:
Certifications
-Terraform Associate or Professional
Splunk Professional
-AWS DevOps Engineer Pro or GCP DevOps Engineer
HashiCorp Vault Associate
-Certified Kubernetes Administrator (CKA)
ITIL Foundation or Practitioner
Must-Have Technical Skills
-Observability SRE: Splunk— SLI/SLO/error budget engineering, OpenTelemetry, ELK/Splunk pipeline reliability, and observability-as-code practices
-IaC Reliability: Terraform — state backend health, drift detection automation, module registry SRE, and policy-as-code pipeline reliability across AWS and GCP
-CI/CD SRE: GitHub Actions, ArgoCD — pipeline health monitoring, runner auto-scaling, deployment success rate SLOs, and automated rollback engineering
-Vault Reliability: HA cluster monitoring, seal/unseal automation, certificate lifecycle management, and lease renewal automation for secrets infrastructure
-LLMOps Reliability: Model API health monitoring (OpenAI, Anthropic, Gemini), prompt execution SLOs, token/cost anomaly alerting, and AI pipeline auto-remediation
- RAG Platform SRE: Vector database availability (Pinecone, Weaviate, ChromaDB), retrieval latency SLOs, embedding pipeline health, and index freshness monitoring
-Automation & Toil Reduction: Python — SRE automation scripting, event-driven remediation, infrastructure SDK integration (boto3, GCP client), and operational workflow engineering
-Incident Management: P1/P2 bridge leadership, blameless post-mortems, structured RCA, error budget reviews, and SLA-governed resolution in FSI environments
-Performance & Capacity: Platform capacity trending, SLO burn rate alerting, cardinality management, and proactive capacity interventions across observability and AI tooling
Preferred / Advantageous
– Experience with chaos engineering or game day exercises for platform tooling resilience — validating failure modes in observability, CI/CD, or AI pipeline infrastructure
– Familiarity with eBPF-based observability (Cilium, Pixie) for deep platform telemetry and service mesh reliability engineering
– Exposure to model serving infrastructure — Triton, vLLM, or similar — for AI/ML platform reliability beyond API-based LLM tooling
– Background in SRE or platform engineering within financial services or other highly regulated industries


Additional Information:
– SLOs for all platform pillars — observability, IaC, CI/CD, ITSM, and AI tooling — are consistently met, with error budgets actively managed and reliability improving measurably quarter-on-quarter
– Toil across the tooling estate decreases consistently — manual intervention patterns are replaced by automated, observable workflows and the team's time shifts toward reliability engineering rather than repetitive operations
– Major platform incidents are managed with clear ownership, rapid mobilisation, blameless RCA outputs, and systemic fixes that prevent recurrence
– LLMOps and AI-augmented operations tooling meets production SLOs — model API failures, agent workflow breakdowns, and RAG pipeline degradation are detected early, remediated automatically where possible, and escalated with full context when not
– Engineering teams across Cloud, Network, Security, Database, and Voice towers rely on platform tooling that is observable, self-healing, and consistently available — the Tools & Platforms SRE is the reason it stays that way
- The candidate should have minimum 5 years of experience in Splunk Enterprise Architecture and Design.
- A 15 years full time education is required.

15 years full time education

About Accenture

Accenture is a leading global professional services company that helps the world’s leading businesses, governments and other organizations build their digital core, optimize their operations, accelerate revenue growth and enhance citizen services—creating tangible value at speed and scale. We are a talent- and innovation-led company with approximately 791,000 people serving clients in more than 120 countries. Technology is at the core of change today, and we are one of the world’s leaders in helping drive that change, with strong ecosystem relationships. We combine our strength in technology and leadership in cloud, data and AI with unmatched industry experience, functional expertise and global delivery capability. Our broad range of services, solutions and assets across Strategy & Consulting, Technology, Operations, Industry X and Song, together with our culture of shared success and commitment to creating 360° value, enable us to help our clients reinvent and build trusted, lasting relationships. We measure our success by the 360° value we create for our clients, each other, our shareholders, partners and communities.

Visit us at www.accenture.com 

Equal Employment Opportunity Statement


We believe that no one should be discriminated against because of their differences. All employment decisions shall be made without regard to age, race, creed, color, religion, sex, national origin, ancestry, disability status, military veteran status, sexual orientation, gender identity or expression, genetic information, marital status, citizenship status or any other basis as protected by applicable law. Our rich diversity makes us more innovative, more competitive, and more creative, which helps us better serve our clients and our communities.

Skills

PythonAWSGCPKubernetesTerraformAnsibleCI/CDGitHubJiraConfluenceServiceNowSplunkDevOpsSREComplianceITIL

Similar Jobs

30

Operations Engineer

Accenturefederalservices · Colorado Springs, CO +1 · Onsite

Today

Operations Engineer

Accenture · Bengaluru, BDC7C, India · Hybrid

Yesterday

Operations Engineer

Accenture · Bengaluru, BDC7C, India · Hybrid

Yesterday

Operations Engineer

Accenture · Bengaluru, BDC7C, India

Yesterday

Operations Engineer

Accenture · Bengaluru, BDC7C, India

Yesterday

Operations Engineer

Accenture · Bengaluru, BDC7C, India

Yesterday

Operations Engineer

Accenture · Bengaluru, BDC7C, India

Yesterday

Operations Engineer

Accenture · Bengaluru, BDC7C, India

Yesterday

Operations Engineer

Accenture · Bengaluru, BDC7C, India · Remote, Hybrid

Yesterday

Operations Engineer

Accenture · Bengaluru, BDC7C, India

Yesterday

Operations Engineer

Accenture · Bengaluru, BDC7C, India

Yesterday

Operations Engineer

Accenture · Bengaluru, BDC7C, India

Yesterday

Operations Engineer

Accenture · Bengaluru, BDC7C, India · Remote, Hybrid

Yesterday

Operations Engineer

Accenture · Bengaluru, BDC7C, India

Yesterday

Operations Engineer

Accenture · Bengaluru, BDC7C, India

Yesterday

Operations Engineer

Accenture · Bengaluru, BDC7C, India

Yesterday

Operations Engineer

Accenture · Bengaluru, BDC7C, India

Yesterday

Operations Engineer

Accenture · Bengaluru, BDC7C, India

Yesterday

Operations Engineer

Accenture · Bengaluru, BDC7C, India

Yesterday

Operations Engineer

Accenture · Bengaluru, BDC7C, India

Yesterday

Operations Engineer

Accenture · Bengaluru, BDC7C, India

Yesterday

Operations Engineer

Accenture · Bengaluru, BDC7C, India · Hybrid

Yesterday

Operations Engineer

Accenture · Bengaluru, BDC7C, India

Yesterday

Operations Engineer

Accenture · Bengaluru, BDC7C, India

Yesterday

Operations Engineer

Accenture · Bengaluru, BDC7C, India · Remote, Hybrid

Yesterday

Operations Engineer

Accenture · Bengaluru, BDC7C, India

Yesterday

Operations Engineer

Accenture · Bengaluru, BDC7C, India

Yesterday

Operations Engineer

Accenture · Bengaluru, BDC7C, India

Yesterday

Operations Engineer

Accenture · Bengaluru, BDC7C, India

Yesterday

Operations Engineer

Accenture · Bengaluru, BDC7C, India

Yesterday
Operations Engineer at Accenture | Hiring.Camp