Hiring.Camp

IND Staff Engineer, Reliability

Thehartford

·

Yesterday

Location
India GCC-Puppalaguda Village
Type
Full-time
Department
Engineering
Seniority
Senior
Experience
8+ years
Education
Bachelor
Closing date
Today
Source
Workday

Description

IND Staff Engineer, Reliability - GCC070

We’re determined to make a difference and are proud to be an insurance company that goes well beyond coverages and policies. Working here means having every opportunity to achieve your goals – and to help others accomplish theirs, too. Join our team as we help shape the future.

Position Summary

We are seeking a highly skilled T7 AI Operations & Site Reliability Engineer to join our engineering team in Hyderabad, India. This role is laser-focused on the availability, reliability, and performance of our production AI systems. You will own the operational health of AI-powered products — ensuring LLM-based services, agentic workflows, RAG pipelines, and ML inference platforms maintain enterprise-grade uptime while scaling to meet demand. You will build the observability, automation, and incident response capabilities that keep our AI products running 24/7.

Level: T7 (Senior Engineer)

Location: Hyderabad, India

Employment Type: Full-Time

Key Responsibilities

AI Platform Reliability & Uptime

  • Own the end-to-end reliability of production AI systems including LLM services, RAG pipelines, agentic workflows, and inference endpoints
  • Define and maintain SLOs/SLIs/SLAs for AI products — latency, availability, error rates, token throughput, and response quality
  • Design and implement high-availability architectures for AI workloads: multi-region failover, load balancing, auto-scaling, and graceful degradation
  • Build circuit breakers, retry logic, fallback models, and rate-limiting strategies to ensure AI services remain available under stress
  • Drive availability targets of 99.9%+ for critical AI-powered products
  • Establish disaster recovery procedures and regularly test backup/restore for AI data stores and model artifacts

Observability & Monitoring

  • Build and maintain comprehensive observability stacks for AI systems — metrics, logs, traces, and AI-specific signals (hallucination rates, model drift, token costs)
  • Implement real-time dashboards and alerting for AI service health, model performance, and infrastructure utilization
  • Design anomaly detection and proactive alerting to identify degradation before users are impacted
  • Monitor LLM provider dependencies (GCP Vertex AI, OpenAI, etc.) and implement automated failover when external services degrade
  • Track and optimize cost-per-inference, token utilization, and resource efficiency across AI workloads

Incident Management & Response

  • Lead incident response for AI system outages and degradations — triage, mitigate, resolve, and communicate
  • Build and maintain runbooks for common AI failure modes: model timeouts, context window overflows, embedding pipeline failures, vector DB issues
  • Establish on-call rotations and escalation procedures tailored to AI system failure patterns
  • Automate incident detection and remediation where possible — self-healing pipelines and auto-rollback

AI Infrastructure & Platform Operations

  • Operate and scale cloud-native AI infrastructure (GCP, AWS) including model serving platforms, Kubernetes containers, and vector databases
  • Implement and maintain Infrastructure-as-Code (Terraform) for AI platform environments
  • Automate deployment pipelines for model updates, configuration changes, and infrastructure scaling

Collaboration & Documentation

  • Partner closely with AI Engineers to ensure new features are built with operability, observability, and reliability in mind
  • Define production-readiness criteria for AI services — ensuring all systems meet reliability standards before launch
  • Maintain comprehensive operational documentation: architecture diagrams, runbooks, playbooks, and SOPs
  • Contribute to architecture reviews with a reliability lens — identifying single points of failure, blast radius, and operational risk
  • Participate in on-call rotations and drive continuous improvement of operational practices

Required Qualifications

  • Experience: 8+ years of professional experience in software engineering, DevOps, or site reliability engineering, with 1+ year operating AI/ML systems in production
  • Education: Bachelor's degree in Computer Science, Software Engineering, or related field (or equivalent experience)
  • SRE Fundamentals:
    • Deep understanding of SRE principles: SLOs, error budgets, toil reduction, incident management, and capacity planning
    • Proven track record maintaining high availability (99.9%+) for production systems at scale
    • Experience integrating into observability platforms (Prometheus, Grafana, Datadog, Splunk, or equivalent)
    • Strong incident response skills with experience leading war rooms and post-incident reviews
  • AI/ML Operations:
    • Understanding of AI-specific failure modes: model drift, hallucination spikes, token limit errors, embedding pipeline failures, and provider outages
    • Familiarity with LLM providers and platforms (GCP Vertex AI, OpenAI, AWS Bedrock) from an operational perspective
  • Cloud & Infrastructure: Advanced-level experience with cloud platforms (GCP, AWS), Kubernetes, containerization, and Infrastructure-as-Code (Terraform)
  • Programming: Strong proficiency in Python and at least one systems language; comfortable writing automation scripts, custom exporters, and operational tooling
  • Networking & Security: Solid understanding of networking, load balancing, DNS, TLS, and security best practices for cloud-native systems
  • CI/CD: Experience building and maintaining deployment pipelines (Jenkins, GitHub Actions, ArgoCD) with automated rollback capabilities
  • Communication: Excellent communication skills for incident coordination, stakeholder updates, and cross-team collaboration

Preferred Qualifications

  • Knowledge of AI cost optimization strategies — model routing, caching, batching, and tiered inference
  • Experience in regulated industries (insurance, finance, healthcare) with compliance and audit requirements
  • Cloud certifications (GCP Professional Cloud Architect, AWS Solutions Architect, CKA/CKAD)
  • Experience with AIOps — using AI/ML to improve operational intelligence and automated remediation

About Us | Our Culture | What It’s Like to Work Here

Skills

PythonAWSGCPKubernetesTerraformJenkinsCI/CDGitHubSplunkDevOpsSRECompliance

Similar Jobs

4

IND Staff Software Engineer

Thehartford · India GCC-Puppalaguda Village

3 months ago

IND Staff Software Engineer

Thehartford · India GCC-Puppalaguda Village

3 months ago

IND Staff Software Engineer

Thehartford · India GCC-Puppalaguda Village · Remote, Hybrid

3 months ago

IND Senior Staff Software Engineer

Thehartford · India GCC-Puppalaguda Village

3 months ago