- Location
- Hyderabad, India
- Type
- Full-time
- Department
- Engineering
- Experience
- 6+ years
- Education
- Bachelor
- Source
- Workday
Description
About Rimini Street, Inc.
Rimini Street, Inc. (Nasdaq: RMNI), a Russell 2000® Company, is a proven, trusted global provider of end-to-end, mission-critical enterprise software support, managed services and innovative Agentic AI ERP solutions, and is the leading third-party support provider for Oracle, SAP and VMware software.
Our comprehensive portfolio of unified solutions help run, manage, support, customize, configure, connect, protect, monitor, and optimize enterprise application, database and technology software, enabling our clients to achieve better business outcomes, significantly reduce costs and reallocate resources towards strategic projects.
The Company has signed thousands of contracts with Fortune Global 100, Fortune 500, midmarket, public sector and government organizations who selected Rimini Street as their trusted, proven mission-critical enterprise software solutions provider and achieved better operational outcomes, realized billions of US dollars in savings and funded AI and other innovation investments.
We are actively seeking a Sr. Site Reliability Engineer. This role is based in India, Hyderabad.
About Rimini Street, India, GCC.
Rimini Street Inc, HQ : Las Vegas, NV, USA a disruptor in third party ERP support services, established undisputed leadership and as a natural progression, entered India with Rimini Street, India GCC India kick starting operations in Hyderabad, in 2013 with Global Client Onboarding Services, IT shared services and Global Service Development. In no time, Rimini Street, India GCC started Bengaluru operations going up the value chain with more complex product development (Oracle, SAP, Peoplesoft, JDE etc.) & advanced services (Managed services, Professional services, Security Managed Services etc).
Rimini Street, India GCC gained valuable share in bringing the reputation to Rimini Street Inc of being a global provider of unified support and managed service solutions for enterprise software. Today, Rimini Street, India GCC is a family of about 800+ full time talented individuals, thanks to the remarkable talent that has supported the expansion.
Rimini Street, India has nicely emerged as Global Capability Centre (GCC), and proudly says, “if you are best of the best, you belong to Rimini”. We are on a mission to contribute significantly to our “Rimini ONE” program, a turnkey Rimini Street service program that offers a comprehensive set of unified, integrated services that can run, manage, support, customize, configure, connect, protect, monitor, and optimize your Oracle and SAP ERP, database, and technology software.
Position Summary
The Senior Site Reliability Engineer is responsible for the reliability, availability, performance, observability and operational excellence of Rimini Street's Agentic AI ERP Platform.
The role covers production operations, incident management, resilience engineering, capacity planning and disaster recovery across Kubernetes infrastructure, AI Agents, MCP Servers, AI/LLM Gateways, Temporal workflows, RAG services and related enterprise AI platform components.
The SRE works closely with Platform Engineering, AI Engineering, Product Engineering, DevOps and Security teams to ensure that platform services are scalable, secure, supportable and ready for enterprise production use across cloud-hosted, customer-hosted and air-gapped environments.
Key Responsibilities
Reliability and Production Operations
- Own service reliability, availability and operational health across the Agentic AI Platform.
- Define, measure and maintain Service Level Indicators, Service Level Objectives and Error Budgets.
- Conduct production-readiness and reliability reviews for new services and major platform changes.
- Identify recurring operational risks and drive engineering improvements that prevent recurrence.
Observability and Incident Management
- Build and maintain monitoring, logging, distributed tracing, alerting and service-health dashboards.
- Lead or support production incident response, service restoration, Root Cause Analysis and corrective actions.
- Develop actionable alerts, operational runbooks and escalation procedures.
- Track reliability trends, service degradation and opportunities to reduce detection and recovery time.
Kubernetes, Cloud and Resilience
- Operate and support production Kubernetes environments, including deployments, upgrades, scaling and troubleshooting.
- Improve high availability, fault tolerance, backup, recovery and disaster-recovery capabilities.
- Perform capacity planning, performance analysis and infrastructure optimization.
- Support cloud-hosted, customer-hosted and air-gapped deployment models.
Agentic AI Platform Reliability
- Operate and monitor AI Agents, Agentic Workflows, MCP Servers, AI/LLM Gateways, Temporal workflows and RAG services.
- Improve latency, throughput, scalability, resilience and recovery across AI platform components.
- Monitor gateway routing, authentication, quotas, rate limits, failover and service health.
- Build automation and self-healing capabilities to reduce operational toil and improve platform stability.
Mandatory Skills
- Production Kubernetes operations using Helm, Docker and container orchestration.
- SRE practices, including SLIs, SLOs, Error Budgets, production-readiness reviews and reliability engineering.
- Production incident management, Root Cause Analysis, corrective actions and recurrence prevention.
- Observability using Prometheus, Grafana, OpenTelemetry, ELK, OpenSearch or equivalent tools.
- Infrastructure as Code using Terraform, Pulumi or equivalent tooling.
- Production operations on at least one major cloud platform: Azure, AWS or GCP.
- Operational automation and scripting using Python, Go, Bash or equivalent.
- High availability, scalability, capacity planning, backup and disaster recovery.
- Secrets management, certificates, access controls and operational security.
- Hands-on experience operating or supporting AI Agents and Agentic Workflows.
- Hands-on experience deploying, operating or supporting MCP Servers and MCP-based integrations.
- Hands-on experience operating or supporting AI/LLM Gateways such as LiteLLM, Azure AI Gateway, Azure API Management, Kong AI Gateway or equivalent.
- Experience supporting GenAI, RAG or LLM-powered applications in production-oriented environments.
- Strong troubleshooting skills across infrastructure, application and AI platform layers.
Preferred Skills
- Temporal or another durable workflow orchestration platform.
- GitOps practices using Argo CD, Flux or equivalent.
- AI observability tools such as Langfuse, LangSmith, OpenLIT or Arize.
- Vector databases such as Qdrant, Pinecone, Weaviate, Chroma or pgvector.
- Service mesh technologies such as Istio or Linkerd.
- PostgreSQL or other production database operations.
- Multi-region, customer-hosted or air-gapped deployment experience.
- Experience supporting enterprise SaaS or other business-critical platforms.
Experience and Qualifications
- 6 to 10 years of experience in Site Reliability Engineering, Platform Operations, DevOps, Cloud Engineering or Infrastructure Engineering.
- Demonstrated experience operating business-critical production environments.
- Experience supporting Kubernetes-based platforms at scale.
- Bachelor's degree in Computer Science, Engineering or a related field.
- CKA, CKAD, cloud platform or SRE-related certifications are advantageous.
Skills and Competencies
- Strong production ownership and a disciplined approach to reliability and recovery.
- Ability to troubleshoot across infrastructure, applications, workflows and AI services.
- Ability to convert recurring incidents into automation, platform improvements and clear runbooks.
- Clear written and verbal communication across distributed, multi-time zone teams.
- Ability to balance reliability, security, performance, scalability and customer impact.
Ideal Candidate
The ideal candidate is a hands-on SRE with strong Kubernetes and production operations experience who understands the operational characteristics of Agentic AI systems. They can reliably operate AI Agents, MCP Servers, AI/LLM Gateways and supporting services while improving observability, resilience, performance, scalability and incident response across enterprise deployment environments.
Why Rimini Street?
We are looking for talented, passionate people to help us build our future at Rimini Street. We hire only the best, the most extraordinary professionals and provide compensation, bonuses, and benefits to match the skills of our top-performing team members. Do you thrive in a fast-paced environment, enjoy growing together, and get excited about learning new skills? Are you looking for an opportunity to make a true impact as part of a team of extraordinary professionals? This is the place for you.
Our work is challenging and meaningful. We start and end each day with a sense of achievement and purpose guided by our core values, the Four Cs:
- Company
- We dream big and innovate boldly.
- Colleagues
- We work with extraordinary people who create a culture of mutual respect and collaboration.
- Clients
- We relentlessly pursue solutions that help clients achieve their goals. Our unmatched client care is rooted in our passion for exceptional service.
- Community
- We believe in leaving the world a better place than we found it. With the Rimini Street Foundation, we’ve made positive impacts in six continents for over 425 charities.
Accelerating Company Growth
- Nasdaq-listed under ticker symbol RMNI since October 2017
- Over 6,300+ signed contracts to date, including Fortune 500 and Global 100 companies
- Over 2,000 team members in 23 countries
- US and international recognition for industry leadership and philanthropic efforts. See all of our awards and recognitions here: https://www.riministreet.com/company/awards/
Rimini Street is committed to creating a diverse and inclusive environment and is proud to be an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to age, race, color, religion, national origin, sexual orientation, gender or gender identity, disability, protected veteran status, or any other characteristic protected by law.
To learn more about how Rimini Street is redefining the enterprise software support industry, visit http://www.riministreet.com
Please Note: Rimini Street does not accept resumes submitted by recruiting/staffing firms unless specifically requested by Human Resources. Unsolicited resumes will be ineligible for referral fees.