- Location
- Chennai Embassy Tower Office, India · HYDERABAD, IND
- Workplace
- Remote, Hybrid
- Type
- Full-time
- Department
- IT
- Seniority
- Manager
- Experience
- 10+ years
- Source
- Workday
Description
About NCR VOYIX
NCR Voyix Corporation (NYSE: VYX) is a global platform-powered leader in unified commerce for shopping and dining. Combining a flexible, intelligent platform with end-to-end payments capabilities and services developed through its deep industry experience, NCR Voyix empowers retailers and restaurants to accelerate new possibilities for their operations, experiences and business outcomes. NCR Voyix is headquartered in Atlanta, Georgia, and serves customers in more than 35 countries worldwide.
Position Overview
We are seeking an experienced Observability Platform Manager to lead the strategy, implementation, and continuous evolution ofenterprise observability capabilities across cloud and hybrid environments. The ideal candidate has recent hands-on experience implementing observability tools at scale, a strong foundation in cloud engineering, and the leadership skills to drive platform adoption, standardization, and operational excellence across engineering teams. Familiarity with AI applications and AI-enabled observability capabilities is a strong plus.
Key Responsibilities
Platform Strategy & Leadership
- Define and execute the vision, roadmap, and operating model for the enterprise observability platform, including logs, metrics, traces, dashboards, and alerting.
- Lead and mentor engineers responsible for building, operating, and improving observability capabilities across business-critical platforms and applications.
- Establish platform standards, governance, onboarding patterns, and success measures to drive consistent adoption at scale.
- Partner with SRE, platform engineering, cloud engineering, infrastructure, and application teams to align observability strategy with reliability and business goals.
Observability Platform Implementation
- Lead recent and large-scale implementations of observability tools and platforms across multi-team or enterprise environments.
- Evaluate, select, and optimize observability tooling such as Splunk, Datadog, Dynatrace, Grafana, Prometheus, Elasticsearch, or equivalent solutions based on scale, cost, and business needs.
- Drive implementation of telemetry standards, ingestion pipelines, access models, dashboard frameworks, and alerting practices that improve signal quality and reduce operational noise.
- Manage vendor relationships, platform lifecycle decisions, and cost/performance tradeoffs for observability capabilities.
Cloud Engineering & Reliability
- Bring strong cloud engineering experience across AWS, Azure, or GCP, with an understanding of cloud-native architectures, resilience patterns, and scalable operations.
- Guide teams on instrumentation and monitoring for microservices, containers, Kubernetes, serverless workloads, and distributed systems using modern telemetry standards such as OpenTelemetry.
- Partner with engineering teams to embed observability into platform design, CI/CD pipelines, incident response, and service reliability practices.
Operational Excellence & Enablement
- Establish service-level indicators, objectives, reporting, and operational health reviews to improve platform reliability and engineering outcomes.
- Develop enablement programs, best practices, and self-service patterns that help teams adopt observability consistently and effectively.
- Drive continuous improvement in incident detection, triage, root-cause analysis, and post-incident learning through better telemetry and platform workflows.
Automation & AI Applications
- Promote automation-first practices for instrumentation, alert tuning, dashboard provisioning, and operational workflows.
- Identify opportunities to apply AI-enabled capabilities such as anomaly detection, event correlation, intelligent alerting, and operational insights.
- Familiarity with AI applications, AI platforms, or AI-supported engineering workflows is a plus.
Required Qualifications
- 10+ years of experience in observability, SRE, platform engineering, cloud engineering, or related disciplines, including recent experience implementing observability tools at scale.
- Proven experience leading or managing observability platforms, programs, or engineering teams in complex enterprise environments.
- Hands-on experience with enterprise observability and monitoring tools such as Splunk, Datadog, Dynatrace, Grafana, Prometheus, Elastic, or comparable platforms.
- Strong cloud engineering experience with AWS, Azure, or GCP, including cloud-native services, architecture patterns, and operational best practices.
- Knowledge of distributed systems, microservices, containers, Kubernetes, CI/CD, and infrastructure-as-code practices.
- Familiarity with OpenTelemetry, distributed tracing, service health models, and telemetry data management.
- Ability to define meaningful KPIs, SLIs/SLOs, alerting strategies, and executive-ready reporting that tie technical health to business outcomes.
- Strong collaboration and communication skills, with the ability to influence stakeholders across engineering, operations, and leadership teams.
Preferred Qualifications
- Familiarity with AI applications, AI engineering workflows, or AI-enhanced observability capabilities.
- Knowledge of container ecosystems and orchestration platforms (Kubernetes, AKS/EKS/GKE).
- Experience working with event-driven architectures and microservices environments.
- Strong scripting or programming skills (Python, PowerShell, Bash, etc.).
- Relevant certifications (e.g., Splunk Architect, Dynatrace Professional, Cloud certifications).
Soft Skills
- Excellent communication and stakeholder management skills.
- Ability to lead technical strategy and influence architectural decisions.
- Strong analytical, troubleshooting, and problem-solving abilities.
- Adaptability and curiosity about new technologies and evolving observability trends.
Offers of employment are conditional upon passage of screening criteria applicable to the job
EEO Statement
Integrated into our shared values is NCR Voyix’s commitment to equal employment opportunity. All qualified applicants will receive consideration for employment without regard to sex, age, race, color, creed, religion, national origin, disability, sexual orientation, gender identity, veteran status, military service, genetic information, or any other characteristic or conduct protected by law. NCR Voyix is committed to being a globally inclusive company where all people are treated fairly, recognized for their individuality, promoted based on performance and encouraged to strive to reach their full potential. We believe in understanding and respecting differences among all people. Every individual at NCR Voyix has an ongoing responsibility to respect and support a globally diverse environment.
Statement to Third Party Agencies
To ALL recruitment agencies: NCR Voyix only accepts resumes from agencies on the preferred supplier list. Please do not forward resumes to our applicant tracking system, NCR Voyix employees, or any NCR Voyix facility. NCR Voyix is not responsible for any fees or charges associated with unsolicited resumes
“When applying for a job, please make sure to only open emails that you will receive during your application process that come from a @ncrvoyix.com email domain.”