- Location
- Bengaluru, India
- Workplace
- Onsite
- Type
- Full-time
- Department
- Engineering
- Seniority
- Manager
- Experience
- 3+ years
- Source
- Workday
Description
Company Description
Vialto Partners is a market leader in Global Mobility Services. Our purpose is to
“Connect the World.” We are unique and the only stand-alone global mobility business.
This presents a rare opportunity for our clients, stakeholders and colleagues.
Our teams help companies streamline and effectively manage their global mobility programs in a cost-efficient and compliant manner. Our services focus on providing cross-border compliance and risk assessment for tax, immigration, business travel, rewards and compensation, and remote work.
Working at Vialto Partners is about getting the chance to be part of a global and dynamic team. Globally, Vialto Partners has over 8,000 staff worldwide and continues to grow. You will work with clients from a range of industries and different geographical locations. We believe in connecting the world and supporting our colleagues to do the same in their careers by undertaking assignments and opportunities globally that broaden their skills and ultimately benefit our clients.
About Vialto Labs (VLabs)
Vialto Labs (VLabs) is responsible for redesigning how work is delivered in the tax and immigration service lines, as well as driving operational efficiency
across Vialto’s functional areas using AI. The team builds and deploys novel AI-enabled solutions that directly improve productivity and increase delivery quality for our clients. VLabs is accountable for rapidly turning innovative experiments into production-ready deliverables at scale and embedding them into day-to-day operations. This team focuses on the highest-impact workflows, creating standardized, repeatable capabilities that can be deployed globally. Operating with a mandate for speed and measurable outcomes, VLabs works alongside service line, product, and platform leaders.
About the Role
We are seeking a skilled Ops / SRE Support Engineer to support and operate cloud-native, AI-enabled applications and platforms in production. This role focuses on ensuring system reliability, observability, incident response, and operational excellence across Azure-based environments. The ideal candidate has hands-on experience in production support, monitoring, incident management, and automation, and is comfortable working with distributed systems, containers, and modern observability tools. You will play a key role in maintaining uptime, reducing operational noise, and
enabling proactive system health through data-driven insights. In addition, you are a strong problem-solver who brings debugging skills and an ability to work in fast-paced,
high-availability environments with good communication and collaboration skills.
Key Responsibilities
Production Support & Operations
• Monitor and support production systems across applications, infrastructure, and cloud services
• Respond to incidents, troubleshoot issues, and drive timely resolution
• Perform root-cause analysis and contribute to post-incident reviews
• Maintain operational runbooks and standard operating procedures
Observability & Monitoring
• Work with logs, metrics, traces, and alerts to ensure system visibility
• Build and maintain dashboards and alerts using modern monitoring tools
• Improve signal quality through alert tuning, deduplication, and noise reduction
• Support Open Telemetry (OTEL) based instrumentation and telemetry
collection from modern monitoring tools like Tempo, Loki, Prometheus, Grafana, etc.
Incident Management & Reliability
• Participate in on-call rotations and incident response processes
• Support triage, escalation, and communication during incidents
• Help improve MTTR (Mean Time to Resolve) and MTTD (Mean Time to Detect)
• Contribute to defining and tracking SLIs, SLOs, and error budgets
Automation & Efficiency
• Develop scripts and automation for operational tasks using Python, Bash, or
similar
• Assist in automating incident response, alert handling, and remediation
workflows
• Support CI/CD pipelines and deployment processes
Cloud & Platform Support
• Operate and support systems running on Microsoft Azure
• Support containerized environments using Docker and Kubernetes
• Work with cloud-native monitoring tools such as Azure Monitor and Log Analytics
AI-Enabled Operations Support
• Assist in using AI-assisted tools for incident investigation and troubleshooting
• Support correlation of events, anomaly detection, and intelligent alerting
• Help maintain telemetry pipelines and operational data flows
Required Qualifications / Background
• 3–6+ years of experience in operations, production support, SRE, or DevOps
roles
• Hands-on experience supporting production systems in cloud environments
(Azure preferred)
• Experience with monitoring and observability tools (logs, metrics, traces,
alerting)
• Strong understanding of incident management and troubleshooting practices
• Experience with scripting or automation (Python, Bash, etc.)
• Familiarity with Docker and Kubernetes
• Experience working with logs, metrics, traces, and alerts
• Experience in incident response, triage, and root-cause analysis
• Experience with monitoring dashboards and alerting systems
• Understanding of SRE principles (SLIs, SLOs, error budgets)
• Ability to troubleshoot distributed systems and service dependencies
Preferred Qualifications / Background
• Experience with: Azure Monitor, Application Insights, Prometheus, Grafana
OpenTelemetry (OTEL) instrumentation ,CI/CD tools such as Azure DevOps or
GitHub Actions
• Familiarity with: AIOps concepts (alert correlation, anomaly detection,
automation)
• Query languages like KQL or SQL
• Experience working in high-scale or enterprise production environments
Additional Information:
• Location: Bangalore
• We are an equal opportunity employer that does not discriminate on the basis of
any legally protected status.
• Please note, AI is used as part of the application process.