- Location
- India, Pune COE
- Workplace
- Hybrid
- Type
- Full-time
- Department
- Management
- Seniority
- Senior
- Closing date
- Today
- Source
- Workday
Description
For more than 40 years, Accelya has been the industry’s partner for change, simplifying airline financial and commercial processes and empowering the air transport community to take better control of the future. Whether partnering with IATA on industry-wide initiatives or enabling digital transformation to simplify airline processes, Accelya drives the airline industry forward and proudly puts control back in the hands of airlines so they can move further, faster.
Critical Incident Manager (CIM)
Location: Pune / Hybrid
Role Purpose
As an Critical Incident Manager (CIM), you will lead the end-to-end management of business-critical and high-visibility production incidents, ensure rapid restoration of services and minimizing customer impact. You will direct incident response activities, coordinate cross-functional teams, and drive effective communication strategies throughout the incident lifecycle.
The successful candidate will be responsible for owning and managing P1/P2 incidents, ensuring timely resolution within SLA commitments, driving structured incident governance, and maintaining alignment between technical teams, business stakeholders, and customers.
This position will be responsible for incident management, stakeholder communication, operational governance, root cause analysis, and continuous service improvement. The focus will be on ensuring operational resilience, maintaining customer trust, improving service reliability, and driving proactive incident prevention initiatives in line with Accelya's global business strategy, values, and mission.
Duties & Responsibilities:
Incident Ownership & Lifecycle Management
Own and manage the lifecycle of critical P1/P2 and high-priority incidents from detection through resolution and closure.
Act as the primary escalation point for major incidents, ensuring timely engagement of appropriate stakeholders and technical teams.
Perform impact assessments, severity classification, and incident prioritization to ensure effective response execution.
Drive production fixes, workarounds, and mitigation plans required to restore services.
Incident Coordination & Bridge Management
Assemble and lead Incident Response Teams (IRT) across Engineering, Cloud Operations, Infrastructure, DevOps, and Support functions.
Initiate, facilitate, and manage incident bridge calls, ensuring clear accountability, ownership, and progress tracking.
Remove operational roadblocks and maintain resolution momentum throughout the incident lifecycle.
Stakeholder & Customer Communication
Provide timely, accurate, and structured communications to customers, leadership teams, and internal stakeholders.
Ensure communication consistency across customer notifications, management updates, bridge calls, and service status channels.
Manage communications for high-impact incidents affecting strategic customers and business-critical operations.
Status Page & Customer Transparency
Own customer-facing status page updates for all major incidents.
Publish clear incident updates outlining customer impact, affected services, mitigation activities, and recovery progress.
Maintain defined communication cadence throughout incident management and restoration activities.
Ensure alignment of messaging across all communication channels.
RCA & Problem Management
Drive Root Cause Analysis (RCA) creation, review, quality assurance, and timely delivery.
Identify recurring issues and collaborate with Product, Engineering, and Operations teams to eliminate underlying causes.
Track corrective and preventive actions through to closure.
Monitoring, Governance & Continuous Improvement
Proactively monitor service health and operational alerts to identify potential service disruptions.
Lead Daily Operations Reviews (DORs), providing updates on incident trends, risks, and operational performance.
Analyze incident patterns and service trends to identify stability improvements and automation opportunities.
Support enhancements in monitoring, alerting, reporting dashboards, and operational tooling.
Maintain audit-ready incident documentation across Zendesk, Jira, RCA repositories, and reporting systems.
Cross-Functional Collaboration
Partner with Engineering, Infrastructure, Cloud Operations, Product, DevOps, and Support teams to improve service reliability and operational efficiency.
Coordinate with external vendors and technology partners during critical incidents when required.
Ensure accountability, ownership, and adherence to escalation processes across all stakeholder groups.
Operational Excellence
Drive accountability throughout incident management, ensuring timely service restoration and closure.
Participate in a 24x7 rotational on-call support model to provide uninterrupted major incident coverage.
Foster a culture of continuous improvement, operational excellence, and customer-first service delivery.
Required Experience & Skills
Proven experience in Major Incident Management, Service Operations, Production Support, IT Service Management, or Site Reliability Operations.
Strong experience managing critical P1/P2 incidents in complex enterprise or SaaS environments.
Demonstrated experience leading cross-functional technical teams during high-pressure incident situations.
Hands-on experience with incident management processes aligned with ITIL best practices.
Experience conducting incident triage, impact assessments, severity classification, and escalation management.
Strong stakeholder management skills with the ability to communicate effectively with senior leadership, customers, and technical teams.
Experience driving Root Cause Analysis (RCA), corrective action tracking, and problem management initiatives.
Knowledge of cloud-based environments, distributed systems, and modern application architectures.
Experience with ticketing and incident management tools such as Jira, Zendesk, ServiceNow, or equivalent platforms.
Strong analytical, organizational, and decision-making skills.
Excellent written and verbal communication skills.
Preferred Experience & Skills
ITIL Foundation or higher certification.
Experience supporting airline, travel technology, SaaS, or mission-critical enterprise platforms.
Knowledge of AWS, Azure, or cloud-native operational environments.
Familiarity with observability and monitoring platforms such as Datadog, New Relic, Grafana, Splunk, or Dynatrace.
Experience driving automation initiatives within operational support environments.
Exposure to SRE, DevOps, Problem Management, and Change Management practices.
Experience creating executive dashboards and operational reporting metrics.
What do we offer?
Open culture and challenging opportunity to satisfy intellectual needs
Flexible working hours
Smart working: hybrid remote/office working environment
Work-life balance
Excellent, dynamic and multicultural environment
About Accelya
Accelya is a leading global software provider to the airline industry, powering 200+ airlines with an open, modular software platform that enables innovative airlines to drive growth, delight their customers and take control of their retailing.
Owned by Vista Equity Partners long-term perennial fund and with 2K+ employees based around 10 global offices, Accelya are trusted by industry leaders to deliver now and deliver for the future.
The company's passenger, cargo, and industry platforms support airline retailing from offer to settlement, both above and below the wing. Accelya is proud to deliver leading-edge technologies to its customers, including through its partnership with AWS and through the pioneering NDC expertise of its Global Product teams.
We are proud to enable innovation-led growth for the airline industry and put control back in the hands of airlines.
For more information, please visit www.accelya.com.
What does the future of the air transport industry look like to you? Whether you’re an industry veteran or someone with experience from other industries, we want to make your ambitions a reality!