- Location
- Dublin, Ireland
- Type
- Full-time
- Source
- Workday
Description
We are seeking an experienced Incident Response Analyst to strengthen our global Intelligent Operations Centre and provide reliable frontline support across incident monitoring, triage, response and service reliability.
The role will be responsible for monitoring infrastructure and service alerts, identifying potential operational risks, supporting incident response and change activities, coordinating with technical teams, and contributing to root cause analysis and continuous service improvement.
The successful candidate will be comfortable working in a structured production environment, managing operational issues independently and collaborating with multiple technical and business stakeholders to protect service availability and minimise disruption.
Key Responsibilities
Incident Monitoring & Response
Monitor server, infrastructure and service-related alerts and events.
Perform initial triage, assessment and prioritisation of incidents.
Respond to incidents in accordance with defined operational procedures and service levels.
Identify abnormal behaviour, potential service degradation and emerging operational risks.
Escalate incidents to the appropriate technical or operational teams based on severity and ownership.
Support the coordination of incidents across multiple technical teams and stakeholders.
Maintain accurate and timely incident records, including actions, decisions, escalations and outcomes.
Support major incident response activities when required.
Change & Operational Monitoring
Monitor infrastructure and server changes to identify potential risks or unexpected impacts.
Support technical teams during planned change windows and operational events.
Identify and escalate abnormal conditions arising during or following changes.
Assist with cross-functional coordination where changes may have a wider or horizontal impact.
Contribute to post-change reviews and lessons learned where required.
Root Cause Analysis & Service Reliability
Support structured investigation of service interruptions, outages and recurring incidents.
Conduct preliminary root cause analysis and gather relevant technical evidence.
Work with engineering and infrastructure teams to identify contributing factors and corrective actions.
Document findings and support the production of incident and RCA reports.
Track follow-up actions through to completion.
Identify recurring issues and opportunities to improve service reliability.
Governance & Continuous Improvement
Participate in operational governance activities and service reliability initiatives.
Maintain and improve SOPs, MOPs, operational procedures and knowledge documentation.
Contribute to incident management process improvements and operational standardisation.
Identify opportunities to improve monitoring, alerting, escalation and response processes.
Support service improvement initiatives based on incident trends and operational data.
Ensure operational activities are performed in accordance with established policies and procedures.
On-Call & Operational Coverage
Participate in an operational on-call rota as required.
Provide frontline support during out-of-hours operational events.
Support BMC/on-call activities and downtime analysis where applicable.
Participate in structured handovers between regional teams to ensure continuity of operational coverage.
Candidate Profile
Required Experience
2–5 years of relevant experience in incident management, infrastructure operations, technical support, service operations, monitoring or a similar function.
Practical experience in incident triage, escalation and resolution within a production environment.
Good understanding of server operations, infrastructure monitoring and alert management.
Experience working with structured incident management and escalation processes.
Experience supporting change activities and identifying operational risks.
Ability to conduct structured investigations and contribute to root cause analysis.
Strong analytical and problem-solving skills.
Ability to work independently and make appropriate operational decisions within defined processes.
Strong written and verbal communication skills.
Strong documentation and reporting capabilities.
Ability to collaborate effectively with technical teams across different functions and regions.
Willingness to participate in on-call and operational support rotations.
Preferred Qualifications
Experience with infrastructure monitoring, observability or alerting platforms.
Familiarity with ITIL or formal incident management frameworks.
Experience supporting major incidents or high-severity operational events.
Knowledge of server hardware, operating systems and infrastructure environments.
Experience with change management and operational risk management.
Exposure to service reliability, operational governance or continuous improvement initiatives.
Previous experience with BMC-related operational support.
Experience working within a global or follow-the-sun operations model.
Mandarin language capability is advantageous, particularly for China-region responsibilities.
Key Skills
Incident Management
Incident Triage & Response
Infrastructure Monitoring
Alert Management
Root Cause Analysis
Change Monitoring
Service Reliability
Operational Governance
Problem Solving
Technical Investigation
Cross-Functional Coordination
Documentation & Reporting
Escalation Management
Continuous Improvement
Working Environment & Coverage
This is a global operational role supporting regional Intelligent Operations Centre teams across Singapore, Dublin and San Jose.
The operating model provides continuous regional coverage, with each region working a 9-hour shift comprising 8 hours of work and a 1-hour lunch break, with approximately one hour of overlap between regions to facilitate operational continuity and handover.
What We Are Looking For
We are looking for an operationally minded Incident Response Analyst who can combine technical awareness, structured incident management and strong communication skills.
The ideal candidate will be comfortable working under pressure, able to quickly assess and prioritise operational issues, and capable of coordinating effectively with multiple technical teams. They should also have a continuous-improvement mindset and be motivated to strengthen monitoring, incident response and service reliability processes.
Salary Range
45,280.00 - 56,600.00 EUR (Annual)- Please note that the salary information provided herein is base pay only (gross); it does not include other forms of compensation which may or may not apply to this specific position, namely, performance-based bonuses, benefits-related payments, or other general incentives - none of which are guaranteed, may be subject to specific eligibility requirements, and are wholly within the discretion of Astreya to remit.
- Further, the salary information noted above is a range that consists of a minimum and maximum rate of pay for this specific position. Where an applicant or employee is placed on this range will depend and be contingent on objective, documented work-related considerations like education, experience, certifications, licenses, preferred qualifications, among other factors.