- Location
- Hyderabad, IN
- Type
- Full-time
- Department
- Operations
- Experience
- 2+ years
- Closing date
- Today
- Source
- iCIMS
Description
Job Description
Job Purpose
Working for Intercontinental Exchange (ICE) as part of the ICE Data Services Systems Operations team will provide you with the opportunity to experience supporting one of the most famous and widely known, publicly visible companies in the world. The technology we work with day to day is necessarily on the bleeding edge. Our triumphs and shortcomings make the news and give this position a level of excitement and importance not available in a typical operational/support role.
We are seeking a highly motivated DevOps / Site Reliability Engineering (SRE) Support Operations Analyst to join our team. This role is focused on ensuring the reliability, stability, and performance of production systems through proactive monitoring, incident management, and continuous operational improvements.
The ideal candidate will work closely with engineering teams to maintain high system availability, reduce operational toil, and enhance automation across infrastructure and application environments.
The Operations Support Analyst role is part of a highly specialized support organization that is responsible for the daily operations of multiple industry leading trading exchanges, clearing systems, and data distribution services. This position requires identifying, troubleshooting, and resolving both internal system problems, as well as external customer-related IT issues. The role requires a blend of general technical and business knowledge, as well as a comprehensive understanding of DevOps/SRE function. This position is not a Systems Administrator or Network Engineering position; however, prior experience in these areas is desirable.
Responsibilities
- Production Support & Operations
- Monitor system health, application performance, and infrastructure using observability tools
- Respond to alerts, incidents, and service degradation issues in a timely manner
- Perform initial triage, troubleshooting, and escalation for production issues
- Participate in on-call rotations and support 24/7 operations as needed
- Incident & Problem Management
- Lead and/or support incident response and resolution efforts
- Coordinate with cross-functional teams during major incidents
- Conduct root cause analysis (RCA) and ensure corrective actions are implemented
- Track and improve key reliability metrics (e.g., MTTR, incident frequency)
- Automation & Continuous Improvement
- Identify repetitive operational tasks and automate them using scripting/tools
- Contribute to infrastructure-as-code and configuration management practices
- Improve system observability, logging, and monitoring capabilities
- Support continuous improvement initiatives to enhance reliability and efficiency
- System Reliability & Availability
- Ensure services meet defined SLAs, SLOs, and SLIs
- Proactively identify risks and implement preventive measures
- Support capacity planning and performance tuning efforts
- Collaborate with engineering teams to improve resiliency and fault tolerance
- Collaboration & Communication
- Work closely with development, QA, and platform teams
- Document operational procedures, runbooks, and troubleshooting guides
- Provide regular updates on incidents, risks, and operational improvements
Knowledge and Experience
- Bachelor’s degree in computer science, Engineering, or related field (or equivalent experience)
- 2–5 years of experience in DevOps, SRE, or production support roles
- Hands-on experience with Windows, Linux/Unix systems administration
- Strong understanding of incident management and troubleshooting methodologies
- Experience with monitoring tools (e.g. Spectrum, Grafana, Splunk)
- Knowledge of scripting languages (e.g., Python, Bash, or PowerShell)
- Familiarity with cloud platforms such as AWS, Azure.
- Understanding of CI/CD pipelines and deployment processes
Preferred Knowledge and Experience
- Experience with containerization and orchestration tools (Docker, Kubernetes)
- Familiarity with Infrastructure as Code tools (Terraform, CloudFormation)
- Knowledge of networking fundamentals (DNS, TCP/IP, load balancing)
- Experience with version control systems (Git) and Agile methodologies
- Exposure to SRE principles (error budgets, reliability engineering practices)
Key Skills and Competencies
- Strong analytical and problem-solving skills
- Ability to perform under pressure in high-impact situations
- Proactive mindset with a focus on automation and efficiency
- Excellent communication and collaboration skills
- Attention to detail and commitment to quality