Hiring.Camp

Reliability Engineer (Trading Platforms)

Io Tech Solutions Limited

·

Today

Location
Hong Kong
Type
Full-time
Department
Engineering
Closing date
Today
Source
CareersPage

Description

Job Description:

  • Automate repeatable triage workflows to help first-line teams respond faster and more consistently (e.g., alert enrichment, routing, correlation, and operational runbooks).
  • Identify monitoring/alerting gaps and drive improvements in visibility and alert quality.
  • Track reliability and availability across critical trading applications and their dependencies. Partner with users, development teams, and IT to pinpoint where service levels are degrading.
  • Triage incoming alerts, issues, and escalations—assessing impact, urgency, and ownership.
  • Determine when incident criteria are met, declare incidents, and act as Incident Commander.
  • Coordinate responders and stakeholders; keep incident calls focused on facts, mitigation, and recovery.
  • Maintain clear timelines, actions, and status updates throughout the incident lifecycle.
  • Recover and stabilize systems using approved runbooks. Escalate cleanly through the defined support/development path when the issue exceeds documented recovery steps.
  • Support post-incident review (PIR) follow-ups and recurring issue reviews.
  • Ensure smooth handovers across EMEA, AMER, and APAC using a single global model: one incident standard and one handover process.

Requirements:

  • Experience in production operations, SRE, NOC/command center, trading operations, or a comparable first-line technical role—ideally in a trading, financial services, or other latency-sensitive environment.
  • Strong triage and prioritization skills: you can separate facts from assumptions under pressure and keep the response moving.
  • Clear communication (verbal and written): status updates are understandable to both traders and engineers.
  • Broad technical understanding (not just deep specialist knowledge): enough to collaborate effectively across domains and interfaces.
  • Solid Linux and networking fundamentals, plus the ability to quickly interpret alerts, logs, dashboards, and symptoms.
  • Working knowledge of common operational tasks across adjacent teams (application support, infrastructure, connectivity, data).
  • Familiarity with incident and observability tooling (e.g., PagerDuty or equivalent, Jira Service Management or equivalent, Grafana, Prometheus, log search).
  • Scripting/automation skills (Python preferred; Bash and Go are a plus), applied to triage, enrichment, routing, and correlation (not product code).
  • Exposure to containerized/cloud-hosted production environments (Kubernetes, Docker, GCP) is a plus.

Skills

PythonGCPDockerKubernetesLinuxJiraSRE

Similar Jobs

9

Senior Lead Site Reliability Engineer, Electronic Colo Trading

JPMorgan Chase · Tokyo-To, Japan, JP

3 weeks ago

Senior Lead Site Reliability Engineer, Electronic Colo Trading

JP Morgan Chase · Tokyo-To, Japan, JP

3 weeks ago

Site Reliability Engineer - - Electronic Trading Team

LSEG · St. Loui, Missouri, United States of America +1 · onsite

2 months ago

Site Reliability Engineer - Algorithmic Trading

DRW · Chicago +1

3 months ago

Site Reliability Engineer - Algorithmic Trading

DRW · Tel Aviv +1

3 months ago

Lead Site Reliability Engineer, Electronic Trading Services

JPMorgan Chase · Singapore, SG

4 months ago

Lead Site Reliability Engineer, Electronic Trading Services

JP Morgan Chase · Singapore, SG

4 months ago

Trading Systems Reliability Engineer

Quberesearchandtechnologies · London +1

11 months ago

Trading Systems Reliability Engineer (C++)

Quberesearchandtechnologies · Hong Kong +1

1+ year ago