Hiring.Camp

Application Reliability Engineer

Io Tech Solutions Limited

·

Today

Location
Hong Kong
Type
Full-time
Department
Engineering
Closing date
Today
Source
CareersPage

Description


The Opportunity

We are seeking a world-class Reliability Engineer to join a premier High-Frequency Trading (HFT) firm. In our world, we measure success in microseconds and nanoseconds. Downtime isn't just a ticket—it's a direct, measurable hit to P&L by the minute.

This is not a conventional "keeping the lights on" role. You will sit shoulder-to-shoulder with traders, quantitative researchers, and core systems engineers, acting as the critical linchpin that keeps the global trading engine firing on all cylinders. You won't just react to problems; you will actively engineer resiliency into the fabric of one of the fastest trading environments on the planet.

Why You'll Love This Role

  • Massive P&L Impact: Your decisions directly protect (and unlock) millions in daily revenue. Every second of uptime you preserve is a tangible win for the firm.
  • Elite Compensation: We pay at the top of the market to attract the best. Your base salary and performance-based bonuses reflect the critical nature of this role.
  • Unmatched Autonomy: You own the room. As Incident Commander, your decisions hold authority—even when the call is filled with senior engineers, quants, or managing directors. You coordinate, delegate, and dictate the strategy.
  • Cutting-Edge Complexity: Manage ultra-low-latency architectures, globally distributed Kubernetes clusters, and highly advanced observability stacks at a scale and speed that few firms can match.
  • Zero Bureaucracy: We operate a flat structure. You have the standing to push back on development teams, infrastructure leads, or traders when operational standards slip.

What You Will Do

Proactive Resilience:

  • Automate the repetitive parts of triage (alert enrichment, routing, and correlation) so your first-line responders are 10x faster.
  • Obsess over monitoring gaps. If it can't be observed, it can't be traded. You will define service levels and push teams to meet rigorous SLAs.

Command the Response:

  • Take full control when things break. You assess the impact, assemble the right responders, and run the entire incident lifecycle under our Global Incident Management framework.
  • Use your deep technical breadth (Linux, Networking, Logs) to read symptoms instantly, stabilize systems using runbooks, and escalate cleanly when issues exceed documented steps.

Global Ownership:

  • Seamlessly hand over between EMEA, AMER, and APAC under one unified incident standard. You are part of a 24/7 elite global force.

What You Need to Succeed

  • Proven Experience: Background in Production Operations, SRE, NOC/Command Centre, or Trading Operations—ideally within HFT, financial services, or other extreme latency-sensitive environments.
  • Command Presence: A track record of coordinating major incidents. You aren't afraid to take the microphone and guide a room of senior stakeholders toward resolution.
  • Elite Triage Skills: You cut through assumptions under pressure. You know when to push forward and exactly when to pull in a specialist.
  • Technical Breadth (Not Just Depth): You are dangerous enough across all domains (Apps, Infrastructure, Data, Connectivity) to be useful everywhere.
  • Solid Fundamentals: Strong Linux and networking knowledge. You can read a dashboard, parse a log file, and spot the anomaly in seconds.
  • Tooling Mastery: Hands-on with PagerDuty, Jira Service Management, Grafana, Prometheus, and log aggregation tools.
  • Automation Mindset: Scripting proficiency (Python preferred; Bash/Go are a bonus) applied to operational workflows—not just product code.
  • Bonus: Exposure to containerized, cloud-hosted, and bare-metal production systems (Kubernetes, Docker, GCP) is highly desirable.


Skills

PythonGCPDockerKubernetesLinuxJiraSRE

Similar Jobs

14

Application Reliability Engineer

Graviton Research Capital LLP · Gurugram, Haryana, India

1 week ago

Production Engineer, Site Reliability (Application Software)

Spacex · Hawthorne, CA +1

3 weeks ago

Software Engineer, Site Reliability Engineering (Application Software)

Spacex · Hawthorne, CA +1

3 weeks ago

Site Reliability Engineer (Application Software)

Spacex · Hawthorne, CA +1

3 weeks ago

Application & System Reliability Engineer

Eaton · Coraopolis, PA,US, US

4 weeks ago

Senior Site Reliability Engineer (Application)

Guidewire · Kuala Lumpur Office, Malaysia · Hybrid

1 month ago

Senior Application Support Engineer / Site Reliability Engineer (SRE)

Depository Trust Company · Boston, MA, United States, US

1 month ago

Sr. Site Reliability Engineer (Application Software)

Spacex · Hawthorne, CA +1

1 month ago

Senior Site Reliability Engineer (Application Support)

Depository Trust Company · Jersey City, NJ, United States, US

2 months ago

Senior Application Manager, Site Reliability Engineer (w/m/div.)

Robert Bosch · Dresden, SN, Germany

2 months ago

Site Reliability Engineer III (SRE) - Guidewire Cloud Platform (Application)

Guidewire · Dublin Office, Ireland · Hybrid

3 months ago

Site Reliability Engineer II (SRE) - Guidewire Cloud Platform (Application)

Guidewire · Krakow Office, Poland · Hybrid

5 months ago

Application Support Engineer, Service Reliability Engineering

Ciena · Remote

7 months ago

Lead, Site Reliability Engineering (Application Support)

Omers · Head Office Toronto, Canada

3 weeks ago