- Location
- Hong Kong
- Type
- Full-time
- Department
- Engineering
- Closing date
- Today
- Source
- CareersPage
Description
The Opportunity
We are seeking a world-class Reliability Engineer to join a premier High-Frequency Trading (HFT) firm. In our world, we measure success in microseconds and nanoseconds. Downtime isn't just a ticket—it's a direct, measurable hit to P&L by the minute.
This is not a conventional "keeping the lights on" role. You will sit shoulder-to-shoulder with traders, quantitative researchers, and core systems engineers, acting as the critical linchpin that keeps the global trading engine firing on all cylinders. You won't just react to problems; you will actively engineer resiliency into the fabric of one of the fastest trading environments on the planet.
Why You'll Love This Role
- Massive P&L Impact: Your decisions directly protect (and unlock) millions in daily revenue. Every second of uptime you preserve is a tangible win for the firm.
- Elite Compensation: We pay at the top of the market to attract the best. Your base salary and performance-based bonuses reflect the critical nature of this role.
- Unmatched Autonomy: You own the room. As Incident Commander, your decisions hold authority—even when the call is filled with senior engineers, quants, or managing directors. You coordinate, delegate, and dictate the strategy.
- Cutting-Edge Complexity: Manage ultra-low-latency architectures, globally distributed Kubernetes clusters, and highly advanced observability stacks at a scale and speed that few firms can match.
- Zero Bureaucracy: We operate a flat structure. You have the standing to push back on development teams, infrastructure leads, or traders when operational standards slip.
What You Will Do
Proactive Resilience:
- Automate the repetitive parts of triage (alert enrichment, routing, and correlation) so your first-line responders are 10x faster.
- Obsess over monitoring gaps. If it can't be observed, it can't be traded. You will define service levels and push teams to meet rigorous SLAs.
Command the Response:
- Take full control when things break. You assess the impact, assemble the right responders, and run the entire incident lifecycle under our Global Incident Management framework.
- Use your deep technical breadth (Linux, Networking, Logs) to read symptoms instantly, stabilize systems using runbooks, and escalate cleanly when issues exceed documented steps.
Global Ownership:
- Seamlessly hand over between EMEA, AMER, and APAC under one unified incident standard. You are part of a 24/7 elite global force.
What You Need to Succeed
- Proven Experience: Background in Production Operations, SRE, NOC/Command Centre, or Trading Operations—ideally within HFT, financial services, or other extreme latency-sensitive environments.
- Command Presence: A track record of coordinating major incidents. You aren't afraid to take the microphone and guide a room of senior stakeholders toward resolution.
- Elite Triage Skills: You cut through assumptions under pressure. You know when to push forward and exactly when to pull in a specialist.
- Technical Breadth (Not Just Depth): You are dangerous enough across all domains (Apps, Infrastructure, Data, Connectivity) to be useful everywhere.
- Solid Fundamentals: Strong Linux and networking knowledge. You can read a dashboard, parse a log file, and spot the anomaly in seconds.
- Tooling Mastery: Hands-on with PagerDuty, Jira Service Management, Grafana, Prometheus, and log aggregation tools.
- Automation Mindset: Scripting proficiency (Python preferred; Bash/Go are a bonus) applied to operational workflows—not just product code.
- Bonus: Exposure to containerized, cloud-hosted, and bare-metal production systems (Kubernetes, Docker, GCP) is highly desirable.