Enter a job title or keyword

Application Reliability Engineer


Job Location:

Hong Kong - Hong Kong

Monthly Salary: Not provided by the employer
Posted: 22 August 2026 (21 hours ago)
Application Deadline: 19 November 2026
Vacancies: 1 Vacancy

Job Summary


The Opportunity

We are seeking a world-class Reliability Engineer to join a premier High-Frequency Trading (HFT) our world we measure success in microseconds and nanoseconds. Downtime isnt just a ticketits a direct measurable hit to P&L by the minute.

This is not a conventional keeping the lights on role. You will sit shoulder-to-shoulder with traders quantitative researchers and core systems engineers acting as the critical linchpin that keeps the global trading engine firing on all cylinders. You wont just react to problems; you will actively engineer resiliency into the fabric of one of the fastest trading environments on the planet.

Why Youll Love This Role
  • Massive P&L Impact: Your decisions directly protect (and unlock) millions in daily revenue. Every second of uptime you preserve is a tangible win for the firm.
  • Elite Compensation: We pay at the top of the market to attract the best. Your base salary and performance-based bonuses reflect the critical nature of this role.
  • Unmatched Autonomy: You own the room. As Incident Commander your decisions hold authorityeven when the call is filled with senior engineers quants or managing directors. You coordinate delegate and dictate the strategy.
  • Cutting-Edge Complexity: Manage ultra-low-latency architectures globally distributed Kubernetes clusters and highly advanced observability stacks at a scale and speed that few firms can match.
  • Zero Bureaucracy: We operate a flat structure. You have the standing to push back on development teams infrastructure leads or traders when operational standards slip.
What You Will Do

Proactive Resilience:

  • Automate the repetitive parts of triage (alert enrichment routing and correlation) so your first-line responders are 10x faster.
  • Obsess over monitoring gaps. If it cant be observed it cant be traded. You will define service levels and push teams to meet rigorous SLAs.

Command the Response:

  • Take full control when things break. You assess the impact assemble the right responders and run the entire incident lifecycle under our Global Incident Management framework.
  • Use your deep technical breadth (Linux Networking Logs) to read symptoms instantly stabilize systems using runbooks and escalate cleanly when issues exceed documented steps.

Global Ownership:

  • Seamlessly hand over between EMEA AMER and APAC under one unified incident standard. You are part of a 24/7 elite global force.
What You Need to Succeed
  • Proven Experience: Background in Production Operations SRE NOC/Command Centre or Trading Operationsideally within HFT financial services or other extreme latency-sensitive environments.
  • Command Presence: A track record of coordinating major incidents. You arent afraid to take the microphone and guide a room of senior stakeholders toward resolution.
  • Elite Triage Skills: You cut through assumptions under pressure. You know when to push forward and exactly when to pull in a specialist.
  • Technical Breadth (Not Just Depth): You are dangerous enough across all domains (Apps Infrastructure Data Connectivity) to be useful everywhere.
  • Solid Fundamentals: Strong Linux and networking knowledge. You can read a dashboard parse a log file and spot the anomaly in seconds.
  • Tooling Mastery: Hands-on with PagerDuty Jira Service Management Grafana Prometheus and log aggregation tools.
  • Automation Mindset: Scripting proficiency (Python preferred; Bash/Go are a bonus) applied to operational workflowsnot just product code.
  • Bonus: Exposure to containerized cloud-hosted and bare-metal production systems (Kubernetes Docker GCP) is highly desirable.