Enter a job title or keyword

Reliability Engineer


Job Location:

Hong Kong - Hong Kong

Monthly Salary: Not provided by the employer
Posted: 22 August 2026 (7 hours ago)
Application Deadline: 19 November 2026
Vacancies: 1 Vacancy

Job Summary

Job Description:

  • Automate repeatable triage workflows to help first-line teams respond faster and more consistently (e.g. alert enrichment routing correlation and operational runbooks).
  • Identify monitoring/alerting gaps and drive improvements in visibility and alert quality.
  • Track reliability and availability across critical trading applications and their dependencies. Partner with users development teams and IT to pinpoint where service levels are degrading.
  • Triage incoming alerts issues and escalationsassessing impact urgency and ownership.
  • Determine when incident criteria are met declare incidents and act as Incident Commander.
  • Coordinate responders and stakeholders; keep incident calls focused on facts mitigation and recovery.
  • Maintain clear timelines actions and status updates throughout the incident lifecycle.
  • Recover and stabilize systems using approved runbooks. Escalate cleanly through the defined support/development path when the issue exceeds documented recovery steps.
  • Support post-incident review (PIR) follow-ups and recurring issue reviews.
  • Ensure smooth handovers across EMEA AMER and APAC using a single global model: one incident standard and one handover process.

Requirements:

  • Experience in production operations SRE NOC/command center trading operations or a comparable first-line technical roleideally in a trading financial services or other latency-sensitive environment.
  • Strong triage and prioritization skills: you can separate facts from assumptions under pressure and keep the response moving.
  • Clear communication (verbal and written): status updates are understandable to both traders and engineers.
  • Broad technical understanding (not just deep specialist knowledge): enough to collaborate effectively across domains and interfaces.
  • Solid Linux and networking fundamentals plus the ability to quickly interpret alerts logs dashboards and symptoms.
  • Working knowledge of common operational tasks across adjacent teams (application support infrastructure connectivity data).
  • Familiarity with incident and observability tooling (e.g. PagerDuty or equivalent Jira Service Management or equivalent Grafana Prometheus log search).
  • Scripting/automation skills (Python preferred; Bash and Go are a plus) applied to triage enrichment routing and correlation (not product code).
  • Exposure to containerized/cloud-hosted production environments (Kubernetes Docker GCP) is a plus.