Site Reliability Engineer – SRE Data & Platform
Posted:
8 October 2026 (Yesterday)
Application Deadline:
5 January 2027
Vacancies:
1 Vacancy
Job Summary
We are looking for a Site Reliability Engineer with strong production reliability and observability experience together with hands-on exposure to data pipelines streaming or data platforms.
You will work in a fast-paced technology environment supporting high-performance production systems with a focus on reliability monitoring automation and system performance.
Key Responsibilities
- Own and improve the reliability availability and performance of production systems.
- Build and enhance observability monitoring dashboards and alerting for critical services.
- Troubleshoot production issues perform root-cause analysis and drive reliability improvements.
- Support Kubernetes/cloud-based infrastructure and production workloads.
- Develop and maintain CI/CD and GitOps workflows.
- Work with data pipelines streaming workloads or data platform infrastructure to support reliable and scalable data processing.
- Monitor data flow system performance and potential issues such as latency failures capacity constraints and data loss.
- Develop automation and operational tools using Python or other programming/scripting languages.
- Participate in production support and on-call activities.
Requirements
- 37 years of experience in SRE DevOps Platform Engineering Production Engineering or a related role.
- Strong hands-on experience with production systems monitoring/observability and incident troubleshooting.
- Must have practical exposure to data-related infrastructure such as:
- Data pipelines / data ingestion
- Streaming platforms
- ETL / ELT
- Kafka / Kinesis or similar technologies
- Spark / Flink or similar processing frameworks
- Data platforms / data lakes / OLAP databases
- Experience with Kubernetes and cloud environments such as AWS GCP or Azure.
- Experience with CI/CD GitOps or infrastructure automation.
- Proficiency in Python or another programming/scripting language.
- Good understanding of system performance reliability and troubleshooting.
- Strong communication and problem-solving skills.
- Hong Kong-based candidates are highly preferred.