Lead Engineer – Reliability & Observability
Job Summary
We are looking for a Lead Engineer to own and significantly advance our Reliability & Observability platform. We already have an established observability stack including ELK Grafana Loki and monitoring dashboards. As our multi-tenant platform scales across multiple deployment sites and environments you will drive the next phase of its evolution making it easier for engineering teams to understand monitor and improve the reliability of the systems they build.
This is a hands-on engineering leadership role. You will own and evolve our observability stack establish reliability engineering standards and drive automation across our production environment. You will work closely with product/application engineering cloud infrastructure and security teams to build resilient self-operating systems across an increasingly complex distributed infrastructure.
- Observability Platform: Own scale and continuously improve our existing metrics logging distributed tracing dashboards and alerting infrastructure to support increasing platform scale multiple deployment sites and multi-tenant environments.
- Reliability Engineering: Establish SLO frameworks error-budget reporting production-readiness standards and reliability engineering best practices.
- Developer Enablement: Build self-service observability capabilities standardized instrumentation reusable dashboards and actionable alerting for engineering teams.
- Incident Management: Establish incident-management tooling escalation processes runbooks and post-incident review practices. Drive improvements that prevent recurring incidents.
- Automation: Eliminate operational toil through automation intelligent alerting automated remediation and resilient system design.
- Technical Leadership: Own the architecture and roadmap for Reliability & Observability mentor engineers and collaborate with other engineering teams on cross-cutting reliability challenges.
Application teams own their services and cloud infrastructure engineers own the underlying infrastructure. Your role is to provide the shared capabilities standards and engineering expertise that help them operate reliably.
- 7 years of software engineering platform engineering or SRE experience including operating business-critical production systems at high scale and complexity.
- Strong experience with observability technologies such as Elasticsearch Logstash Kibana (ELK) Grafana Loki OpenTelemetry Jaeger or equivalent platforms.
- Hands-on experience with AWS Kubernetes Linux networking and distributed systems.
- Strong programming and scripting skills in Go Java Python or similar languages with an automation-first mindset.
- Experience designing scaling and improving production observability platforms effective monitoring alerting SLOs and incident-management practices.
- A solid understanding of distributed-system failure modes scalability fault tolerance and production debugging.
- Experience with infrastructure as code CI/CD and cloud-native architectures.
- Ability to drive technical initiatives independently make sound architectural decisions and influence engineering teams without relying on organizational authority.
- Experience operating high-availability multi-tenant SaaS platforms particularly in fintech or other regulated environments.
- Experience designing observability architectures for geographically distributed infrastructure multiple Kubernetes clusters and isolated tenant environments.
- Experience with large-scale telemetry pipelines observability cost optimization and high-cardinality metrics.
- Experience with chaos engineering resilience testing and automated incident remediation.
- Experience building internal developer platforms or self-service engineering tools.
Above all we are looking for an engineer who enjoys building reliable platforms not just maintaining monitoring tools.
- Build a high-impact platform in the financial services domain
- Opportunity to work with some of the best brains in fintech and a strong product engineering team
- Grow exponentially by working in small and transparent teams and influence tech culture
- Increase your geek quotient by attending meetups and conferences.
Cybrilla is a financial infrastructure company aiming to disrupt the way mutual funds work. We work with some of the largest financial institutions as well as some of the fastest growing fintechs in the country. We are decentralizing distribution and changing the face of the industry. We are a SEBI registered RTA and are co-authoring the Mutual Fund protocol on ONDC. More details:
Our product Fintech Primitives (FP) is an API platform that provides solutions to the problem statements of the Indian Mutual Fund domain. The APIs abstract domain regulatory and technical complexities to enable customers to build different distribution use cases in a short time. More details:
Required Experience:
Senior IC