Observability Engineer & Terraform

Synechron

Job Location:

Bengaluru - India

Monthly Salary: Not Disclosed

Posted on: 30+ days ago

Vacancies: 1 Vacancy

Job Summary

Job Summary
Synechron is seeking an experienced Observability Consultant to enhance and support our enterprise observability platform. This role involves designing implementing and scaling solutions that monitor system metrics logs and traces ensuring optimal system performance reliability and faster incident response. The ideal candidate will collaborate with application teams to onboard services develop dashboards and establish monitoring standards supporting enterprise operational excellence and continuous improvement initiatives.

Software Requirements

Required Software Proficiency:

Prometheus (latest version) Grafana for metrics collection dashboarding and alerting support (latest version)
Loki Tempo OpenTelemetry for log management and trace collection supporting distributed systems monitoring
Cloud-native architectures: Kubernetes supporting scalable and resilient environments supporting observability in cloud ecosystems
Monitoring and alerting tools support: support for enterprise alerting incident management and operational dashboards supporting high-availability environments
CI/CD integration tools supporting self-service and automation of observability workflows support

Preferred Software Skills:

Cloud platform integrations (AWS Azure GCP) supporting observability in cloud environments (preferred)
Support for infrastructure as code: Terraform support for automation and reproducibility (preferred)
Incident management tools (e.g. PagerDuty ServiceNow) supporting operational workflows (preferred)

Overall Responsibilities

Design build and support scalable observability solutions across metrics logs and traces for enterprise systems
Collaborate with application platform and DevOps teams to onboard services develop dashboards and establish monitoring standards
Develop and optimize dashboards alerts and instrumentation to enable proactive incident detection and response
Troubleshoot complex performance security and reliability issues supporting high-availability environments
Support and improve the deployment and integration of observability tools within cloud and on-premises infrastructures
Automate data collection alerting and incident workflows supporting DevOps and SRE best practices
Maintain detailed documentation on system architecture dashboards and operational procedures supporting audits and compliance
Stay current with emerging trends best practices and tools supporting enterprise observability and automation support

Technical Skills (By Category)

Monitoring & Observability Tools (Essential):
- Prometheus Grafana support for metrics collection dashboarding and alert configuration
- Loki Tempo OpenTelemetry support for log and trace collection in distributed systems

Cloud & Infrastructure:
- AWS Azure or GCP supporting cloud-native observability and monitoring integration (preferred)

Automation & Infrastructure as Code:
- Terraform supporting environment and tool deployment support (preferred)

Support Tools & Platforms:
- PagerDuty ServiceNow or equivalent incident response support tools supporting alert escalation and operational workflows

Experience Requirements

5 years supporting enterprise observability monitoring and incident response environments supporting high-availability systems
Proven experience designing implementing and scaling metrics logs and trace collection in cloud or hybrid infrastructures
Strong background supporting automation alert management and incident resolution support in enterprise environments
Experience working with cloud-native monitoring practices supporting scalable resilient services supported by Kubernetes or similar platforms
Industry experience supporting regulated environments (finance healthcare or large-scale enterprise) with compliance standards (preferred)

Day-to-Day Activities

Deploy configure and support scalable observability tools across cloud and on-premises systems
Develop and maintain dashboards alerts and instrumentation supporting proactive monitoring
Collaborate with application platform and DevOps teams to onboard services customize dashboards and refine monitoring processes
Troubleshoot performance issues logging failures and security incidents supporting high availability
Automate data collection alerting and incident workflows to support agile response and resolution
Support infrastructure scaling system upgrades and environment configuration supporting operational efficiency
Document system architectures dashboards alerting rules and operational procedures
Review incident logs analyze system health metrics and support continuous improvement initiatives

Qualifications

Bachelors degree in Computer Science Information Technology or a related technical field
5 years supporting enterprise observability monitoring and incident response in cloud or hybrid environments
Certifications supporting cloud platforms or monitoring tools (e.g. Prometheus Grafana certifications or cloud certifications) are advantageous
Proven experience supporting high-availability secure and scalable observability environments supporting enterprise standards

Professional Competencies

Strong analytical and troubleshooting skills supporting complex system monitoring and incident response activities
Leadership qualities to guide teams enforce best practices and support operational excellence
Excellent stakeholder communication and documentation skills supporting operational and compliance reporting
Adaptability to evolving technologies security and compliance standards supporting enterprise resilience
Strategic thinking to develop scalable effective and automated observability solutions supporting operational goals
Organizational skills to manage multiple environments dashboards and incident response activities efficiently

SYNECHRONS DIVERSITY & INCLUSION STATEMENT

Diversity & Inclusion are fundamental to our culture and Synechron is proud to be an equal opportunity workplace and is an affirmative action employer. Our Diversity Equity and Inclusion (DEI) initiative Same Difference is committed to fostering an inclusive culture promoting equality diversity and an environment that is respectful to all. We strongly believe that a diverse workforce helps build stronger successful businesses as a global company. We encourage applicants from across diverse backgrounds race ethnicities religion age marital status gender sexual orientations or disabilities to apply. We empower our global workforce by offering flexible workplace arrangements mentoring internal mobility learning and development programs and more.

All employment decisions at Synechron are based on business needs job requirements and individual qualifications without regard to the applicants gender gender identity sexual orientation race ethnicity disabled or veteran status or any other characteristic protected by law.

Candidate Application Notice

Required Experience:

Job SummarySynechron is seeking an experienced Observability Consultant to enhance and support our enterprise observability platform. This role involves designing implementing and scaling solutions that monitor system metrics logs and traces ensuring optimal system performance reliability and faster...

Software Requirements

Required Software Proficiency:

Prometheus (latest version) Grafana for metrics collection dashboarding and alerting support (latest version)
Loki Tempo OpenTelemetry for log management and trace collection supporting distributed systems monitoring
Cloud-native architectures: Kubernetes supporting scalable and resilient environments supporting observability in cloud ecosystems
Monitoring and alerting tools support: support for enterprise alerting incident management and operational dashboards supporting high-availability environments
CI/CD integration tools supporting self-service and automation of observability workflows support

Preferred Software Skills:

Cloud platform integrations (AWS Azure GCP) supporting observability in cloud environments (preferred)
Support for infrastructure as code: Terraform support for automation and reproducibility (preferred)
Incident management tools (e.g. PagerDuty ServiceNow) supporting operational workflows (preferred)

Overall Responsibilities

Design build and support scalable observability solutions across metrics logs and traces for enterprise systems
Collaborate with application platform and DevOps teams to onboard services develop dashboards and establish monitoring standards
Develop and optimize dashboards alerts and instrumentation to enable proactive incident detection and response
Troubleshoot complex performance security and reliability issues supporting high-availability environments
Support and improve the deployment and integration of observability tools within cloud and on-premises infrastructures
Automate data collection alerting and incident workflows supporting DevOps and SRE best practices
Maintain detailed documentation on system architecture dashboards and operational procedures supporting audits and compliance
Stay current with emerging trends best practices and tools supporting enterprise observability and automation support

Technical Skills (By Category)

Monitoring & Observability Tools (Essential):
- Prometheus Grafana support for metrics collection dashboarding and alert configuration
- Loki Tempo OpenTelemetry support for log and trace collection in distributed systems

Cloud & Infrastructure:
- AWS Azure or GCP supporting cloud-native observability and monitoring integration (preferred)

Automation & Infrastructure as Code:
- Terraform supporting environment and tool deployment support (preferred)

Support Tools & Platforms:
- PagerDuty ServiceNow or equivalent incident response support tools supporting alert escalation and operational workflows

Experience Requirements

5 years supporting enterprise observability monitoring and incident response environments supporting high-availability systems
Proven experience designing implementing and scaling metrics logs and trace collection in cloud or hybrid infrastructures
Strong background supporting automation alert management and incident resolution support in enterprise environments
Experience working with cloud-native monitoring practices supporting scalable resilient services supported by Kubernetes or similar platforms
Industry experience supporting regulated environments (finance healthcare or large-scale enterprise) with compliance standards (preferred)

Day-to-Day Activities

Deploy configure and support scalable observability tools across cloud and on-premises systems
Develop and maintain dashboards alerts and instrumentation supporting proactive monitoring
Collaborate with application platform and DevOps teams to onboard services customize dashboards and refine monitoring processes
Troubleshoot performance issues logging failures and security incidents supporting high availability
Automate data collection alerting and incident workflows to support agile response and resolution
Support infrastructure scaling system upgrades and environment configuration supporting operational efficiency
Document system architectures dashboards alerting rules and operational procedures
Review incident logs analyze system health metrics and support continuous improvement initiatives

Qualifications

Bachelors degree in Computer Science Information Technology or a related technical field
5 years supporting enterprise observability monitoring and incident response in cloud or hybrid environments
Certifications supporting cloud platforms or monitoring tools (e.g. Prometheus Grafana certifications or cloud certifications) are advantageous
Proven experience supporting high-availability secure and scalable observability environments supporting enterprise standards

Professional Competencies

Strong analytical and troubleshooting skills supporting complex system monitoring and incident response activities
Leadership qualities to guide teams enforce best practices and support operational excellence
Excellent stakeholder communication and documentation skills supporting operational and compliance reporting
Adaptability to evolving technologies security and compliance standards supporting enterprise resilience
Strategic thinking to develop scalable effective and automated observability solutions supporting operational goals
Organizational skills to manage multiple environments dashboards and incident response activities efficiently

SYNECHRONS DIVERSITY & INCLUSION STATEMENT

Candidate Application Notice

Required Experience:

Apply Now

About Company

Synechron

Chez Synechron, nous croyons en la puissance du numérique pour transformer les entreprises en mieux. Notre cabinet de conseil mondial combine la créativité et la technologie innovante pour offrir des solutions numériques de premier plan. Les technologies progressistes et les stratégie ... View more

View Profile View Profile

AI AutoApply

Apply to 100+ jobs with one click