We are looking for aSeniorObservability/ Monitoring Engineerto design implement and optimizeobservabilitysolutions for large-scale enterprise platforms. This role will play a critical part in enabling proactive monitoring faster incident detection and improved system reliability across Salesforce and Microsoft Azure environments. The ideal candidate will have strong expertise in metrics logs traces alerting strategies andobservabilitytooling along with hands-on experience supporting production environments.
Key Responsibilities -ObservabilityEngineering:
Design and implement end-to-endobservabilityframeworks across Salesforce and Azure platforms
Establish unified monitoring across logs metrics and distributed tracing
Define and standardizeobservabilitybest practices dashboards and alerting strategies
Enable proactive detection of issues through intelligent alerting and anomaly detection Monitoring & Tooling
Implement and manage tools such as Azure Monitor Application Insights Splunk Datadog Grafana Prometheus or similar
Build actionable dashboards for operations SRE and business stakeholders
Optimize alert noise reduction and improve signal-to-noise ratio
Continuously enhance monitoring coverage across applications and infrastructure Incident Support & Reliability
Support incident management by providing deep insights usingobservabilitydata
Perform root cause analysis (RCA) leveraging logs traces and metrics
Collaborate with SRE and engineering teams to improve system reliability and performance
Contribute to post-incident reviews and continuous improvement initiatives Automation & Integration
Automate monitoring setup and configuration using Infrastructure as Code (IaC)
Integrateobservabilitytools with CI/CD pipelines and DevOps workflows
Develop scripts or tools to enhance monitoring capabilities and data collection Platform & Integration Support
Monitor and optimize Salesforce applications including integrations and APIs
Support Azure-based services ensuring visibility across compute storage and networking layers
Ensure end-to-endobservabilityacross integrated systems and middleware Governance & Compliance
Ensureobservabilitypractices align with security and compliance requirements (e.g. SOX)
Maintain documentation runbooks and monitoring standards
Support audits and governance reviews as required
Required Skills & Qualifications Technical Skills
Strong experience inobservability monitoring or SRE roles
Hands-on experience with Azure (Azure Monitor Application Insights)
Understanding of cloud-native and microservices architectures Operational Excellence
Experience in incident management and RCA
Ability to analyze system performance and recommend improvements
Strong troubleshooting and analytical skills Soft Skills
Strong communication and collaboration skills
Ability to work with cross-functional and global teams
Proactive mindset with a focus on continuous improvement
We may use artificial intelligence (AI) tools to support parts of the hiring process such as reviewing applications analyzing resumes or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed please contact us.
Required Experience:
Senior IC
SeniorObservability/ Monitoring EngineerRole OverviewWe are looking for aSeniorObservability/ Monitoring Engineerto design implement and optimizeobservabilitysolutions for large-scale enterprise platforms. This role will play a critical part in enabling proactive monitoring faster incident detection...
SeniorObservability/ Monitoring Engineer
Role Overview
We are looking for aSeniorObservability/ Monitoring Engineerto design implement and optimizeobservabilitysolutions for large-scale enterprise platforms. This role will play a critical part in enabling proactive monitoring faster incident detection and improved system reliability across Salesforce and Microsoft Azure environments. The ideal candidate will have strong expertise in metrics logs traces alerting strategies andobservabilitytooling along with hands-on experience supporting production environments.
Key Responsibilities -ObservabilityEngineering:
Design and implement end-to-endobservabilityframeworks across Salesforce and Azure platforms
Establish unified monitoring across logs metrics and distributed tracing
Define and standardizeobservabilitybest practices dashboards and alerting strategies
Enable proactive detection of issues through intelligent alerting and anomaly detection Monitoring & Tooling
Implement and manage tools such as Azure Monitor Application Insights Splunk Datadog Grafana Prometheus or similar
Build actionable dashboards for operations SRE and business stakeholders
Optimize alert noise reduction and improve signal-to-noise ratio
Continuously enhance monitoring coverage across applications and infrastructure Incident Support & Reliability
Support incident management by providing deep insights usingobservabilitydata
Perform root cause analysis (RCA) leveraging logs traces and metrics
Collaborate with SRE and engineering teams to improve system reliability and performance
Contribute to post-incident reviews and continuous improvement initiatives Automation & Integration
Automate monitoring setup and configuration using Infrastructure as Code (IaC)
Integrateobservabilitytools with CI/CD pipelines and DevOps workflows
Develop scripts or tools to enhance monitoring capabilities and data collection Platform & Integration Support
Monitor and optimize Salesforce applications including integrations and APIs
Support Azure-based services ensuring visibility across compute storage and networking layers
Ensure end-to-endobservabilityacross integrated systems and middleware Governance & Compliance
Ensureobservabilitypractices align with security and compliance requirements (e.g. SOX)
Maintain documentation runbooks and monitoring standards
Support audits and governance reviews as required
Required Skills & Qualifications Technical Skills
Strong experience inobservability monitoring or SRE roles
Hands-on experience with Azure (Azure Monitor Application Insights)
Understanding of cloud-native and microservices architectures Operational Excellence
Experience in incident management and RCA
Ability to analyze system performance and recommend improvements
Strong troubleshooting and analytical skills Soft Skills
Strong communication and collaboration skills
Ability to work with cross-functional and global teams
Proactive mindset with a focus on continuous improvement
We may use artificial intelligence (AI) tools to support parts of the hiring process such as reviewing applications analyzing resumes or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed please contact us.
Brillio is a global leader in Enterprise Digital Transformation Solutions, providing strategic consulting services and solutions using emerging technologies.