Senior Lead Observability – Dynatrace
Job Summary
About Northern Trust
As a global leader in innovative wealth management asset servicing asset management and banking services Northern Trust (Nasdaq: NTRS) is proud to guide the worlds most successful individuals families corporations and institutions.
Since 1889 we have aligned our efforts with our three guiding Principles That Endure: Service Expertise and Integrity. Together they reflect the three cornerstones of business conduct which we strive to instill in our employees whom we call partners and to provide to our clients and the communities we serve worldwide.
With more than 135 years of financial experience and over 24000 partners we serve the worlds most sophisticated clients using leading technology and exceptional service.
We are seeking a highly skilled Senior Observability Engineer to support and enhance enterprise observability capabilities across infrastructure applications cloud and business services. This is an individual contributor role focused on hands-on engineering platform operations monitoring standardization automation and continuous improvement across observability platforms including Dynatrace Microsoft SCOM ServiceNow ITOM & Event Management and related monitoring integrations.
The successful candidate will play a key role in improving monitoring effectiveness reducing alert noise enhancing event correlation enabling faster incident response and supporting operational resilience through automation and observability best practices. The role requires strong technical expertise analytical thinking and the ability to work closely with platform application infrastructure and operations teams to deliver reliable and scalable monitoring solutions.
This role builds on the observability focus outlined in the reference document including enterprise visibility across applications infrastructure cloud and business services through Dynatrace ServiceNow ITOM Event Management AI-driven observability and automation-led efficiency improvements.
- Implement maintain and continuously improve enterprise observability capabilities across Dynatrace SCOM ServiceNow ITOM & Event Management and supporting monitoring tools.
- Configure and support monitoring for infrastructure applications services databases middleware cloud and hybrid environments to ensure end-to-end visibility and operational stability.
- Develop tune and optimize monitoring alerts dashboards thresholds synthetic checks anomaly detection and service-level views to improve signal quality and reduce false positives.
- Drive alert hygiene activities including duplicate alert reduction threshold tuning suppression logic event enrichment and monitoring standardization across technology teams.
- Support integrations between observability platforms and ServiceNow Event Management ensuring events are enriched correlated deduplicated and routed effectively for incident response.
- Manage and enhance monitoring integrations from tools such as Dynatrace SCOM infrastructure monitoring sources and application monitoring platforms into ServiceNow Event Management.
- Use Dynatrace capabilities such as service flow distributed tracing Davis AI anomaly detection problem correlation dashboards management zones tags and alerting profiles to improve root cause analysis and operational insights.
- Support Microsoft SCOM monitoring activities including management pack configuration alert rule tuning agent health checks monitoring coverage and operational troubleshooting.
- Build automation scripts and reusable solutions using Python PowerShell REST APIs YAML/JSON CI/CD pipelines and other automation frameworks to improve onboarding monitoring configuration health checks reporting and operational efficiency.
- Contribute to observability-as-code practices by supporting standardized repeatable and automated monitoring onboarding patterns.
- Partner with application infrastructure cloud and support teams to onboard new applications and services into enterprise monitoring platforms.
- Troubleshoot monitoring gaps integration failures agent issues event flow problems and alerting defects across observability systems.
- Support operational reporting and KPI tracking related to alert volume noise reduction MTTR improvement event quality monitoring coverage and automation adoption.
- Maintain technical documentation operational runbooks configuration standards troubleshooting guides and onboarding procedures.
- Participate in incident reviews and problem management discussions to identify opportunities for monitoring improvement and proactive detection.
- Apply ITIL practices and event lifecycle management principles to improve incident quality operational response and service reliability.
- 10 years of overall IT experience with strong hands-on experience in observability monitoring event management infrastructure operations application support or platform engineering.
- Strong practical experience with Dynatrace including OneAgent dashboards alerts management zones synthetic monitoring service flow problem detection tagging and Davis AI capabilities.
- Hands-on experience with Microsoft SCOM including alert configuration management packs agent monitoring rule tuning infrastructure monitoring and operational troubleshooting.
- Strong working knowledge of ServiceNow Event Management / ITOM including event ingestion event rules alert correlation deduplication enrichment service mapping awareness and incident integration.
- Experience integrating monitoring tools with ServiceNow or similar ITSM platforms.
- Good understanding of observability concepts including metrics logs traces events topology service health synthetic monitoring and full-stack monitoring.
- Strong scripting and automation experience using Python PowerShell REST APIs shell scripting or similar technologies.
- Experience with CI/CD tools source control configuration files automation workflows and infrastructure-as-code or observability-as-code practices.
- Familiarity with cloud and hybrid environments such as Azure AWS VMware Windows Linux databases middleware and enterprise infrastructure platforms.
- Understanding of ITIL processes especially Incident Management Problem Management Change Management Event Management and operational support models.
- Strong analytical and troubleshooting skills with the ability to identify monitoring gaps alert quality issues and event flow problems.
- Ability to work independently as an individual contributor while collaborating effectively with engineering operations application and support teams.
- Experience with OpenTelemetry log analytics AIOps machine learning-based monitoring predictive monitoring or GenAI-assisted incident diagnosis.
- Exposure to enterprise observability platforms beyond Dynatrace and SCOM such as Splunk Elastic Azure Monitor Grafana Prometheus or other monitoring ecosystems.
- Experience building dashboards operational reports executive metrics and service health views.
- Knowledge of CMDB service mapping application dependency mapping and event-to-CI relationships in ServiceNow.
- Experience supporting large-scale enterprise monitoring environments in financial services or regulated industries.
- Familiarity with Ansible Terraform GitHub Azure DevOps Jenkins or other automation and DevOps tooling.
- Ability to contribute to monitoring governance standards best practices and platform maturity improvements.
Preferred certifications include:
- Dynatrace Associate Professional or equivalent certification.
- Microsoft SCOM or Microsoft infrastructure/platform certifications.
- ServiceNow ITOM / Event Management certification.
- Microsoft Azure or AWS certification.
- ITIL Foundation or higher certification.
The individual in this role will be measured on the ability to deliver measurable improvements in observability maturity monitoring quality automation and operational effectiveness. Key success measures include:
- Improved incident quality and faster triage through actionable alerts and contextual event information.
- Increased monitoring coverage across applications infrastructure and services.
- Faster and more consistent onboarding of applications into observability platforms.
- Increased use of automation for repeatable monitoring and operational tasks.
- Improved dashboard quality service health visibility and operational reporting.
- Stronger alignment with enterprise monitoring standards and ITIL event management practices.
- Reduction in duplicate noisy unactionable or low-value alerts.
- Improved signal-to-noise ratio across monitoring platforms.
- Better event correlation and enrichment through ServiceNow Event Management.
Working with Us
As a Northern Trust partner you will be part of a flexible and collaborative work culture which has a strong history of financial strength and stability. Movement within the organization is encouraged senior leaders are accessible and you can take pride in working for a company committed to an inclusive workplace and assisting the communities we serve.
Philanthropy is deeply rooted in Northern Trusts history and is an essential element of our culture. Employees around the world give their time and talent to work for the greater good of their communities.
Reasonable Accommodation
Northern Trust is committed to working with and providing adjustments to individuals with health conditions and disabilities. If you need a reasonable accommodation for any part of the employment process please email our HR Service Center at or alternatively you can discuss your individual requirements with the recruiter you are working with.
About Our Pune Office
The Northern Trust Pune office established in 2016 is now home to over 3000 employees. The office handles various functions including Operations for Asset Servicing and Wealth Management as well as delivering critical technology solutions that support business operations across the globe.
Our Pune team takes our commitment to service to 2024 they volunteered more than 10000 hours into the communities where they live and work. Learn more.
Required Experience:
Senior IC