Staff Engineer (Datadog Engineer)
Department:
Job Summary
Requirements
- Experience : 5.5 years
- Strong experience in enterprise monitoring observability or cloud infrastructure engineering.
- Strong hands-on experience with Datadog across Infrastructure Monitoring APM Synthetic Monitoring Database Monitoring Real User Monitoring (RUM) and Log Management.
- Expertise in Microsoft Azure DevOps Pipeline Strategy including designing and implementing CI/CD pipelines for monitoring and observability solutions.
- Strong experience with Terraform for Infrastructure as Code (IaC) automating Datadog configuration and deployment.
- Strong knowledge of ITIL processes including Incident Problem Change and Service Management.
- Hands-on experience in incident management root cause analysis troubleshooting and production support.
- Experience creating and maintaining Datadog dashboards monitors alerts log pipelines and service-level monitoring.
- Strong understanding of cloud platforms such as AWS Azure or Google Cloud Platform.
- Experience with automation tools and scripting to streamline monitoring deployment and operational processes.
- Knowledge of application performance monitoring distributed tracing infrastructure monitoring and observability best practices.
- Familiarity with ServiceNow and ITSM workflows is preferred.
- Exposure to other enterprise monitoring and observability tools is an advantage.
- Strong analytical troubleshooting and problem-solving skills with the ability to resolve complex production issues.
- Excellent verbal and written communication skills with the ability to collaborate across cross-functional teams.
- Ability to manage multiple priorities in a fast-paced enterprise environment.
Responsibilities
- Design implement and manage Datadog monitoring solutions across infrastructure applications databases synthetic monitoring and Real User Monitoring (RUM).
- Configure and maintain Datadog dashboards monitors alerts log pipelines and observability frameworks to provide comprehensive operational visibility.
- Develop and maintain Terraform modules to automate Datadog configuration deployment and infrastructure provisioning.
- Design and implement Azure DevOps CI/CD pipelines for monitoring configuration automation and continuous delivery.
- Collaborate with application infrastructure cloud and DevOps teams to ensure end-to-end monitoring coverage across enterprise platforms.
- Implement monitoring standards instrumentation and best practices for cloud-native and enterprise applications.
- Support incident problem and change management processes while ensuring adherence to ITIL best practices.
- Perform root cause analysis troubleshoot monitoring issues and optimize platform performance to improve service reliability.
- Develop monitoring strategies for cloud environments across AWS Azure and hybrid infrastructure.
- Configure log collection parsing enrichment and routing to support operational monitoring and analytics.
- Build automated monitoring and alerting solutions to proactively identify performance availability and infrastructure issues.
- Integrate Datadog with enterprise tools ITSM platforms and automation frameworks to improve operational efficiency.
- Maintain technical documentation monitoring standards operational procedures and deployment guides.
- Participate in production support release activities platform upgrades and continuous improvement initiatives.
- Work closely with stakeholders to enhance observability capabilities optimize monitoring coverage and improve overall platform reliability and operational excellence.
Qualifications :
Bachelors or masters degree in computer science Information Technology or a related field.
Remote Work :
No
Employment Type :
Full-time
About Company
Nagarro helps future-proof your business through a forward-thinking, fluidic, and CARING mindset. We excel at digital engineering and help our clients become human-centric, digital-first organizations, augmenting their ability to be responsive, efficient, intimate, creative, and susta ... View more