Application Production Support Engineer
Job Summary
The Application Production Support Engineer is a detail-oriented and technically skilled professional responsible for managing and supporting critical business applications in a production environment with a strong focus on delivering high-quality IT services.
This role is responsible for ensuring the stability availability performance and reliability of applications within the teams scope. The professional will also manage incident resolution in a timely manner collaborating closely with internal stakeholders (Development Infrastructure and Architecture teams) as well as external partners and service providers to drive sustainable and long-term solutions.
Main Responsibilities
Application Stability & Availability (Expert)
- Monitor maintain and support business-critical applications ensuring high availability stability and optimal performance.
- Actively participate in Incident and Problem Management processes including:
- Participation in Major Incident (P1/P2) Situation Rooms.
- Root Cause Analysis (RCA) activities.
- Identification of incident trends and recurring issues.
- Contribution to permanent corrective actions and preventive measures.
- Ensure compliance with ITIL governance practices operational procedures and established Service Level Agreements (SLAs).
- Execute application deployments releases and change requests following ITIL and DevOps methodologies.
- Proactively identify analyze and resolve technical issues to support uninterrupted business operations.
- Participate in on-call support rotations and provide 24/7 support coverage for critical applications when required.
Technical Support & Collaboration (Expert)
- Serve as a primary point of contact between Production Support and Development teams for troubleshooting and issue resolution.
- Collaborate closely with Scrum and DevOps teams to design deploy maintain and continuously improve application services.
- Implement upgrades patches configuration changes and new functionalities while minimizing business impact and ensuring service continuity.
- Contribute to the continuous improvement of operational processes automation and platform reliability.
Documentation & Knowledge Sharing (Expert)
- Create maintain and continuously update technical documentation including operational procedures system configurations troubleshooting guides and runbooks.
- Promote knowledge sharing and best practices across global support teams to improve operational efficiency and service quality.
- Support the development and maintenance of knowledge bases to facilitate issue resolution and onboarding activities.
Platform Monitoring & Observability (Expert)
- Implement maintain and optimize monitoring and observability solutions across production environments.
- Leverage platforms such as Dynatrace and other observability tools to ensure proactive monitoring and rapid incident detection.
- Collaborate with Development teams and Centers of Expertise to define and enhance observability standards and monitoring strategies.
- Promote an observability-first mindset enabling early identification and resolution of potential service disruptions.
- Continuously improve monitoring dashboards alerting mechanisms and operational metrics.
General Responsibilities
- Complete all mandatory training required for the effective operation of the IT Production area and compliance with company policies.
- Perform additional activities when required to support business objectives and ensure the proper functioning of the Center of Expertise.
- Contribute to continuous improvement initiatives within the Production Support organization.
Qualifications :
API Application Servers & Kubernetes (Expert)
- Strong knowledge of Java Application Servers particularly Red Hat JBoss EAP.
- Understanding of Java application troubleshooting including:
- Heap dump analysis.
- Thread dump analysis.
- Performance tuning and optimization.
- Hands-on experience with OpenShift and Kubernetes-based platforms.
- Knowledge of Cloud-native architectures and containerized environments.
- Experience with API Gateway solutions such as Axway and Apigee.
RHEL Linux Operating System (Expert)
- Advanced administration and troubleshooting experience in Red Hat Enterprise Linux (RHEL) environments.
Tooling & Automation (Expert)
- Monitoring and observability platforms:
- Dynatrace
- Grafana
- Prometheus
- ELK Stack
- Jaeger
- CI/CD tools and practices:
- GitLab
- Jenkins
- Argo CD
- Nexus Repository (Sonatype)
- Infrastructure and configuration management:
- Ansible
- Terraform
Databases (Expert)
- Strong knowledge of relational databases including:
- Microsoft SQL Server
- PostgreSQL
- Experience with database monitoring performance analysis and troubleshooting.
Agile Methodologies
Scrum Framework (Practitioner)
- Experience working within Agile and Scrum environments collaborating effectively with cross-functional teams.
Additional Information :
English mandatory.
Remote Work :
No
Employment Type :
Full-time
About Company
Inetum is a European leader in digital services. Inetums team of 28,000 consultants and specialists strive every day to make a digital impact for businesses, public sector entities and society. Inetums solutions aim at contributing to its clients performance and innovation as well ... View more