Site Reliability Engineer (SRE) DevOps Engineer
Job Summary
Who are we
Fulcrum Digital is an agile and next-generation digital accelerating company providing digital transformation and technology services right from ideation to implementation. These services have applicability across a variety of industries including banking & financial services insurance retail higher education food healthcare and manufacturing.
Were looking for a Kafka Messaging / SRE Engineer to join our growing platform engineering team and help build operate and scale missioncritical messaging services.
- Own and operate Kafka-based messaging platforms in production environments
- Apply SRE principles to improve reliability availability and performance
- Drive DevOps & automation initiatives to reduce toil and manual operations
- Build and enhance automation using Ansible scripts and CI/CD pipelines
- Perform incident management RCA capacity planning and operational readiness
- Collaborate closely with application and platform engineering teams
- Contribute to Java-based tooling and platform enhancements
- What Were Looking For
- 36 years of experience working with Kafka / messaging systems
- Strong understanding of Kafka architecture (brokers topics partitions replication)
- Hands-on experience with SRE / DevOps practices
- Proven skills in automation (Ansible scripting CI/CD)
- Java development background (ability to debug enhance or build platform tools)
- Experience with Linux distributed systems monitoring & alerting
- Exposure to incident response production support and operational excellence
Required Skills:
Understanding of event-driven architectures Distributed systems - How clusters are formed Quorum management Failure handling. 3 to 5 years of hands-on Experience in MQ or NATS broker or similar messaging solutions. Understanding of Kafka clustering would be good to have. Knows Client-Server communication aspects - sockets TLS protocol etc Understands the concept of region and AZs. Provide L2 support production systems like application database middleware components infrastructure and network components. Manage production incidents end-to-end within defined SLAs with focus on resolution rather than who caused it. Interact with various stakeholders such as Release managers program leads service managers development and test leads Review operational readiness requirements such as monitoring and alerting log rotation and resilience of the components and report the gaps Provide pre-implementation support with activities such as release notes review and implementation dry runs. Protect production components by running health checks monitoring latency and memory utilization. Automate day-to-day activities and propose changes that improve reliability Participate in CAB and provide feedback on change requests Support the DevOps team in testing the promoted pipelines and suggest automation of configuration items. Practice incident management best practices and perform RCA. Participate in disaster recovery tests and operational acceptance tests Analyze the technology stack that makes up the product and optimize recovery time objective. Work with team members spread across and time zones Share knowledge document improvements and mentor junior resources It is good to have skills using Jenkins to orchestrate builds and link to Sonar Maven etc. to build out the CI/CD pipeline. Support deployments of code into multiple lower environments. Supporting current processes needed with an emphasis on automating everything as soon as possible. It is good to have skill to design Implement and enhance our deployment automation based on Chef. We need proven experience designing and implementing an overall release and deployment process. It is good to have skill to design and implement a Git based code management strategy that will support multiple environment deployments in parallel. Experience with automation for Branch management code promotions and version management. Engage in and improve the whole lifecycle of servicesfrom inception and design through deployment operation and refinement. Requirements MQ/EB Understanding of event-driven architectures Distributed systems - How clusters are formed Quorum management Failure handling. 3 to 5 years of hands-on Experience in MQ or NATS broker or similar messaging solutions. An understanding of Kafka clustering would be good to have. Knows Client-Server communication aspects - sockets TLS protocol etc Understand the concept of region and AZs. Deployments MTF/Prod Maintenance items (including stop/start Disaster Recovery-related activities etc.) CR for changes in MTF/Prod Good knowledge on Nginx Tools - Log Monitoring Tool - Splunk Application Monitoring tool - Dynatrace Ticketing incident/problem management tool - Remedy Dev-ops Basics - CI-CD Basics Overview of Git Bit-bucket SonarQube Ansible/Chef Skills - Linux & Shell Scripting ITIL / ITSM PL/SQL Troubleshooting Jenkins - CI/CD Groovy Scripting/Yaml Ansible/Chef Nginx Java / JEE Event-Driven Architectures MQ or NATS broker or similar messaging solutions. Kafka Client-server communication aspects - sockets TLS protocol Understand the concept of region and AZs.