Consultant, Site Reliability Engineering
Job Summary
Location: Toronto ON
Onsite Flexibility: Hybrid Onsite at CIBC Square once per week (Wednesdays) with an additional onsite requirement every third Friday of the month.
- Position Type: Contract
- Contract Duration: 6 months (with potential for extension or conversion to FTE)
- Pay Rate: C$65.00C$72.00 / Hour (CAD)
- Shift / Schedule: 37.5 hours/week 9:00 AM 5:00 PM Monday to Friday; overtime potential: Yes
As a member of the Application Reliability Engineering Platform team the Consultant Site Reliability Engineering will play a key role in Digital and Client Experience Technology to support the Banking Centre Line of Business ensuring the seamless operation and ongoing improvement of application platforms across the organization. You will work closely with cross-functional teams providing technical expertise to resolve application and infrastructure technology issues on medium to highly complex projects in compliance with service standards policies and procedures. The contractor will be responsible for building and migrating CI/CD pipelines into a new environment transitioning from Jenkins to GitHub Actions to streamline deployment processes and modernize automation workflows. Youll have the flexibility to manage your work activities within a hybrid work arrangement where youll spend 13 days per week on-site while other days will be remote.
Technical Expertise
- Manage and optimize automated application deployments across multiple environments ensuring consistency and reliability.
- Monitor system health performance metrics and alerts to proactively identify and resolve issues before client impact.
- Participate in and as required lead incident response and post-mortem analysis to identify root causes and implement preventive measures.
Partnership
- Continuously evaluate and improve systems processes and tools to enhance reliability scalability and efficiency.
- Define and maintain observability strategies to ensure visibility into application behaviors and performance.
- Participate in the on-call rotation for application support responding to incidents and resolving technical issues.
- Balance feature development speed and system reliability by adhering to well-defined service level objectives (SLOs).
- Support AIOps initiatives to reduce manual intervention and improve operational intelligence through automation.
Continuous Improvement
- Contribute to the development and execution of observability and telemetry strategies.
- Provide input to best practices standards and documentation for site reliability engineering (SRE) methodologies and technologies.
- Participate in defining standard reliability and resilience requirements for infrastructure and application components.
- Act as a point of escalation for production application issues providing advanced troubleshooting and resolution.
- 510 years of experience in SRE DevOps or Application Support environments.
- Strong Linux/Unix administration and troubleshooting skills.
- Hands-on experience with Azure OpenShift (OCP) and containerized environments.
- Strong application support incident management and problem-solving experience.
- Experience building and supporting CI/CD pipelines using Jenkins and/or GitHub Actions.
- Proficiency in scripting using Python Bash PowerShell or similar.
- Experience with monitoring observability and automation tools.
- Knowledge of file transfer protocols networking fundamentals and application security best practices.
- Demonstrated expertise in automating application deployments incident response workflows and operational tasks using scripting and orchestration techniques.
- Solid understanding of AIOps principles and hands-on experience applying automation and machine learning to improve observability anomaly detection and incident resolution.
- Proficiency in scripting and programming languages (e.g. Python PowerShell Bash JavaScript) for building automation related to application monitoring deployment and recovery.
- Familiarity with application security best practices including authentication authorization secrets management and secure coding principles.
- Experience working with continuous integration and continuous delivery (CI/CD) pipelines and version control systems (e.g. GitHub) with a focus on application delivery and release automation.
- Working knowledge of observability strategies including telemetry logging tracing and alerting to maintain application health and performance.
- Understanding of networking fundamentals as they relate to application performance and reliability.
- Experience with pipeline-integrated security scanning and compliance tools to support secure application delivery.
- Strong analytical and problem-solving abilities.
- Excellent verbal and written communication skills.
- Demonstrated leadership skills with the ability to influence and collaborate across teams.
- CIBC or banking industry experience.
- JavaScript or other scripting/automation experience.
- Bachelors degree or equivalent in Computer Science or a Technical discipline.
- 510 years of experience in site reliability engineering application support or DevOps roles with a strong focus on application lifecycle automation and operational excellence.
- Medical Vision and Dental Insurance Plans
- 401k Retirement Fund
- Interview process: Two rounds of interviews. The first round will be conducted remotely and candidates who are successful will be invited to a second-round interview onsite at CIBC Square.
- The candidate may begin their assignment while employment verification and if applicable professional credential checks are in progress. The supplier is required to have the checks completed within 30 days of the start date.
This client is a leading financial services and banking institution operating across Canada with a significant presence in major urban centres including Toronto. Employing tens of thousands of professionals the organization ranks among the top tier of Canadian banks and serves millions of customers through an extensive branch network digital platforms and contact centre operations. The technology teams supporting this institution include DevOps engineers site reliability engineers business analysts and enterprise program specialists who collaborate across agile cross-functional environments to support critical banking applications and digital infrastructure.
GTT is a minority-owned staffing firm and a subsidiary of Chenega Corporation a Native American-owned company in Alaska. We highly value diverse and inclusive workplaces and support Fortune 500 organizations across banking financial services technology life sciences biotech utilities and retail sectors throughout the U.S. and Canada.
Job Number: 26-14698
#LI-Hybrid #LI-GTT
Required Experience:
Contract