Enter a job title or keyword

Site Reliability Engineering

Truist Bank


Job Location:

Atlanta, GA - USA

Monthly Salary: Not provided by the employer
Posted: 21 August 2026 (Yesterday)
Application Deadline: 18 November 2026
Vacancies: 1 Vacancy

Job Summary

The position is described below. If you want to apply click the Apply Now button at the top or bottom of this page. After you click Apply Now and complete your application youll be invited to create a profile which will let you see your application status and any communications. If you already have a profile with us you can log in to check status.

Need Help

If you have a disability and need assistance with the application you can request a reasonable accommodation. Send an email to Accessibility (accommodation requests only; other inquiries wont receive a response).

Regular or Temporary:

Regular

Language Fluency: English (Required)

Work Shift:

1st shift (United States of America)

Please review the following job description:

The Site Reliability Engineering Lead role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud and on-premises environments. This senior technical leader drives improvements in automation observability and incident management while collaborating across multiple business and technology teams.

Responsibilities include leading major incident responses driving problem management and implementing automation to reduce service downtime.

The role involves standardizing observability practices mentoring SRE team members and contributing to enterprise-wide reliability frameworks.

Candidates require 7 years of experience expertise in distributed systems Kubernetes automation scripting and strong leadership in incident management.

ESSENTIAL DUTIES AND RESPONSIBILITIES
Following is a summary of the essential functions for this job. Other duties may be performed both major and minor which are not mentioned below. Specific activities may change from time to time.
1. Implements software architecture and engineering approaches for complex initiatives within the job area contributing to technical plans and working to achieve operational targets with major impact on results.
2. Adopts and refines advanced software engineering standards practices and governance mechanisms for the job area influencing how multiple teams improve quality reliability and delivery.
3. Collaborates with senior engineers product partners and architecture teammates to shape technology approaches for the domain providing deep technical insight and proposing solution patterns that inform local roadmaps and priorities.
4. Leads the end-to-end technical design and implementation of scalable secure and highly available software solutions for the job area producing patterns and examples that other technical professionals can follow.
5. Independently troubleshoots and resolves complex technical issues in the area of responsibility designing innovative architectures and performance reliability and scalability improvements that advance business objectives.
6. Provides ongoing technical guidance coaching and training to other engineers delegating and reviewing work from lower-level technical professionals and raising the technical bar through design reviews and knowledge sharing.
7. Evaluates emerging technologies and techniques relevant to the job area building prototypes and solution concepts that contribute measurable input into new features products or capabilities.
8. Contributes to the development of long-term technical goals and plans for the area of responsibility through well-reasoned recommendations design proposals and implementation experience.
9. Leads large or complex initiatives within the job area coordinating and delegating technical work that may span outside the immediate team and ensuring cohesive high-quality outcomes with limited supervision.

Qualifications
Required Qualifications
The requirements listed below are representative of the knowledge skill and/or ability required. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions.
1. Bachelors degree in Computer Science Software Engineering or related field.
2. Minimum of 7 years of professional experience in software development.
3. Deep knowledge of multiple programming languages software architecture and design principles.
4. Deep understanding of software development lifecycle testing deployment and security practices.

Preferred Qualifications
1. Advanced degree in Computer Science or related technical discipline.
2. Professional certifications such as Certified Software Development Professional (CSDP) or equivalent.
3. Deep expertise in cloud-native architectures microservices container orchestration and DevOps.
4. Strong familiarity with Agile frameworks continuous integration/continuous deployment (CI/CD) and enterprise innovation management.

5. 7 years of experience in Site Reliability Engineering DevOps Platform Engineering or Infrastructure Operations.

6. Deep handson experience with distributed systems container orchestration (Kubernetes) and cloud-native operational tooling.

7. Proficiency with automation and scripting languages (Python Go PowerShell Ansible).

8. Strong understanding of observability platforms (Splunk Dynatrace) and event-driven monitoring.

9. Proven leadership in major incident management and cross-team technical coordination.

10. Strong grasp of networking Linux/Unix internals and modern infrastructure patterns.

11. Excellent communication skills including executive-level situational awareness during critical incidents.

12. Demonstrated ability to influence technical roadmaps and drive adoption of reliability best practices.

Preferred Qualifications

  • Financial services or regulated industry experience.

  • Experience enabling large-scale SRE transformations or modernization initiatives.

  • Familiarity with chaos engineering resilience assessments and service failure modeling.

  • Exposure to hybrid-cloud and multi-cloud operational frameworks.

  • Experience contributing to or leading Center for Enablement functions or Communities of Practice.

Key Responsibilities

Incident & Problem Management Leadership

  • Lead major and high-severity incident response effortsfocusing on diagnosing technical rootcausestherein and drivingmulti-team technical resolution.

  • Drive problem management to closure ensuring systemic fixes replace recurring operational risks.

  • Establish and maintain standardized incident playbooks escalation paths and communication frameworks.

Reliability Engineering & Automation

  • Architect and deliver automation solutions that eliminate toil reduce MTTR and increase service resilience.

  • Implement intelligent alerting anomaly detection and event correlation leveragingAI andAIOps tools.

  • Guide and enforce SLO/SLI adoption across product teams ensuring metrics inform decision-making and prioritization.

Observability & Operational Excellence

  • Enhance telemetry coverage across logs metrics traces and events using platforms such as Dynatrace and Splunk.

  • Define and standardize enterprise observability practices dashboards and KPIs.

  • Ensure operational readiness of applications and platforms through resiliency testing chaos engineering and failure-mode validation.

Cross-Functional Leadership & Influence

  • Partner with Delivery Architecture Security and Risk teams to embed reliability and resilience into design and execution.

  • Act as a change agent to elevate operational maturity and drive transformative improvements acrossWholesale.

  • Lead workshops maturity assessments and enablement sessions through the SRE C4E and Communities of Practice.

Standardization & Documentation

  • Develop maintain and enforce runbooks response playbooks and automated recovery patterns.

  • Contribute to enterprise SRE frameworks templates and maturity models.

  • Promote consistent adoption of best practices across domains and lines of business.

Mentorship & Technical Development

  • Coach and mentor Associate Professional and Senior SREs to build technical depth and operational discipline.

  • Provide thought leadership in SRE methodologies cloud-native operational patterns and automated reliability engineering.

For this opportunity Truist will not sponsor an applicant for work visa status or employment authorization nor will we offer any immigration-related support for this position (including but not limited to H-1B F-1 OPT F-1 STEM OPT F-1 CPT J-1 TN-1 or TN-2 E-3 O-1 or future sponsorship for U.S. lawful permanent residence status.)

Candidate must be willing to work onsite Monday - Friday at either office in Charlotte NC Raleigh NC or Atlanta GA.

General Description of Available Benefits for Eligible Employees of Truist Financial Corporation: All regular teammates (not temporary or contingent workers) working 20 hours or more per week are eligible for benefits though eligibility for specific benefits may be determined by the division of Truist offering the offers medical dental vision life insurance disability accidental death and dismemberment tax-preferred savings accounts and a 401k plan to teammates. Teammates also receive no less than 10 days of vacation (prorated based on date of hire and by full-time or part-time status) during their first year of employment along with 10 sick days (also prorated) and paid holidays. For more details on Truists generous benefit plans please visit our Benefits site. Depending on the position and division this job may also be eligible for Truists defined benefit pension plan restricted stock units and/or a deferred compensation plan. As you advance through the hiring process you will also learn more about the specific benefits available for any non-temporary position for which you apply based on full-time or part-time status position and division of work.

Truist is an Equal Opportunity Employer that does not discriminate on the basis of race gender color religion citizenship or national origin age sexual orientation gender identity disability veteran status or other classification protected by law. Truist is a Drug Free Workplace.

EEO is the Law E-Verify IER Right to Work


Required Experience:

IC


About Company

Company Logo

Your journey to better banking starts with Truist. Checking and savings accounts, credit cards, mortgages, small business, commercial banking, and more.

View Profile View Profile