SRE Architect
Job Summary
We fuse together exceptional talent who deliver outstanding software solutions. Our approach has helped us grow 60% in 2021 94% in 2022 while in 2023 we joined forces with Insight a Fortune 500 company and a leading solutions and systems integrator. With exciting growth plans and cutting-edge projects there has never been a better time to join our incredible team.
SRE Architect / Principal Engineer
About the role
We are looking for a highly experiencedSRE Architect / Principal Engineerto help lead the infrastructure and reliability architecture forthetechnology landscape as we accelerate our modernisation journey.
Our broader architectural direction is already taking shape including DDDmicrofrontendsand Backend-for-Frontend (BFF) while our infrastructure direction is standardising around AWS Terraform GitHub Actions and ECS/Fargate. However important decisionsremainaround areas such as gateways and routing scalability observability resilience deployment architecture and how existing systems should progressively move toward the target state.
Your primary responsibility will be toconsolidatethis direction into a coherent infrastructure and reliability architecture and define a pragmatic path for its adoption.
You willoperatebetween Enterprise Architecture and SRE/engineering teams translating broader architectural direction into practical patterns standards reference implementations and modernisation strategies.
This is ahands onPrincipal level individual contributor rolewith significant technical influence acrossinfrastructure. You will be expected to challenge existing decisions whereappropriatevalidateimportant architectural choices through proofs of concept and reference implementations and provide the technical direction that enables engineering teams to implement and adopt the target architecture successfully.
Whatyoullwork on
Infrastructure modernisation & target architecture
- Define and evolve the target infrastructure and reliability architecture.
- Consolidatearchitectural decisions already underway into a coherent scalable and maintainable target state.
- Define pragmatic modernisation strategies for existing systems balancing business value technical risk cost and migration effort.
- Assess systems and recommend whether they should be incrementally modernised aligned with the target architecture temporarilyretained or eventually replaced.
- Define transition patterns that allow teams to modernise without unnecessarylarge scalerewrites.
- Establish the target architecture and infrastructure patterns as the default for new modules.
- Identifyarchitectural gaps risks and cross-system dependencies across landscape.
AWS cloud & platform architecture
- Define scalable resilient secure and cost-conscious architectures using AWS.
- Define approaches to service-to-service communication external API exposure routing and gateway strategy.
- Guide architectural decisions around scalability workload and capacity management fault isolation availability and graceful degradation.
- Defineappropriate environment networking and deployment strategies for services.
- Work closely with security and enterprise platform teams to ensure alignment with Pearson-wide standards.
Infrastructure as Code & CI/CD
- Establish Infrastructure as Code standards using Terraform including reusable patterns modules and conventions that engineering teams can adopt consistently.
- Shape CI/CD architecture using GitHub and GitHub Actions as the standard delivery platform.
- Define reusable deployment patterns and approaches to environment promotion rollback and safe releases.
- Reduce infrastructure and CI/CD divergence between engineering teams through reusable standards and automation.
Reliability observability & production readiness
- Shape and evolve wide approaches to reliability resilience observability and production readiness.
- Shape standards for metrics logs traces dashboards and alerting across distributed systems.
- Helpestablishmeaningful SLIs SLOs and reliability targets whereappropriate.
- Guide architectural approaches to disaster recovery failure handling backups and recovery strategies.
- Help teams design systems thatremainoperable and cost-effective as usage and complexity grow.
Architecture standards & AI-enabled engineering
- Translate infrastructure and SRE architecture decisions into reusable standards reference implementations Terraform patterns and CI/CD practices.
- Collaborate with teams evolving AI enabled engineering framework so agreed infrastructure patterns and guardrails can be incorporated into engineering workflows.
- Create clear architecture decision records reference architectures and implementation guidance that teams can apply consistently.
Technical leadership
- Act as a senior technical authority for SRE and infrastructure architecture.
- Operate as a bridge between Enterprise Architects and the SRE and engineering teams responsible for implementation.
- Validate important architectural decisions through proofs of concept and reference implementations whereappropriate.
- Review major infrastructure designs and provide technical direction across teams.
- Mentor senior engineers Tech Leads and SREs and help drive alignment on cross-team technical decisions.
Required skills & experience
Architecture & technical leadership
- Extensive professional experience designing andoperatinglarge-scale distributed systems in production.
- Proven experienceoperatingatSREArchitect / Principal Engineeror equivalent senior technical leadership level.
- Strongtrack recorddefining cloud and infrastructure architecture across multiple teams or services.
- Experience leading or shaping modernisation across technology estatescontainingboth legacy and modern systems.
- Ability to define a target architecture while creating realistic incremental migration paths toward it.
- Strong architectural judgment and the ability to balance technical quality delivery speed risk cost and organisational constraints.
- Strong technical depth and willingness to create prototypes or reference implementations when needed tovalidatearchitectural decisions.
AWS infrastructure & delivery
- Deepexpertisein AWS and cloud-native architecture.
- Strong experience designing andoperatingcontainerised workloads on AWS particularly ECS/Fargate.
- Strong understanding of AWS networking load balancing routing and secure connectivity.
- Experience designing scalable API gateway and ingress architectures.
- Strong experience with relational data platforms such as Amazon RDS/Aurora.
- Good understanding of event-driven and asynchronous architecture patterns.
- Strong experience with Terraform in production environments.
- Strong experience with GitHub and GitHub Actions including reusable CI/CD workflows and deployment automation.
- Strong understanding of cloudnative autoscaling workload capacity management and scaling strategies for containerised and event-driven systems.
SRE & reliability engineering
- Strong understanding of Site Reliability Engineering principles.
- Experience designing observability approaches for distributed production systems.
- Strong understanding of metrics logging tracing alerting and production monitoring.
- Experience working with SLIs SLOs reliability targets and production-readiness practices.
- Strong knowledge of resilience patterns failure modes scalability and disaster recovery.
- Ability to reason about trade-offs between reliability performance complexity and cost.
Security
- Strong understanding of AWS security fundamentals including IAM secrets management encryption and secure infrastructure patterns.
- Ability to work effectively with specialist security teams and translate organisational requirements into practical engineering approaches.
General & soft skills
- Strong technical leadership skills without relying on formal line-management authority.
- Ability to influence engineering direction across multiple teams and organisational boundaries.
- Comfortable challenging existing approaches constructively and explaining the trade-offs behind alternative solutions.
- Pragmatic approach to modernisation with the judgment to distinguish valuable change from unnecessary technical churn.
- Excellent communication skills with engineers Tech Leads Engineering Managers Enterprise Architects and other stakeholders.
- Ability to turn ambiguous technical problems into clear decisions and actionable next steps.
Nice to have
- Experience with Backend-for-Frontend andmicrofrontendarchitectures.
- Experience with event-driven architectures and Kafka/MSK.
- Experience withOpenTelemetryand modern observability platforms.
- Experience with progressive delivery approaches such as canary or blue/green deployments.
- Experience with cloud cost optimisation or FinOps practices.
- Experience defining architecture standards paved roads or engineering golden paths across multiple autonomous teams.
- Experience working in EdTech digital learning or another large global technology organisation.
To see more roles click here.
Required Experience:
Staff IC
About Company
Unleash limitless possibilities with Amdaris, the premier provider of Extended Delivery Teams! Software Engineering, Product Design, Data Science & Engineering and beyond.