Cloud Engineer ( Azure Platform Engineer) Remote Ontario
Job Summary
The Azure Platform Engineer is responsible for designing deploying and operating production-grade Azure Kubernetes Service (AKS) clusters with a strong focus on cluster internals networking security hardening and reliability. This role owns the platform end-to-end from cluster architecture and CI/CD through to performance tuning observability and incident response for critical systems. Acting as the SME for Kubernetes deployments blockers and production issues. The successful candidate works closely with application security and platform teams to keep AKS environments reliable observable and production-ready.
- Act as the subject-matter expert (SME) for Kubernetes deployments troubleshooting and production issues across all environments.
- Design and deploy Azure Kubernetes Service (AKS) clusters with private cluster configurations managed identities and RBAC.
- Own the health scaling and lifecycle management of production AKS clusters including upgrades node pool management autoscaling and capacity planning.
- Configure AKS networking including Azure CNI internal load balancers and ingress controllers (NGINX Traefik).
- Design and maintain integrations between AKS with Azure Container Registry (ACR) Key Vault via CSI driver Azure Monitor for containers Azure SQL Kafka/Event Hubs Azure Storage and other client-facing dependencies (DNS resolution firewall rules private endpoints and service connectivity).
- Manage containerized application deployments using Docker and Helm; maintain reusable chart and templating standards namespaces resource quotas and Azure Policy for AKS.
- Harden AKS environments through policy enforcement network policies and image scanning.
- Own container and cluster vulnerability management: scanning triage prioritization and remediation coordination with engineering teams.
- Contribute to Terraform-based infrastructure as code for provisioning and managing Azure resources.
- Support Azure DevOps (or equivalent) CI/CD pipelines including GitOps workflows (Flux/ArgoCD) to reduce deployment risk and improve release velocity.
- Design implement and manage observability across the Grafana stack (Prometheus Loki Tempo) and Azure-native tooling (Azure Monitor Log Analytics Workspace Application Insights) for critical systems; define SLIs/SLOs and tune alerting to reduce noise.
- Collaborate with performance engineering and application teams to identify diagnose and resolve performance bottlenecks spanning the AKS platform and its dependent services (database messaging network) to right-size node pools and workloads based on observed performance and utilization trends.
- Contributing to Disaster Recovery and Business Continuity Planning (DR/BCP) procedures for AKS-hosted workloads including cross-team failover drill participation and RTO/RPO validation.
- Provide escalation support for production Kubernetes and infrastructure incidents; participate in on-call rotation lead root-cause analysis and drive preventative follow-up actions.
- Document runbooks and post-incident reviews and operational knowledge; maintain a living knowledge base to reduce tribal knowledge.
- 5 years of hands-on Kubernetes experience including at least 2 years running AKS in production.
- Strong knowledge of Kubernetes internals scheduling networking storage RBAC.
- Proficiency in Azure CNI networking and AKS private cluster configuration.
- Hands-on experience integrating AKS with Azure PaaS services (ACR AKV Azure SQL Kafka and Managed Identities) and troubleshooting network-layer dependencies (DNS firewall private endpoints).
- Experience with Helm and GitOps workflows (Flux/ArgoCD).
- Working knowledge of Terraform for infrastructure as code.
- Working knowledge of Azure DevOps or similar CI/CD tooling.
- Hands-on experience implementing and operating observability platforms: Grafana stack (Prometheus Loki Tempo) and Azure-native monitoring (Azure Monitor Log Analytics Application Insights); experience defining SLIs/SLOs for production systems.
- Practical experience with container/cluster vulnerability management and remediation workflows.
- Scripting skills in Bash Python.
- Certified Kubernetes Administrator (CKA) Or CKAD certification and Microsoft Certified: Azure Administrator (AZ-104) required.
- Excellent written and verbal communication skills; able to convey technical issues clearly to both technical and non-technical stakeholders.
This position is a new role created to support Smiles continued growth and commitment to operational excellence.
Required Experience:
IC
About Company
Built around the visionary HL7 FHIR standard and powered by HAPI, Smile Digital Health is the most proven FHIR implementation in the world.