Enter a job title or keyword

AI Operations Engineer

Collinson


Job Location:

Cape Town - South Africa

Monthly Salary: Not provided by the employer
Posted: 16 September 2026 (19 hours ago)
Application Deadline: 14 December 2026
Vacancies: 1 Vacancy

Job Summary

Collinson is the global privately-owned company dedicated to helping the world to travel with ease and confidence. The group offers a unique blend of industry and sector specialists who together provide market-leading airport experiences loyalty and customer engagement and insurance solutions for over 400 million consumers.

Collinson is the operator of Priority Pass the worlds original and leading airport experiences programme. Travellers can access a network of 1500 lounges and travel experiences including dining retail sleep and spa in over 650 airports in 148 countries helping to elevate the journey into something special. We work with the worlds leading payment networks over 1400 banks 90 airlines and 20 hotel groups worldwide.

We have been bringing innovation to the market since inception from launching the first independent global VIP lounge access Programme Priority Pass to being the first to sell direct travel insurance in the UK through Columbus Direct and creating the first loyalty agency of its kind in the travel sector with ICLP. Today we still invest heavily in innovation to ensure that we continue to deliver superior customer experiences.

Key clients include Mastercard American Express Cathay Pacific British Airways LATAM Flying Blue Accor EasyJet HSBC Chase HDFC.

Our mission is focused on doing good beyond profit which for us means we seek out opportunities for our people to share in our success and that we give back to the communities and people within which we work.

Never short of ambition the success of our business is delivered through the diverse and talented team of over 2200 global colleagues.

Purpose of the role


As an AI Operations Engineer you will help us run AI-powered workflows applications and agents reliably and securely in production. You will take an operational engineering perspective across the service lifecycle helping new AI capabilities move safely into production and ensuring they remain dependable once live.

You will work closely with the Hyperautomation Lead and AI Engineers who design and build AI-enabled solutions. Your role is to make sure those solutions are production-ready and supportable with appropriate deployment environments access monitoring resilience and recovery arrangements.

This is a hands-on engineering role. You will configure cloud services automate deployments manage environments and access implement monitoring troubleshoot production issues and improve service reliability over time. Because AI services often depend on multiple platforms APIs and enterprise systems you will also help coordinate operational dependencies across Technology Security and other teams.

You do not need to be an AI model developer. We are looking primarily for strong production engineering skills combined with an interest in how AI-enabled applications behave in real-world environments.

Key Responsibilities

Own the operational health of AI services. Help establish clear support arrangements dependencies and operational standards for AI-powered workflows applications and agents and ensure services remain reliable secure and supportable once live.

Make new services production-ready. Work alongside the Hyperautomation Lead and AI Engineers to ensure new and changed services have appropriate deployment monitoring access controls failure handling recovery support arrangements and documentation before go-live.

Engineer deployment and cloud environments. Build and maintain repeatable deployment and environment patterns using CI/CD infrastructure-as-code and configuration management and configure the cloud services identity secrets and connectivity needed to operate AI services securely.

Implement observability and improve reliability. Create and maintain useful logs metrics dashboards alerts and health checks. Use telemetry incidents and recurring operational issues to improve resilience error handling recovery and automation and to reduce manual operational effort.

Manage incidents and operational problems. Act as a technical responder for production incidents troubleshoot issues across cloud application identity integration and workflow layers and coordinate service restoration where other teams or vendors are involved. Contribute to root-cause analysis and ensure appropriate corrective actions are identified and followed through.

Operate integrations and technical dependencies. Maintain visibility over APIs connectors authentication and external platform dependencies troubleshoot integration issues and work with internal teams to resolve problems that affect production services.

Maintain strong operational documentation. Produce and keep up-to-date technical documentation runbooks support procedures recovery processes and operational playbooks covering how services are deployed monitored supported and restored.

Implement operational governance and controls. Ensure appropriate security access auditability data-handling and AI operational controls are built into production services including permissions boundaries logging approval controls and recovery mechanisms where appropriate.

Coordinate technically with platform providers and vendors. Act as an operational and technical counterpart to providers such as AWS Microsoft Salesforce Anthropic and Mindflow coordinating technical support and escalations and assessing platform changes that may affect reliability security or supportability.

Continuously improve how AI services are operated. Feed operational learning back to the Hyperautomation Lead AI Engineers and wider Technology teams help improve reusable operational patterns and remove recurring sources of support effort or production risk.

Knowledge skills and experience required

Must-have

We are looking for someone with strong production engineering and operational ownership experience. You do not need to be an AI model developer.

Production cloud engineering

Hands-on experience operating applications or services in AWS Azure or a comparable cloud environment including configuration troubleshooting access monitoring and production support.

Deployment infrastructure and automation

Practical experience with:

CI/CD and Git-based delivery practices

Infrastructure-as-code such as Terraform CloudFormation CDK or equivalent

Environment and configuration management

Automating repeatable operational tasks

Programming APIs and integrations

Practical programming or scripting experience using Python TypeScript/JavaScript Bash PowerShell or similar with the ability to automate tasks and diagnose production issues.

Good working knowledge of APIs HTTP JSON authentication and system integrations including troubleshooting dependencies across different systems.

Observability incident and problem management

Experience:

Working with production logs metrics dashboards and alerts

Troubleshooting live services

Responding to and coordinating production incidents

Contributing to root-cause analysis and corrective actions

Improving monitoring and reliability based on operational experience

You should be comfortable working through incidents that span several systems teams or external providers rather than only troubleshooting a single application.

Security access and operational controls

Good understanding of:

Identity and access management

Secrets and credential management

Least-privilege access

Secure configuration

Auditability and operational controls

You should be comfortable applying security and governance requirements as part of normal production engineering rather than treating them as a separate activity.

Operational ownership and documentation

Experience taking responsibility for how production services are operated and supported including creating and maintaining:

Runbooks and operational playbooks

Recovery and troubleshooting procedures

Deployment and support documentation

Service dependencies and escalation paths

Strong written documentation skills are important for this role.

Cross-team and vendor technical coordination

Ability to work effectively with engineers infrastructure and security teams business stakeholders and external technology providers.

You should be comfortable:

Coordinating technical resolution where several teams are involved

Working with vendors during incidents or technical escalations

Explaining technical risks and operational issues clearly

Following issues and corrective actions through to resolution

Nice to have

Experience in any of the following would be useful but is not required:

Operating AI-enabled applications LLM applications agents or workflow automation in production

Understanding AI-specific operational considerations such as model/API dependencies rate limits latency retries tool calls structured outputs and failure handling

Mindflow n8n or similar workflow/orchestration platforms

Amazon Bedrock Anthropic OpenAI or other model platforms

Microsoft Entra AWS IAM or Salesforce

CloudWatch DataDog or similar observability tooling

Containers or Kubernetes

Distributed tracing and SRE practices such as SLIs and SLOs

MCP RAG agentic AI or tool/function calling

Secure cloud networking for enterprise integrations

Building reusable deployment templates monitoring patterns infrastructure modules or operational tooling

Candidates from Site Reliability Engineering DevOps Platform Engineering Cloud Engineering Production Engineering or MLOps backgrounds may be particularly well suited to the role.

We welcome strong production engineers who may not yet have extensive AI operations experience but are interested in developing that capability.

Collinson is an equal opportunity employer and welcomes differences in all their forms including: colour race ethnicity gender identity sexual orientation neurodivergence family status age individuals with disabilities and people from all backgrounds cultures and experiences as we strongly believe this contributes to our on-going success.

We are focused on continually evolving our purpose driven high performing culture providing an environment where our people have the opportunity to achieve their full potential and do interesting and meaningful work. Our company values are: Take Action Do the right thing One team and Be insight led. These help guide everything we do internally in terms of how we think act and interact right through to how we deliver value to our customers and clients.

In your application please feel free to note which pronouns you use (For example - she/her/hers he/him/his they/them/theirs etc).

If you need any extra support throughout the interview process then please email us at


About Company

Company Logo

We use our expertise and products to craft customer experiences. Our range of services helps global brand acquire, engage and retain choice-rich customers.© 2023 Collinson International Limited. Registered in England & Wales under registration No. 2577557Registered address : 3 More Lo ... View more

View Profile View Profile