AI Operations Engineer
Collinson is the global, privately-owned company dedicated to helping the world to travel with ease and confidence. The group offers a unique blend of industry and sector specialists who together provide market-leading airport experiences, loyalty and customer engagement, and insurance solutions for over 400 million consumers.
Collinson is the operator of Priority Pass, the world’s original and leading airport experiences programme. Travellers can access a network of 1,500+ lounges and travel experiences, including dining, retail, sleep and spa, in over 650 airports in 148 countries, helping to elevate the journey into something special. We work with the world’s leading payment networks, over 1,400 banks, 90 airlines and 20 hotel groups worldwide.
We have been bringing innovation to the market since inception – from launching the first independent global VIP lounge access Programme, Priority Pass to being the first to sell direct travel insurance in the UK through Columbus Direct and creating the first loyalty agency of its kind in the travel sector with ICLP. Today we still invest heavily in innovation to ensure that we continue to deliver superior customer experiences.
Key clients include Mastercard, American Express, Cathay Pacific, British Airways, LATAM, Flying Blue, Accor, EasyJet, HSBC, Chase, HDFC.
Our mission is focused on doing good beyond profit, which for us means we seek out opportunities for our people to share in our success and that we give back to the communities and people within which we work.
Never short of ambition, the success of our business is delivered through the diverse and talented team of over 2,200 global colleagues.
Purpose of the role
As an AI Operations Engineer, you will help us run AI-powered workflows, applications and agents reliably and securely in production. You will take an operational engineering perspective across the service lifecycle, helping new AI capabilities move safely into production and ensuring they remain dependable once live.
You will work closely with the Hyperautomation Lead and AI Engineers, who design and build AI-enabled solutions. Your role is to make sure those solutions are production-ready and supportable, with appropriate deployment, environments, access, monitoring, resilience and recovery arrangements.
This is a hands-on engineering role. You will configure cloud services, automate deployments, manage environments and access, implement monitoring, troubleshoot production issues and improve service reliability over time. Because AI services often depend on multiple platforms, APIs and enterprise systems, you will also help coordinate operational dependencies across Technology, Security and other teams.
You do not need to be an AI model developer. We are looking primarily for strong production engineering skills combined with an interest in how AI-enabled applications behave in real-world environments.
Key Responsibilities
· Own the operational health of AI services. Help establish clear support arrangements, dependencies and operational standards for AI-powered workflows, applications and agents, and ensure services remain reliable, secure and supportable once live.
· Make new services production-ready. Work alongside the Hyperautomation Lead and AI Engineers to ensure new and changed services have appropriate deployment, monitoring, access controls, failure handling, recovery, support arrangements and documentation before go-live.
· Engineer deployment and cloud environments. Build and maintain repeatable deployment and environment patterns using CI/CD, infrastructure-as-code and configuration management, and configure the cloud services, identity, secrets and connectivity needed to operate AI services securely.
· Implement observability and improve reliability. Create and maintain useful logs, metrics, dashboards, alerts and health checks. Use telemetry, incidents and recurring operational issues to improve resilience, error handling, recovery and automation and to reduce manual operational effort.
· Manage incidents and operational problems. Act as a technical responder for production incidents, troubleshoot issues across cloud, application, identity, integration and workflow layers, and coordinate service restoration where other teams or vendors are involved. Contribute to root-cause analysis and ensure appropriate corrective actions are identified and followed through.
· Operate integrations and technical dependencies. Maintain visibility over APIs, connectors, authentication and external platform dependencies, troubleshoot integration issues and work with internal teams to resolve problems that affect production services.
· Maintain strong operational documentation. Produce and keep up-to-date technical documentation, runbooks, support procedures, recovery processes and operational playbooks covering how services are deployed, monitored, supported and restored.
· Implement operational governance and controls. Ensure appropriate security, access, auditability, data-handling and AI operational controls are built into production services, including permissions boundaries, logging, approval controls and recovery mechanisms where appropriate.·
Coordinate technically with platform providers and vendors. Act as an operational and technical counterpart to providers such as AWS, Microsoft, Salesforce, Anthropic and Mindflow, coordinating technical support and escalations and assessing platform changes that may affect reliability, security or supportability.
· Continuously improve how AI services are operated. Feed operational learning back to the Hyperautomation Lead, AI Engineers and wider Technology teams, help improve reusable operational patterns and remove recurring sources of support effort or production risk.
Knowledge, skills and experience required
Must-have
We are looking for someone with strong production engineering and operational ownership experience. You do not need to be an AI model developer.
Production cloud engineering
Hands-on experience operating applications or services in AWS, Azure or a comparable cloud environment, including configuration, troubleshooting, access, monitoring and production support.
Deployment, infrastructure and automation
Practical experience with:
· CI/CD and Git-based delivery practices
· Infrastructure-as-code such as Terraform, CloudFormation, CDK or equivalent
· Environment and configuration management
· Automating repeatable operational tasks
Programming, APIs and integrations
Practical programming or scripting experience using Python, TypeScript/JavaScript, Bash, PowerShell or similar, with the ability to automate tasks and diagnose production issues.
Good working knowledge of APIs, HTTP, JSON, authentication and system integrations, including troubleshooting dependencies across different systems.
Observability, incident and problem management
Experience:
· Working with production logs, metrics, dashboards and alerts
· Troubleshooting live services
· Responding to and coordinating production incidents
· Contributing to root-cause analysis and corrective actions
· Improving monitoring and reliability based on operational experience
You should be comfortable working through incidents that span several systems, teams or external providers rather than only troubleshooting a single application.
Security, access and operational controls
Good understanding of:
· Identity and access management
· Secrets and credential management
· Least-privilege access
· Secure configuration
· Auditability and operational controls
You should be comfortable applying security and governance requirements as part of normal production engineering rather than treating them as a separate activity.
Operational ownership and documentation
Experience taking responsibility for how production services are operated and supported, including creating and maintaining:
· Runbooks and operational playbooks
· Recovery and troubleshooting procedures
· Deployment and support documentation
· Service dependencies and escalation paths
Strong written documentation skills are important for this role.
Cross-team and vendor technical coordination
Ability to work effectively with engineers, infrastructure and security teams, business stakeholders and external technology providers.
You should be comfortable:
· Coordinating technical resolution where several teams are involved
· Working with vendors during incidents or technical escalations
· Explaining technical risks and operational issues clearly
· Following issues and corrective actions through to resolution
Nice to have
Experience in any of the following would be useful, but is not required:
· Operating AI-enabled applications, LLM applications, agents or workflow automation in production
· Understanding AI-specific operational considerations such as model/API dependencies, rate limits, latency, retries, tool calls, structured outputs and failure handling
· Mindflow, n8n or similar workflow/orchestration platforms
· Amazon Bedrock, Anthropic, OpenAI or other model platforms
· Microsoft Entra, AWS IAM or Salesforce
· CloudWatch, DataDog or similar observability tooling
· Containers or Kubernetes
· Distributed tracing and SRE practices such as SLIs and SLOs
· MCP, RAG, agentic AI or tool/function calling
· Secure cloud networking for enterprise integrations
· Building reusable deployment templates, monitoring patterns, infrastructure modules or operational tooling
Candidates from Site Reliability Engineering, DevOps, Platform Engineering, Cloud Engineering, Production Engineering or MLOps backgrounds may be particularly well suited to the role.
We welcome strong production engineers who may not yet have extensive AI operations experience but are interested in developing that capability.
Collinson is an equal opportunity employer and welcomes differences in all their forms including: colour, race, ethnicity, gender identity, sexual orientation, neurodivergence, family status, age, individuals with disabilities and people from all backgrounds, cultures and experiences as we strongly believe this contributes to our on-going success.
We are focused on continually evolving our purpose driven, high performing culture, providing an environment where our people have the opportunity to achieve their full potential and do interesting and meaningful work. Our company values are: Take Action, Do the right thing, One team and Be insight led. These help guide everything we do internally in terms of how we think, act and interact, right through to how we deliver value to our customers and clients.
In your application, please feel free to note which pronouns you use (For example - she/her/hers, he/him/his, they/them/theirs, etc).
If you need any extra support throughout the interview process, then please email us at ukrecruitment@collinsongroup.com
- Division
- Technology & Data
- Locations
- Cape Town
- Remote status
- Hybrid
About Collinson
We use our expertise and products to craft customer experiences. Our range of services helps global brand acquire, engage and retain choice-rich customers.
© 2023 Collinson International Limited. Registered in England & Wales under registration No. 2577557
Registered address : 3 More London Riverside, London, SE1 2AQ, United Kingdom.