Hiring.Camp

AI Platform Engineer (Cloud)

Absa

·

Today

Location
15 Alice Lane, South Africa
Workplace
Hybrid
Type
Full-time
Department
Engineering
Experience
2+ years
Education
Bachelor
Closing date
Today
Source
Workday

Description

Empowering Africa’s tomorrow, together…one story at a time.

With over 100 years of rich history and strongly positioned as a local bank with regional and international expertise, a career with our family offers the opportunity to be part of this exciting growth journey, to reset our future and shape our destiny as a proudly African group.

Job Summary

Absa Group’s Chief Data Analytics and Applied AI Office (CDAIO) requires an experienced and technically capable AI Platform Engineer (Cloud) to support the design, deployment, operation, and continuous improvement of the multi-cloud infrastructure powering the bank’s enterprise AI capability.

The role will contribute to the delivery of secure, scalable, reliable, and cost-effective AI platform services across multiple business units and countries. The platform supports AI use cases across Corporate and Investment Banking (CIB), Personal and Private Banking (PPB), Business Banking (BB), and Absa Regional Operations (AR).

The successful candidate will work across technologies such as AWS Bedrock, Databricks AI, Microsoft Azure AI Foundry, Hugging Face, Kubernetes, and GPU-based infrastructure. The role requires practical experience in cloud platform engineering, infrastructure-as-code, AI workload deployment, platform observability, cloud cost optimisation, security controls, and agentic AI infrastructure.

The role includes applying critical thinking, design thinking, and problem-solving skills within an agile engineering environment to address complex platform challenges. The AI Platform Engineer will work closely with senior engineers, architects, security teams, FinOps specialists, and AI Solution Engineers to deliver high-quality platform services in line with Absa’s architecture, risk, security, and responsible AI requirements.
The successful candidate will take accountability for assigned platform components and services while contributing to the broader performance, resilience, and user experience of the enterprise AI platform.

Job Description

Key Focus Areas

  • AI Platform Engineering and Architecture - Support the design, deployment, and operation of enterprise-grade, multi-cloud AI infrastructure across AWS Bedrock, Databricks AI, Microsoft Azure AI Foundry, Hugging Face, and GPU environments.
  • AI FinOps and Compute Cost Optimisation - Monitor AI infrastructure consumption, support cost allocation and reporting, and identify opportunities to optimise token usage, Databricks consumption, provisioned throughput, and GPU utilisation.
  • Platform Observability and Reliability - Implement and maintain monitoring, alerting, dashboards, and operational processes to ensure the availability, performance, and reliability of production AI platform services.
  • AI Security and Zero-Trust Controls - Implement security controls for AI platform APIs, model endpoints, data pipelines, and agentic AI services in line with Absa’s security architecture and regulatory requirements.
  • Agentic AI Infrastructure - Support the deployment and operation of infrastructure enabling AI agents, tool-calling services, autonomous workflows, agent memory, and orchestration frameworks.
  • Agile Engineering and Collaboration - Deliver platform enhancements through agile practices while collaborating with engineers, architects, business units, security teams, risk stakeholders, and third-party technology providers.

 

Accountabilities

Platform Engineering and Architecture

  • Support the design, deployment, configuration, and operation of Absa’s multi-cloud AI platform stack, including AWS Bedrock, Databricks AI, Microsoft Azure AI Foundry, Hugging Face, and GPU clusters.
  • Build and maintain reusable platform components such as AI Gateway configurations, model serving environments, vector databases, API integrations, data pipelines, and containerised workloads.
  • Develop and maintain infrastructure-as-code using technologies such as Terraform, Pulumi, AWS CDK, or equivalent tools.
  • Contribute to repeatable and auditable infrastructure deployments across multiple cloud environments, regions, and operating countries.
  • Configure and support agentic AI infrastructure, including orchestration environments, tool-calling APIs, agent memory, state management, and integration with enterprise systems.
  • Implement cloud-agnostic model serving patterns that improve workload portability across AWS, Azure, Databricks, and Kubernetes-based environments.
  • Support Kubernetes-based AI workloads using Docker, Kubernetes, and Helm.
  • Assist with the evaluation and implementation of new platform technologies, services, and engineering patterns.
  • Participate in architectural reviews, technical design sessions, peer reviews, and platform improvement initiatives.
  • Create and maintain architectural diagrams, configuration documentation, operational procedures, and technical standards.
  • Take accountability for the quality, performance, and operational readiness of assigned platform components.
  • Escalate complex architectural, security, capacity, and operational risks to senior engineers and platform leadership.

AI FinOps and Compute Cost Optimisation

  • Monitor and analyse AI platform consumption across Databricks, AWS, Azure, GPU infrastructure, and third-party services.
  • Support the development and maintenance of chargeback and showback frameworks for business units and individual AI use cases.
  • Assist with cost attribution for Databricks DBU consumption, AWS Bedrock token usage, Azure AI Foundry provisioned throughput, and GPU workloads.
  • Develop and maintain FinOps dashboards and cost reports using tools such as AWS Cost Explorer, Databricks System Tables, Azure Cost Management, and cloud-native monitoring services.
  • Contribute to monthly cost-per-use-case reporting for Finance, platform leadership, and business unit stakeholders.
  • Identify opportunities to optimise AI compute costs through workload scheduling, infrastructure right-sizing, token usage controls, caching, spot instances, and efficient model selection.
  • Support assessments of provisioned throughput versus on-demand consumption for production AI workloads.
  • Monitor spend anomalies and escalate unexpected usage, capacity, or budget risks.
  • Provide technical input into business cases and investment proposals for AI platform services.
  • Work closely with FinOps specialists and senior platform engineers to ensure infrastructure consumption remains within agreed budget parameters.

Platform Observability and SLA Engineering

  • Implement and maintain observability tooling for AI platform infrastructure and production AI services.
  • Build dashboards and alerts covering:
    • Inference latency
    • Platform availability
    • Token throughput
    • API gateway response times
    • Model endpoint health
    • GPU and compute utilisation
    • Databricks workload performance
    • Vector database performance
    • Capacity utilisation
    • Model drift indicators
  • Use tools such as Prometheus, Grafana, Datadog, OpenTelemetry, Databricks Lakehouse Monitoring, or equivalent technologies.
  • Support the implementation and monitoring of AI-specific service-level agreements and operational-level agreements.
  • Participate in incident response, troubleshooting, root-cause analysis, and post-incident reviews for AI platform failures.
  • Develop and maintain operational runbooks, support procedures, escalation paths, and recovery documentation.
  • Investigate platform performance issues and implement corrective or preventative actions.
  • Support release, change, and configuration management processes for AI platform components.
  • Conduct technical validation and operational readiness checks before platform changes are released into production.
  • Use performance and usage data to recommend improvements to platform scalability, resilience, reliability, and cost efficiency.
  • Contribute to initiatives focused on reducing incident volumes and mean time to recovery.

AI Security Architecture and Zero Trust

  • Implement zero-trust security controls for AI platform APIs, services, model endpoints, and agentic AI workloads.
  • Configure and maintain authentication and authorisation controls using:
    • OAuth 2.0
    • OpenID Connect
    • JWT, JWE, and JWS
    • Role-based access control
    • Attribute-based access control
    • Managed identities and service principals
  • Support the implementation of prompt injection prevention, output filtering, content controls, and data loss prevention mechanisms at the AI Gateway layer.
  • Implement controls to reduce the risk of unauthorised access, data exfiltration, insecure tool-calling, and excessive agent permissions.
  • Support data residency and sovereignty controls across Absa’s operating countries.
  • Work with architecture, security, risk, and legal stakeholders to ensure AI workloads comply with applicable data localisation and cross-border transfer requirements.
  • Contribute to AI-specific threat modelling covering model endpoints, agentic workflows, third-party AI providers, model supply chains, APIs, vector stores, and adversarial machine learning risks.
  • Remediate identified security vulnerabilities and configuration risks within agreed timelines.
  • Maintain platform documentation and evidence required for security reviews, audits, architecture approvals, and risk governance processes.
  • Apply Absa’s Enterprise-Wide Risk Management Framework, Group Architecture standards, information security requirements, and AI Responsible Use Policy in all engineering activities.

Agentic AI Infrastructure

  • Support the deployment and operation of agent orchestration technologies such as LangGraph, Microsoft Azure AI Foundry Agent Service, Amazon Bedrock Agents, AutoGen, or equivalent frameworks.
  • Configure infrastructure for agent tools, APIs, memory services, vector stores, workflow engines, and enterprise system integrations.
  • Implement secure tool-calling patterns, including identity propagation, permission controls, audit logging, timeout management, and failure handling.
  • Support agent state management, session persistence, memory controls, and multi-agent communication patterns.
  • Implement monitoring and tracing for agent execution paths, tool calls, latency, errors, and resource consumption.
  • Work with AI Solution Engineers to move agentic AI solutions from development into controlled, production-ready environments.
  • Contribute to platform standards for agent testing, deployment, monitoring, rollback, and lifecycle management.
  • Investigate and resolve infrastructure issues affecting the performance, security, or reliability of agentic AI workloads.

Agile Delivery and Capability Development

  • Participate actively in sprint planning, backlog refinement, daily stand-ups, technical demonstrations, and retrospectives.
  • Estimate engineering effort and deliver assigned platform features within agreed timelines and quality standards.
  • Collaborate with platform engineers, cloud engineers, AI Solution Engineers, architects, security specialists, data engineers, and business unit technology teams.
  • Participate in code reviews, infrastructure reviews, testing, troubleshooting, and technical problem-solving.
  • Contribute to platform engineering standards, reusable templates, automation libraries, and delivery accelerators.
  • Maintain comprehensive technical documentation, architectural decision records, deployment guides, and operational runbooks.
  • Share technical knowledge and provide guidance to junior engineers and other members of the engineering community.
  • Support the development of platform onboarding materials, self-service documentation, and user guides for business unit technology teams.
  • Proactively identify technical dependencies, delivery risks, and operational barriers and escalate these appropriately.
  • Remain current with developments in cloud AI platforms, agentic AI, MLOps, FinOps, AI security, and platform engineering.

 

Qualifications and Experience

Education and Qualifications

  • Bachelor’s degree in Computer Science, Information Technology, Data Science, Mathematics, Statistics, Engineering, or a related quantitative discipline is essential.
  • A postgraduate qualification is advantageous.
  • Relevant practical experience may be considered where supported by a strong record of cloud and platform engineering delivery.

Advantageous Certifications

One or more of the following certifications would be advantageous:

Cloud

  • AWS Certified Solutions Architect
  • AWS Certified Machine Learning Engineer
  • Microsoft Certified: Azure AI Engineer Associate
  • Microsoft Certified: Azure Solutions Architect Expert
  • Databricks Certified Data Engineer or Machine Learning certification

Infrastructure-as-Code

  • HashiCorp Certified: Terraform Associate
  • Equivalent Terraform, Pulumi, or cloud infrastructure certification

FinOps

  • FinOps Certified Practitioner
  • Equivalent cloud cost management or financial operations certification

Security

  • Certified Cloud Security Professional
  • AWS Certified Security
  • Microsoft Security, Compliance, and Identity certification
  • Equivalent cloud or cybersecurity certification

 

Work Experience

  • Approximately 4 to 6 years of relevant experience in cloud engineering, platform engineering, DevOps, MLOps, infrastructure engineering, or AI platform engineering.
  • At least 2 years of practical experience supporting cloud-based data, machine learning, generative AI, or AI platform workloads in a production environment.
  • Production experience with at least two of the following:
    • AWS Bedrock or Amazon SageMaker
    • Databricks
    • Microsoft Azure AI Foundry or Azure Machine Learning
    • Hugging Face
    • Kubernetes-based model serving
  • Practical infrastructure-as-code experience using Terraform, Pulumi, AWS CDK, or an equivalent technology.
  • Experience building or supporting CI/CD pipelines for cloud infrastructure, platform components, data services, or machine learning workloads.
  • Experience with Docker, Kubernetes, Helm, APIs, identity integration, and cloud-native platform services.
  • Experience implementing monitoring, dashboards, alerts, and operational support processes for production platforms.
  • Working knowledge of cloud cost management, cost allocation, capacity monitoring, and infrastructure optimisation.
  • Experience applying cloud security controls, identity and access management, secrets management, and secure API integration.
  • Experience working within enterprise risk, architecture, security, and change management processes.
  • Experience in financial services, telecommunications, healthcare, insurance, or another regulated industry is advantageous.

 

Knowledge and Skills

  • Multi-Cloud AI Platform Engineering - Practical knowledge of designing, deploying, and supporting AI services across AWS, Microsoft Azure, Databricks, Hugging Face, or Kubernetes-based environments.
  • Agentic AI Infrastructure - Working knowledge of agent orchestration frameworks, tool-calling API patterns, agent memory, state management, tracing, and multi-agent workflows.
  • AI FinOps and Cost Management - Knowledge of cloud consumption models, token-based pricing, Databricks DBUs, provisioned throughput, GPU utilisation, chargeback and showback reporting, and spend anomaly detection.
  • AI Security and Zero Trust - Working knowledge of OAuth 2.0, OIDC, JWT, RBAC, ABAC, API security, managed identities, secrets management, prompt injection controls, data loss prevention, and secure agent tool access.
  • Infrastructure-as-Code - Strong practical experience with Terraform, Pulumi, AWS CDK, or equivalent infrastructure automation technologies.
  • Containerisation and Orchestration - Experience with Docker, Kubernetes, Helm, container registries, workload scheduling, resource allocation, and production container operations.
  • Platform Observability - Experience with Prometheus, Grafana, Datadog, OpenTelemetry, cloud-native monitoring tools, or Databricks Lakehouse Monitoring.
  • Cloud-Agnostic Model Serving - Working knowledge of containerised model deployment and serving technologies such as ONNX, BentoML, Triton Inference Server, Kubernetes, or equivalent frameworks.
  • MLOps Tooling - Working knowledge of MLflow, Kubeflow, Airflow, model registries, feature stores, automated testing, and CI/CD for machine learning workloads.
  • GPU Infrastructure - Understanding of GPU workload deployment, capacity management, right-sizing, spot instance strategies, and cost optimisation for model training and inference.
  • Enterprise Risk and Governance - Working knowledge of information security, technology risk, architecture governance, responsible AI, privacy, data residency, and change management requirements within a regulated environment.
  • Agile Delivery - Experience working in agile engineering teams using sprint planning, backlog management, iterative delivery, peer review, testing, and continuous improvement practices.

 

Education

Bachelor's Degree: Information Technology

Absa Bank Limited is an equal opportunity, affirmative action employer. In compliance with the Employment Equity Act 55 of 1998, preference will be given to suitable candidates from designated groups whose appointments will contribute towards achievement of equitable demographic representation of our workforce profile and add to the diversity of the Bank.

Absa Bank Limited reserves the right not to make an appointment to the post as advertised

Skills

AWSAzureDockerKubernetesTerraformCI/CDMachine LearningAirflowDatabricksData ScienceCybersecurityOAuthAgileDevOpsRisk ManagementComplianceLoss PreventionChange ManagementAWS Certified

Similar Jobs

30

AI Platform Engineer

Job Openings · Highland Heights, OH

Today

AI Platform Engineer

Qube Research & Technologies · Mumbai +1

3 days ago

AI Platform Engineer

AstraZeneca is · Spain - Barcelona

4 days ago

AI Platform Engineer

CreateFuture · Leeds, Manchester, Edinburgh, London +4

6 days ago

AI Platform Engineer

EXL Talent Acquisition Team · Dublin, Leinster, Ireland, IE · Hybrid

1 week ago

AI Platform Engineer

EXL · Dublin, Leinster, Ireland, IE · Hybrid

1 week ago

AI Platform Engineer

Supabase · Remote · Remote

1 week ago

AI Platform Engineer

Dialpad · Buenos Aires, Argentina +1

2 weeks ago

AI Platform Engineer

United States Cold Storage · Camden, NJ +1

2 weeks ago

AI Platform Engineer

"Cornelis Networks, Inc."

2 weeks ago

AI Platform Engineer

Massmutual · Boston - 10 Fan Pier Blvd, United States of America +2 · Hybrid

2 weeks ago

Engineer - AI Platform

Ebury · Madrid · Hybrid

2 weeks ago

AI Platform Engineer

Greenberg Traurig · Atlanta Center of Excellence, United States of America +10 · Hybrid

3 weeks ago

AI Platform Engineer

KLA · USA-CA-Milpitas-KLA, United States of America

4 weeks ago

AI Platform Engineer

Uchicago · 6045 Kenwood Building, United States of America · Hybrid

1 month ago

AI Platform Engineer

Givzey · US · Remote

1 month ago

AI Platform Engineer

Wilhelmsen is now · Kuala Lumpur - Level 19, 1 Sentral, Malaysia

1 month ago

AI Platform Engineer

Paypaycard · Hybrid · Remote, Hybrid

1 month ago

AI Platform Engineer

GlobeMed Group · Beirut, Lebanon

1 month ago

AI Platform Engineer

Notable · San Mateo, CA · Hybrid

2 months ago

AI Platform Engineer

Adobe · Bucharest, Romania

2 months ago

AI Platform Engineer

RSC2 · Hanover, MD · Hybrid

2 months ago

AI Platform Engineer

Quberesearchandtechnologies · Wrocław +1

2 months ago

AI Platform Engineer

Ebay · Bangalore, India · Hybrid

3 months ago

AI Platform Engineer

Ebay · Bangalore, India · Hybrid

3 months ago

AI Platform Engineer

Ebay · Bangalore, India · Hybrid

3 months ago

AI Platform Engineer

Xenergy · MD - Gaither Rd., Rockville Corp Hqtrs, United States of America

3 months ago

AI Platform Engineer

Aalo Atomics · Austin, TX

3 months ago

AI Platform Engineer

Alliander · HAARLEM, Netherlands

3 months ago

AI Platform Engineer

Great American Insurance Group · OH Cin-Dixie Terminal So, United States of America · Hybrid

3 months ago
AI Platform Engineer (Cloud) at Absa | Hiring.Camp