Principal AI Platform Engineer
Indexed description
Principal AI Platform Engineer, Cloud Infrastructure & DevOps (Azure / AWS)
Freelance Contract- one year plus extensions
Location Brussels-50% onsite
Role Summary
A principal-level cloud infrastructure and DevOps engineer who builds and operates the governed platform that AI runs on. You will design and run the AI landing zones on Azure and AWS, the model access and AI gateway layer, and the enterprise rollout of AI coding tools, so that every team gets fast, secure, cost-controlled and auditable access to models and agents. You use agentic tools in your own infrastructure work and know them well enough to administer them at enterprise scale.
Key Responsibilities
Governed AI landing zones
- Design, build and operate AI landing zones on Azure and AWS: management-group / OU hierarchy, subscription and account vending, Azure Policy and service control policies, hub-and-spoke networking, private endpoints, no public model endpoints.
- Enforce EU data residency, customer-managed encryption keys, central logging and guardrails as code, so that the compliant path is the default path.
- Ensure cross-cloud resilience for AI services, including multi-region and multi-provider failover for model endpoints.
Model access and AI gateway
- Provision and operate model platforms: Microsoft Foundry (including Azure OpenAI and Claude models), Amazon Bedrock (including Guardrails and AgentCore) and, where data sensitivity requires it, self-hosted open-weight models.
- Build and run a central AI gateway (e.g., Azure API Management AI gateway, LiteLLM, Kong AI Gateway, Claude apps gateway): authentication, per-team token quotas and rate limits, model routing and fallback, semantic caching, logging with PII controls, cost attribution.
- Manage capacity and quota across regions and providers (provisioned vs. pay-as-you-go throughput).
Enterprise AI coding tooling
- Roll out and administer AI coding tools at organisation scale: Claude Code routed through Microsoft Foundry or Amazon Bedrock via the gateway; managed settings and policies (permission modes, allowed tools, MCP server allow-lists); SSO and group-based model access; OpenTelemetry telemetry into the observability stack. Apply the same governance to other approved tools.
- Host, secure and govern MCP servers: registry and allow-list, OAuth-based authorisation, network egress control.
- Provide secure execution environments for autonomous agents: sandboxed dev containers, ephemeral CI runners, least-privilege credentials, egress restrictions.
Operations, security and cost
- Identity and secrets: Microsoft Entra ID, AWS IAM Identity Center, workload identity federation, managed identities, HashiCorp Vault, Keycloak.
- Observability: OpenTelemetry (including GenAI semantic conventions), Azure Monitor, CloudWatch, Elastic / ELK; SLOs for the AI platform.
- FinOps for AI: token and cost dashboards per team and use case, showback / chargeback, budget alerts, anomaly detection.
- AI security: threat modelling (OWASP Top 10 for LLM Applications, MITRE ATLAS), content safety and prompt-injection defences, DLP (e.g., Microsoft Purview), audit trails; evidence for EU AI Act and ISO/IEC 42001 compliance, in cooperation with security and compliance teams.
- Everything as code: Terraform (primary), Bicep / CloudFormation / CDK, policy as code, GitOps, drift detection, fully reproducible environments.
- Use AI agents in platform engineering itself: IaC authoring and review, runbook automation, AI-assisted incident triage integrated with ServiceNow.
Essential Requirements
- 10+ years in cloud infrastructure, platform engineering or DevOps, including 4+ years as principal or lead architect for enterprise environments.
- Expert-level Azure (Cloud Adoption Framework / Azure Landing Zones, networking, Entra ID, Azure Policy, AKS, API Management) and expert-level AWS (Organizations / Control Tower, IAM, VPC, EKS).
- 2+ years provisioning and operating generative-AI platform services in enterprise settings: Azure OpenAI / Microsoft Foundry and Amazon Bedrock.
- Hands-on implementation of an AI / LLM gateway in production (quotas, routing, logging, cost attribution).
- Expert Terraform; Kubernetes in production; CI/CD and GitOps (Argo CD or Flux).
- Hands-on use of agentic coding tools, including Claude Code, both as a daily user and as administrator of enterprise configuration.
- Strong security engineering: zero-trust networking, Private Link, key management, SIEM integration, least-privilege identity design.
- Observability and FinOps practice on both clouds.
- Excellent written and spoken English (C1+); able to produce architecture documentation and defend designs to security and enterprise-architecture boards.
Desirable
- Red Hat OpenShift (including OpenShift on Azure) and OpenShift AI.
- Self-hosted inference: GPU capacity planning, vLLM, KServe.
- HashiCorp Vault, Keycloak, Elastic APM.
- Certifications: Azure Solutions Architect Expert, AWS Certified Solutions Architect – Professional, CKA, HashiCorp Terraform.
- Regulated, aviation or intergovernmental-organisation experience; ISO 27001, EU AI Act, ISO/IEC 42001.
- ServiceNow ITSM integration for AIOps.
- French.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search