Back to search
Sadler Recruitment Linkedin · Posted yesterday

Site Reliability Engineer

United Kingdom

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Role: Site Reliability Engineer

Location: Central Cardiff, South Wales

Working pattern: Hybrid (1 day a week in Cardiff office)

Salary: £45,000 - £50,000 per annum


About the Role


This company provides managed AI operations for technology businesses. The company operates, secures and governs the cloud, observability and AI runtime layer behind mission-critical software.


This is a chance to take real ownership in a Cardiff business on the front line of AI-era cloud operations, with bleeding-edge tech, AI used responsibly, a team worth learning from, and the flexibility to do your best work.


Clients range from start-ups and scale-ups building fast without formal engineering controls, through to regulated, mission-critical platforms in fintech, healthtech and insurance, where the operating model and evidence trail matter as much as uptime. You'll work across traditional cloud platforms like Azure and AWS, as well as newer, AI-native application stacks such as Lovable, and the modern cloud platforms that often sit behind them, like Vercel and Supabase.


The company is a Datadog Advanced Partner (UK) and holds the accolade of being the world's first accredited MSP powered by Datadog. Datadog is a Nasdaq-listed observability platform. The company's technical focus is on observability and LLM observability specifically.


Key Responsibilities


Managed Service Delivery - Support customers across a mix of traditional and modern cloud platforms, Azure and AWS, alongside newer AI-native application stacks and the cloud platforms behind them, such as Lovable, Vercel and Supabase. Work through an ongoing ticket backlog per customer, balancing reactive support with proactive improvement work, and take the lead on triaging and resolving more complex or ambiguous incidents.


Observability & Datadog - Get up to speed on client environments and the Datadog platform quickly, then work independently: identify improvements and either raise them or implement them directly without needing to be told. Tickets may also be generated directly by Datadog alerts, covering both backlog items and live incidents, including acting as an escalation point for trickier ones.


Build & Improvement Projects - Deliver end-to-end build-and-run projects for managed service customers, from scoping through to live operation, with minimal oversight. Design and implement cloud cost optimisation strategies, working through the technical dependencies that can sometimes block them. Manage monitoring and observability infrastructure as estates grow, keeping platforms performant and well-maintained. Build and deploy new infrastructure in Azure and AWS, including infrastructure-as-code, Datadog implementation, and APM setup, covering the full lifecycle from onboarding through go-live stabilisation.


AWS Platform Work - Work hands-on and confidently with containers (ECS/EKS), EC2 and RDS instances across a range of customer environments, from day-to-day configuration and troubleshooting through to cost and performance optimisation and architectural recommendations. Also extends into AWS Bedrock for customers building AI-native products.


Internal Tooling & Process -Contribute to a growing suite of internal AI tools built to improve engineering efficiency, reduce context-switching across platforms, and modernise how we run cloud operations day to day. Help shape best practice and runbooks as the team scales.


Day in the Life

  • Review ticket queues across the managed service accounts, taking ownership of the more complex or customer-sensitive items.
  • Project time is spent on a mix of Azure, AWS and Datadog implementation work, alongside newer platforms like Lovable, Vercel and Supabase. Work includes cost optimisation, infrastructure builds, and Datadog implementation and APM setup on live applications.
  • On AWS, this includes container work, EC2 and RDS configuration, and Bedrock. Python is used throughout. Depending on the customer, this work may also involve GitHub Actions, Azure DevOps pipelines, or PowerShell scripting.
  • Time is also allocated to internal work covering internal AI tools and process improvements.
  • The role involves working across managed service delivery, project builds, observability, and internal tooling within the same week, rather than working on a single area on an ongoing basis.


Experience Required

  • Multiple years of hands-on, production AWS experience, this is not a foundational-familiarity role: you should be comfortable owning EC2, RDS, containerised workloads (ECS/EKS), IAM and networking, and cost/performance optimisation without close supervision, including comfortable Linux troubleshooting (CLI, logs, process/resource issues) on EC2-hosted workloads
  • Demonstrable experience having worked on AWS builds (standing up new infrastructure, IaC, migrations) as well as ongoing operational/support work (troubleshooting, incident response, day-to-day maintenance)
  • Solid working knowledge of observability: metrics, logs, traces, and how to design or tune alerting and dashboards
  • Comfortable scripting and automating: Python, Bash, or similar
  • Experience troubleshooting live production incidents, ideally including on-call
  • Clear written and verbal communication: you'll work directly with customers, including on more technical or sensitive conversations


Nice to have

  • Hands-on Datadog experience
  • Terraform or other IaC tooling
  • Kubernetes or containerised workload experience beyond the basics
  • Working knowledge of Azure alongside AWS
  • Experience in a managed service provider or multi-customer environment
  • Exposure to AI-native platforms or tooling (Bedrock, Vercel, Supabase, Lovable, or similar)
  • Familiarity with ISO 27001 or similar compliance frameworks
  • AWS certification (Solutions Architect, SysOps, or similar)


What's on Offer

Join a growing Cardiff business entering its scaling phase, with real ownership, visibility, and a chance to shape how a fast-moving modern MSP operates AI-era cloud infrastructure.

On-call is documented, structured and remunerated in addition to salary. Training and tooling costs are paid for by the company. Attendance at industry events and courses is encouraged and funded as part of ongoing development. Employees are expected to identify and pursue improvements independently, without pre-defined limits on scope.

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search