Senior AI Engineer
Indexed description
We’re a team of creators. We write code, shape delivery, build go-to-market strategies, develop AI solutions and create the practices that support our people. We work side by side with our clients, challenging what’s not working and helping them to build the future. Our commitment to craft, quality, and culture has helped us scale to over 600 people in just a few years.
Our UK Benefits
- 35 days leave (including bank holidays).
- Private medical insurance.
- Enhanced parental and adoption leave.
- Financial coaching + 5% pension match.
- 40 hours of paid learning and development.
Join us on our journey. Let’s create tomorrow, together, today.
About The Role And Team
You'll join our AI Engineering team with a focus on evaluation, experimentation, and observability — building the systems that tell us whether our AI features actually work, and keep working, in production. This is a hands-on, technical role representing CreateFuture day to day, working across eval platforms, golden datasets, and production monitoring for AI-native systems.
What You'll Be Doing
Technical Delivery & Implementation
- Eval platform: Design and build the evaluation pipelines on a platform such as Braintrust, LangSmith, Arize or Weights & Biases, and own the decision about which one fits the problem.
- Datasets and scoring: Curate golden datasets that represent real user behaviour. Design LLM-as-judge and human-in-the-loop scoring pipelines with a clear view of where each one
- Production observability: Build token-level tracing, quality signal monitoring, hallucination and refusal detection, and per-use-case cost attribution, with alerting that catches degradation before users report it.
- Quality gates: Wire eval gates into CI/CD promotion paths so a regression blocks a release rather than being discovered in it.
- Regression and drift: Establish the baselines, regression suites and drift detection that let the client's teams change prompts, models and retrieval without holding their breath.
- Statistical honesty: Know the difference between a signal and noise, size your eval sets accordingly, and be straight with people when the data can't answer the question they're asking.
- Turning data into decisions: Present eval and observability findings to engineering and product audiences in a way that leads to a decision, and follow through until something changes.
- Setting the standard: Help the client's teams adopt eval-first habits, and challenge the "ship it and see" instinct where you find it, with evidence rather than dogma.
- Delivery ownership: Plan and prioritise your own stream, estimate accurately, manage shifting requirements without dropping quality, and flag risk to timelines early.
- Cost awareness: Understand the cost of what you're measuring. Eval runs and judge models aren't free, so design for a sensible ratio of cost to confidence and be able to justify
- Fitting in fast: Work within the client's existing tooling and release process where it's sound, and make the case for change where it isn't.
- Documentation as you go: Leave the dataset provenance, scoring rationale and runbooks that let the client's engineers maintain and extend the eval suite without you.
- Knowledge transfer: Bring the client's engineers along in evaluation practice, prompt and context engineering, and critical review of model output, so the discipline outlasts the
- Exit readiness: Treat a clean handover as part of the definition of done from the first sprint, not something arranged in the final fortnight.
We're looking for a mix of AI engineering (50%), software engineering (30%), data engineering (20%).
- Python: Strong and production-grade. You write code others maintain.
- Eval and observability platforms: Hands-on production experience with Braintrust, LangSmith, Arize, Weights & Biases or an equivalent, including the parts that didn't work well.
- Scoring design: Demonstrable experience designing LLM-as-judge and human-in-the-loop pipelines, and calibrating them against human judgement.
- Production AI systems: Experience operating LLM-backed features in production, covering tracing, latency, token cost, failure modes and retrieval quality.
- CI/CD: Enough fluency with GitHub Actions, GitLab CI or similar to own an eval gate in a promotion pipeline yourself.
- Data handling: Comfortable building the data pipelines and stores behind datasets, traces and results at volume.
- Communication: The ability to make a technical result land with a non-technical audience, tailoring the framing to the room.
- Regulated industries: Experience where model behaviour has compliance or consumer- protection consequences, such as iGaming, financial services or health, is a strong advantage. Safer gambling and responsible-messaging contexts are directly relevant here.
- Contract and consulting delivery: A track record of arriving on an unfamiliar programme and being productive quickly, including building credibility with engineers who didn't ask for your help.
Evidence
- AWS Certified Machine Learning Engineer (Associate) or AWS Certified AI Practitioner.
- An associate-level AWS or GCP certification for the platform fundamentals around your work.
We offer flexible working, including hybrid and remote options. Our office hubs are located in Edinburgh, Leeds, Manchester, London and Bulgaria, with occasional travel to client sites or CreateFuture offices when needed.
We trust you to manage your time balancing collaboration with client time and focused work. What matters is the impact you have, not how busy you look.
Our hiring process
We try to keep our hiring process clear, fair and respectful of your time. We aim to get back to everyone who applies and we will be upfront about where you are in the process.
It Usually Looks Like This
- Call with our Talent Acquisition Team
- Role specific capability interview
Inclusion at CreateFuture
We believe diverse teams build better workplaces and better products. We want CreateFuture to be a place where people feel able to be themselves and do their best work.
If you need any adjustments or support during the application process, just. We will do what we can to help.
We look forward to your application!
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search