Associate AI Quality Engineer
Indexed description
The Job in a Nutshell💡
You'll build the internal AI systems our engineers work inside every day: agents that generate and maintain tests, pipelines that triage failures before a human sees them, tooling that speeds up code review and debugging, and the evaluation infrastructure that makes our own AI features testable at all.
This is a builder's role. You'll write production code, own systems in CI, and be measured on whether engineers actually use what you ship.
What Will You Do❓
Test automation across the stack
- Backend: API and contract testing, service-level and integration coverage, data setup that doesn't rot, and test design that survives a schema change
- Frontend: web E2E and component-level coverage with Playwright, visual and RTL regression, and suites fast enough to gate a merge rather than a nightly
- Mobile: native and cross-platform coverage with Appium or Maestro, device-farm strategy, offline and sync behaviour, and the payment-peripheral paths that only break on real hardware
- The connective tissue: shared fixtures, environment and test-data management, parallelisation, and CI pipelines where a red build means something
- performance and load testing Experience
- Test generation from specs, code, and production traffic — with the maintenance story solved, not just the first draft
- Failure triage that classifies a red build before a human opens it: real bug, flake, environment, or test rot
- Self-healing locators and suite health tooling — flake detection, quarantine, coverage-gap analysis
- Evaluation infrastructure for AI features across our products: datasets, scoring, and regression detection when a prompt or model changes
- Evaluation for our market specifically — Arabic and English behaviour, RTL interfaces, and region-specific POS, tax, and payment rules. Correctness here is rarely a string match
- Agentic AI that does real work in our pipelines: reads a diff, runs the relevant suite, reproduces a failure, proposes a fix, opens the PR
- Agents that own a quality workflow end to end — exploratory testing against a running build, coverage-gap hunting, release-risk assessment — and know when to escalate to a human
- Orchestration that holds up under load — multi-step planning, tool use, retries, state and memory across steps, sandboxed execution, multi-agent handoffs, and clean boundaries between agentic and deterministic steps
- Integration with the stack we already have (CI, Jira, observability, MCP-style tool interfaces) rather than a parallel system beside it
- The judgement to know when a plain pipeline beats an agent, and to say so
- Test automation: framework design and layering, the test pyramid and where it stops being useful, flake economics, parallel execution, mobile and cross-browser realities, CI/CD gating, Framework: Playwright, Appium, Maestro
- Agentic AI: orchestration and tool use, multi-step planning, memory and state, sandboxed execution, multi-agent patterns, MCP and similar tool-integration standards, and the cost of each
- Context engineering: retrieval strategy, chunking, reranking, caching, and managing long-context behaviour — including where it degrades
- Evaluation: offline and online evals, LLM-as-judge and its failure modes, human-in-the-loop review, statistical significance on small samples, regression gates in CI
- Reliability: structured output, guardrails, fallback and retry design, and handling non-determinism in systems that must not flap
- Operations: tracing and observability for LLM systems, prompt and version management, latency and cost budgeting, model routing, and when fine-tuning or distillation beats a better prompt
- An engineer who ships production software, with recent hands-on work on LLM-backed systems that real users depend on.
- Strong Python; comfortable in at least one of .NET, Java, or TypeScript. Tested, maintained code — not notebooks
- Real automation depth across more than one surface. You've owned a suite that gates releases on backend and on a UI — web or mobile — and you can explain how you kept it green without deleting the hard tests
- Real experience building evaluation systems. You can explain how you knew your system was getting better, with numbers
- Practical depth with the modern LLM toolkit — prompting, structured output, tool use, retrieval, agentic AI orchestration — and a clear sense of the trade-offs
- Credible testing fundamentals. You don't need a QA title, but test design, automation frameworks, and CI/CD shouldn't be new to you
- A bias toward adoption. You measure your work by what other engineers use, not by what you demoed
- We have an inclusive and diverse culture that encourages innovation and flexibility in-office, and hybrid work setups
- We offer highly competitive compensation packages, including bonuses and the potential for shares
- We prioritize personal development and offer regular training and an annual learning stipend to tackle new challenges and grow your career in a hyper-growth environment
- Join a talented team of over 30 nationalities working in 14 countries, and gain valuable experience in an exciting industry
- We offer autonomy, mentoring, and challenging goals that create incredible opportunities for both you and the company
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search