AI Architect
Indexed description
Our Values
We are a passionate and empathetic team that prioritizes human values. Our purpose is to elevate the quality of lives for our family, customers, partners and the community.
Equal Opportunity Statement
CLOUDSUFI is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified candidates receive consideration for employment without regard to race, colour, religion, gender, gender identity or expression, sexual orientation and national origin status. We provide equal opportunities in employment, advancement, and all other areas of our workplace. Please explore more at https://www.cloudsufi.com/
AI Platform Architect
- Role Overview
This is a builder-architect role. The successful candidate will define architecture, make technology decisions, develop reference implementations, review critical code and designs, and guide engineering teams from prototypes to secure, scalable production systems.
- Key Responsibilities
- Design platforms for single-agent and multi-agent systems supporting planning, reasoning, tool use, memory, delegation, validation, and human approval.
- Define orchestration patterns for deterministic, dynamic, event-driven, and long-running AI workflows.
- Establish clear boundaries between LLM reasoning, application logic, quantitative computation, rules, and human decision-making.
- Evaluate and adopt agent frameworks, model providers, tools, and orchestration technologies based on reliability, flexibility, performance, and cost.
- Architect RAG pipelines, document-processing systems, vector search, hybrid retrieval, knowledge graphs, and semantic data layers.
- Integrate structured and unstructured enterprise data from APIs, databases, files, streams, and external platforms.
- Design reusable workflows for research, data collection, transformation, analysis, modelling, validation, and reporting.
- Establish data lineage, provenance, metadata, access controls, freshness, and quality standards.
- Build evaluation frameworks for accuracy, relevance, groundedness, task completion, tool use, safety, latency, and cost.
- Enable systematic experimentation across models, prompts, agents, tools, retrieval strategies, and orchestration patterns.
- Implement versioning and lifecycle management for prompts, agents, workflows, datasets, knowledge bases, evaluations, and model configurations.
- Establish tracing, monitoring, auditability, guardrails, approval workflows, and production-quality diagnostics.
- Define cloud-native architectures using microservices, APIs, event-driven systems, queues, schedulers, and distributed processing.
- Lead Kubernetes-based deployment, containerisation, CI/CD, Infrastructure as Code, environment management, and release automation.
- Design for horizontal scalability, fault tolerance, resilience, security, data privacy, and high availability.
- Optimise model usage, infrastructure, storage, retrieval, and compute for performance, latency, and cost.
- Translate product and business requirements into clear technical designs and implementation plans.
- Build prototypes and reference implementations for high-risk or foundational platform capabilities.
- Review architecture, code, interfaces, data models, infrastructure, and operational readiness.
- Define engineering standards and reusable patterns across AI, backend, data, and platform teams.
- Mentor senior engineers and support teams in resolving complex technical and production issues.
- Required Skills and Experience
- 10+ years of experience in software architecture, platform engineering, distributed systems, data platforms, or AI systems.
- Strong hands-on experience designing and building production-grade AI or data-intensive platforms.
- Deep understanding of LLM applications, tool calling, structured outputs, RAG, embeddings, memory, and agent orchestration.
- Strong experience with cloud platforms, Kubernetes, containers, microservices, APIs, event-driven architecture, CI/CD, and Infrastructure as Code.
- Experience with relational, document, graph, vector, and distributed data systems.
- Practical experience implementing AI evaluation, experimentation, tracing, monitoring, guardrails, and lifecycle management.
- Strong understanding of security, identity, access control, secrets management, data protection, and production reliability.
- Ability to move effectively between architecture, code, infrastructure, debugging, and technical delivery.
- Good to Have
- Experience building enterprise AI copilots, autonomous workflows, research platforms, or analytical systems.
- Experience with knowledge graphs, hybrid search, model gateways, tool gateways, or agent marketplaces.
- Familiarity with LLMOps, MLOps, model serving, feature stores, model registries, and distributed compute.
- Experience supporting real-time and batch data processing at scale.
- Experience comparing and operating multiple commercial and open-source models.
- Prior experience in consulting, client-facing architecture, or complex enterprise platform delivery.
- What We Are Looking For
The right candidate can define platform direction, evaluate trade-offs, validate ideas through implementation, and guide systems into production. They should be equally comfortable discussing distributed architecture, reviewing code, diagnosing workflow failures, designing evaluation systems, and mentoring engineering teams.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search