Senior Software Engineer (Agentic Platform)
Indexed description
About the Role & Product
We are building an Agentic Platform — a platform that helps customers deploy and operate AI agents at production scale quickly and securely. The platform provides ready-to-use services/components that agent developers can quickly integrate instead of building from scratch — significantly shortening the time from idea to a complete agent. The platform also follows open protocols to easily connect with tools, data sources, and other agents.
We believe that a solid agentic platform is built on a software engineering foundation (backend & distributed systems), with AI/agent capabilities layered on top. Therefore, this role is designed with a ratio of 70% traditional software engineering/platform engineering and 30% AI, AI agents, and related protocols.
You will work directly with the engineering team to take platform services from architecture and implementation to operation — ensuring high availability, low latency, security, and scalability.
Key Responsibilities:
1. Platform & Backend Engineering (~70%)
- Design, implement, and maintain backend services/APIs that meet high standards for availability, latency, and security.
- Build platform infrastructure components to support the agent lifecycle — from where agents run, communication between agents/tools/services, state management, to security and authorization.
- Build a secure sandbox runtime where AI agents can execute generated code and tools in isolated environments — using containerization/microVMs (Docker, gVisor, Firecracker), isolation through namespaces & seccomp, resource limits (CPU/memory/network), and strict egress controls to ensure untrusted agent actions never leak to the host or other tenants.
- Design data models and data flows across SQL (MySQL), NoSQL (MongoDB), and vector databases; ensure consistency, throughput, and resilience.
- Integrate message brokers/event-driven backbones (Kafka, RabbitMQ, AWS SQS/SNS) for asynchronous communication between internal platform services.
- Identify and resolve system issues: performance bottlenecks, memory leaks, race conditions, security vulnerabilities, and resource leaks.
- Design and implement observability systems (metrics, logs, distributed tracing) to monitor and debug the platform in production environments.
- Research, experiment with, and evaluate new technology solutions to address technical system challenges.
2. AI & Agentic Capabilities (~30%)
- Design platform services for AI agents; proactively monitor trending AI agent capabilities in the market to ensure customers always have access to the latest capabilities.
- Build and integrate MCP servers as well as adapters for the A2A protocol to connect agents with tools, data sources, and other agents.
- Integrate LLM/agentic frameworks (LangChain, LangGraph, CrewAI, Strands…) into the platform in a framework-agnostic manner; design abstractions so that the platform does not depend on any specific model or framework.
- Design and build agent sandbox capabilities — code/tool execution environments (code interpreter, REPL, tool-call runtime) that give AI agents the ability to run code, evaluate output, and iterate safely; support multi-language execution, persistent session state, file I/O, and deterministic reproducibility so that agents can reliably “think → execute → observe → refine.”
- Define and enforce sandbox security boundaries — permissions per agent/per task (read-only fs, network allow-list, approval gates for sensitive actions), and deeply integrate with the platform’s authorization layer so that actions within the sandbox can be audited, traced, and revoked.
- Write clients (SDK, CLI, AI agent skills) to help customers integrate with and use the platform.
- Apply AI tools (coding assistants, AI code review, AI-assisted testing & debugging) to daily development and operations workflows to improve team productivity.
3. Engineering Practices & Ownership (applies to both areas)
- Write clean code with unit tests and integration tests; maintain quality through code reviews.
- Own features/services from design to production; proactively propose architectural improvements.
- Guide and review junior/mid-level engineers; contribute to building the engineering culture.
- Work closely with product, infra, and stakeholders to translate requirements into clear technical solutions.
Job Requirements
I/ Must Have:
1. Background and Experience
- Bachelor’s degree or higher in Computer Science, Engineering, or a related field.
- Experience building backend services/distributed systems running in production environments; for us, years of experience are only a reference — actual capability is the deciding factor.
- Strong foundation in software architecture, design patterns, distributed systems, and best practices.
- Strong problem-solving, logical thinking, and analytical skills; effective communication and teamwork skills.
2. Backend & Data
- Proficient in at least one of: Java (Spring Boot), Go, Python.
- Proficient in SQL (MySQL) and NoSQL (MongoDB); understanding of indexing, query optimization, transactions, and data modeling.
- Hands-on experience with message brokers: Kafka, RabbitMQ, ActiveMQ, or AWS SQS/SNS.
- Strong understanding of REST API design, authentication/authorization (OAuth, API key, service-to-service auth), and security fundamentals.
3. Infra & DevOps
- Familiar with Git, Docker, CI/CD; understanding of microservices deployment.
- Understanding of observability (metrics, logs, tracing) and experience working with related tools (Prometheus/Grafana, ELK/OpenSearch, OpenTelemetry, Jaeger…).
4. AI & Agentic (30%)
- Understanding of core agent architecture concepts: tool use, memory/context, orchestration loop, guardrails, multi-agent coordination.
- Understanding of and experience working with MCP (Model Context Protocol) — knowing how to write an MCP server is a major advantage.
- Understanding of A2A (Agent-to-Agent) and other agent protocols (ADK, ReAct).
- Experience using agentic frameworks (LangChain, LangGraph, CrewAI…) and/or RAG/vector databases.
- Familiar with using AI tools in daily work.
II/ Nice-to-have
- Experience building or contributing to a real-world agentic platform/AI platform.
- Experience writing MCP servers and integrating them into production systems.
- Experience implementing end-to-end monitoring/observability systems.
- Experience with Kubernetes and operating container workloads at scale.
- Experience with code interpreter sandboxes, identity/authorization for agents, or semantic caching.
- Experience building a secure sandbox/code execution runtime for AI agents or multi-tenant workloads — hands-on experience with container isolation (Docker, gVisor, Firecracker/microVM), namespaces & seccomp, network policies, or serverless code execution platforms.
- Good English listening and speaking skills — an advantage when working with international teams, partners, and technical documentation
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search