Staff Platform Engineer
Indexed description
Staff Platform Engineer (GCP)
$300k Base + 20% Bonus + Stock Options
Staff Platform & Security Engineer, AI Agent Infrastructure
High-growth, venture-backed company · US Remote (hybrid optional)
About the Role
Our client is hiring a Staff-level engineer to own the infrastructure and security foundations of its AI agent platform. As agents move from answering questions to taking real actions across production systems, the hard problems shift to identity, isolation, credentials, and auditability. This role owns them.
This is a platform and security engineering role first. AI is the domain; deep infrastructure and security craft is what they're hiring for.
What You'll Own
Platform & Infrastructure
- Architect and operate a multi-tenant, Kubernetes-based runtime for agent workloads, including custom controllers and operators, CRDs, Helm, and per-tenant isolation.
- Build automation that provisions, scales, and tears down isolated environments on demand.
- Own infrastructure-as-code (Terraform) across cloud environments, along with CI/CD and progressive delivery.
Security & Identity
- Design a security model that treats agent workloads as untrusted.
- Build short-lived, least-privilege credentials and delegated authorization in place of standing secrets.
- Own secrets management and key-managed encryption for sensitive tokens.
- Enforce zero-trust service-to-service communication and deny-by-default networking.
- Threat-model and mitigate confused-deputy attacks, privilege escalation, prompt injection, and data exfiltration.
Reliability & Observability
- Run the platform as a fleet: health monitoring, automated recovery, efficient scaling, and SLOs.
- Build observability and audit pipelines (OpenTelemetry) that give a complete, investigable record of agent activity.
- Lead incident response and balance reliability, performance, and cost.
Agent Safety Controls
- Build human-approval workflows for sensitive or irreversible actions.
- Build protected configuration that agents cannot modify.
- Partner with the AI team so that every agent capability is exposed through governed interfaces.
Technical Leadership
- Set platform and security standards.
- Mentor senior engineers through architecture, not just code review.
- Shape the infrastructure roadmap with Product, Security, and Leadership.
What You Bring
- 8+ years in software, infrastructure, or distributed systems, with end-to-end ownership of security-critical production systems.
- Deep production Kubernetes experience: operators, controllers, multi-tenancy, workload isolation, fleet operations.
- Cloud-native architecture on a major cloud provider, plus strong Terraform.
- Applied security depth: workload identity, service mesh and mTLS, OAuth 2.0 / OIDC, token exchange, secrets management and KMS, zero-trust design.
- Excellent Go and/or Python.
- SRE fundamentals: SLOs, incident management, observability, capacity and cost engineering.
- High autonomy, and comfort taking ambiguous problems from architecture to production.
Nice to Have
- Experience with LLM infrastructure, agent platforms, or AI security.
- Container sandboxing or runtime isolation technologies.
Package: competitive base, equity, bonus, full benefits.
To apply please contact Sam Shinner at Discover International
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search