Senior DevOps Engineer
Indexed description
Our team includes talent from Google, Meta, Amazon, Microsoft, and Stripe, with advisors from OpenAI, Google, HubSpot, and Stripe.
The Role
We’re looking for a Senior DevOps Engineer to own and evolve the infrastructure powering Maven AGI’s AI platform. You’ll design, build, and operate production systems across cloud providers, Kubernetes clusters, on-premises environments, and CI/CD pipelines, ensuring our platform scales reliably as we onboard enterprise customers with complex deployment and security requirements. This is a high-leverage role where your work directly impacts platform availability, developer velocity, and customer trust.
What You'll Do
- Design, implement, and maintain both cloud and on-premise infrastructure (Azure, AWS, datacenter) using infrastructure-as-code (Pulumi, Bicep, Terraform)
- Own Kubernetes cluster operations: deployments, scaling, monitoring, and incident response
- Build and optimize CI/CD pipelines for a large-scale monorepo
- Implement observability across services (metrics, logging, tracing, alerting)
- Drive reliability practices: SLOs, capacity planning, disaster recovery, and runbook development
- Operationalize and scale enterprise AI deployments on-premise, including GPU resource orchestration, model inference performance tuning, and high-concurrency platform management.
- Collaborate with engineering teams to improve developer experience and deployment velocity
- Manage secrets, access controls, and infrastructure security posture
- Evaluate and adopt new tooling to reduce operational toil
- 3-7 years of professional DevOps/SRE/Infrastructure experience
- Deep expertise with Kubernetes in production (AKS, EKS, or GKE)
- Strong infrastructure-as-code skills (Pulumi, Terraform, or Bicep)
- Experience operating CI/CD systems (GitHub Actions, ArgoCD, or Jenkins)
- Proficiency in at least one scripting/programming language (Python, Go, TypeScript, or Bash)
- Solid understanding of IaaS providers, networking, DNS, load balancing, and TLS
- Experience with monitoring and observability stacks (Datadog, Prometheus, Grafana, or similar)
- Experience with multi-cloud or hybrid (cloud + on-prem) deployments
- Strong communication and cross-team collaboration skills
- Organized, great attention to detail, comfortable operating in a ticketing environment
- Thrives in fast-paced startup environments
- Experience with GPU infrastructure and ML/LLM serving workloads (vLLM, TEI)
- Familiarity with Temporal or other workflow orchestration systems
- Security and compliance background (SOC 2, HIPAA, GDPR)
- Cost optimization experience at scale
- We are customer champions. You put users at the center of your thinking, advocate for their needs, and design solutions that make their lives measurably better.
- We are bold in action. You move with urgency and courage. You’re not afraid to challenge convention, take smart risks, and push boundaries in pursuit of meaningful outcomes.
- We are data-driven and insight guided. You make thoughtful decisions grounded in evidence. You’re curious, analytical, and combine data with intuition to guide strategy and execution.
- We are stronger together. You bring others along, value diverse perspectives, and contribute to a culture of trust and shared ownership. You believe the best ideas emerge through open dialogue and collective effort.
- High Impact in cutting-edge field. Be at the vanguard of AI innovation.
- Competitive salary, comprehensive benefits, and meaningful equity stakes.
- A diverse and welcoming work environment where everyone’s voice is heard.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search