Senior DevOps Engineer
Indexed description
Unstract is an open-source, no-code platform that helps enterprises automate complex, document-centric business processes using AI and Large Language Models. Users can design prompts, ingest documents from multiple sources, and build high-throughput extraction and ETL workflows without writing code.We offer the platform in three editions: a fully managed Unstract Cloud, self-hosted On-Prem, and Open-Source. While our cloud customers enjoy a turnkey SaaS experience, our enterprise customers deploy Unstract on their own Kubernetes infrastructure across AWS, Azure, GCP, or bare-metal environments. Unstract is independently audited for SOC 2 Type II, HIPAA, GDPR, and ISO 27001 compliance, making reliability, security, and operational excellence central to everything we build.
The RoleAs a DevOps / Site Reliability Engineer, you'll play a key role in delivering a seamless on-prem deployment experience for our enterprise customers while maintaining our cloud infrastructure. You'll own the deployment journey end to end, work directly with customers to troubleshoot complex deployments, improve deployment tooling and automation, and ensure our cloud and customer-managed environments are secure, reliable, and production ready.
Key Responsibilities- Own the end-to-end on-prem deployment experience for enterprise customers.
- Partner with customers over calls to guide deployments, troubleshoot issues, and unblock installations without direct access to their environments.
- Improve the on-prem deployment experience by building tooling, automation, runbooks, and documentation while incorporating customer feedback.
- Collaborate with Solutions Engineers on complex customer onboarding and enterprise deployments.
- Participate in the on-call rotation (approximately 2 weeks every 3 months), respond to production incidents, lead postmortems, and drive operational improvements.
- Build and maintain Infrastructure as Code using Terraform.
- Manage Kubernetes clusters, Helm charts, ArgoCD, and CI/CD pipelines across cloud and customer-managed environments.
- Enhance observability, security, compliance, and infrastructure reliability using tools such as Prometheus, Grafana, Loki, and Alertmanager.
- Optimize infrastructure performance, scalability, and cost across cloud and on-prem environments.
- 4–6 years of experience in DevOps, SRE, or Cloud Operations roles, with on-call rotation experience.
- Strong hands-on experience managing production Kubernetes clusters (EKS, GKE, AKS, or bare metal).
- Proficiency with Terraform and at least one major cloud provider (AWS preferred).
- Experience with CI/CD tools such as GitHub Actions, GitLab CI, or similar.
- Hands-on experience with Helm and ArgoCD.
- Comfortable working with Linux, container networking, Bash, Python or Go, and observability tools such as Prometheus, ELK, or Loki.
- Excellent troubleshooting, communication, and documentation skills.
- Experience supporting or deploying applications in customer-managed/on-prem Kubernetes environments is highly preferred.
- Contributions to open-source DevOps or CNCF projects.
- Impact & Ownership: Own the deployment experience for enterprise customers and help shape the reliability of a platform processing millions of pages every day.
- Modern Stack: Work with Kubernetes, Terraform, GitOps, AI workloads, and cloud-native technologies.
- Remote-first: Flexible work culture with generous PTO and a home-office stipend.
- Growth: Work alongside experienced founders and engineers at the intersection of DevOps, MLOps, and Document AI.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search