Member of Technical Staff, Infrastructure
Indexed description
We’re looking for folks with experience managing cloud infrastructure, working through various stages of scale, and helping the broader Engineering team be more effective and productive. Some traits that are important to our company culture: customer-obsessed, collaborative, hard-working, and optimistic. And we’re looking for owners, so we hope you’ll help us expand this list.
Responsibilities
- Collaborate with other engineering teams to build and maintain foundational systems that empower developers and support the company's rapid growth.
- Design and implement scalable infrastructure solutions for various deployment models, including SaaS, single-tenant, and private deployments.
- Manage and optimize cloud resources and Kubernetes clusters for cost-effectiveness and performance.
- Enable external customer deployment success through maintaining clear infrastructure boundaries and principles.
- Optimize and improve the release and deployment processes to enhance efficiency and reliability.
- Ensure compliance with relevant regulations and implement robust security measures across different deployment environments.
- Build and operate infrastructure for production LLM applications, including model integrations, inference workloads, evaluation pipelines, observability, and the reliable execution of agentic workflows.
- 8+ years of engineering experience.
- Worked on Platform or Infrastructure teams on significant projects involving infrastructure components (Terraform/CDKTF, Kubernetes, Helm, test infrastructure, release management, observability, etc.)
- Experience in optimizing cloud resource utilization.
- Proficient in tuning Kubernetes clusters and cloud resources for cost and performance efficiency.
- Willing to build LlamaIndex’s engineering culture as we grow.
- You can balance speed and pragmatism and build the appropriate solutions for each stage of the company’s growth.
- Hands-on proficiency with modern LLM tooling and production AI systems, including experience with model APIs, agent or RAG frameworks, evaluation and tracing tools, and the operational characteristics of LLM workloads.
- Experience building out infrastructure from the ground up at a fast-growing startup.
- Experience with observability tools like Prometheus, Grafana, and New Relic.
- Experience with GitOps tools like ArgoCD and Flux for continuous deployment.
- Experience with security compliance and audits in cloud environments such as SOC2.
- Familiar with Python, Postgres, multi-cloud deployments
Why Join Us?
- Impactful Mission: Work on innovative AI products that redefine how knowledge is accessed and utilized.
- Collaborative Team: Join a team of passionate individuals committed to pushing the boundaries of technology.
- Growth Opportunities: Be at the forefront of the AI revolution, with ample opportunities to grow alongside our scaling organization.
- Competitive base salary and equity compensation
- Comprehensive medical/dental/vision coverage for you and your family
- Unlimited paid time off policy
- Daily catered lunch and snacks in the San Francisco office
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search