DevOps Engineer
Indexed description
DevOps Engineer – Hedge Fund / AI Platform
New York City
We are looking for a hands-on DevOps Engineer to own the production foundation supporting a growing suite of internal AI applications, data pipelines, and engineering tools. This person will work closely with software, AI, data, and investment-facing teams to take applications from prototype through production and ensure they remain reliable, observable, secure, and scalable.
This is not a traditional infrastructure support position. The ideal candidate is an engineer who can automate infrastructure, improve deployment patterns, troubleshoot complex production issues, and establish the operational standards needed to move quickly without sacrificing reliability.
Responsibilities
- Design, build, and operate highly available cloud infrastructure supporting internal AI applications, APIs, data pipelines, and engineering platforms.
- Build and maintain infrastructure through Terraform or comparable Infrastructure-as-Code tooling, minimizing manual configuration and improving consistency across environments.
- Own and improve CI/CD pipelines, automated deployments, release processes, environment management, and rollback strategies.
- Build and operate containerized platforms using Docker and Kubernetes, including workload deployment, scaling, networking, configuration, and production troubleshooting.
- Establish comprehensive observability across applications and infrastructure, including metrics, logs, traces, dashboards, alerting, and service-health monitoring.
- Monitor the health and performance of AI applications and associated data ingestion, retrieval, API, and processing pipelines.
- Define and track production SLIs/SLOs, availability targets, latency, throughput, error rates, and other operational metrics.
- Lead or participate in incident response, root-cause analysis, postmortems, and remediation of recurring production issues.
- Design for high availability, disaster recovery, backup/recovery, fault tolerance, and business continuity.
- Own capacity planning and cloud-cost optimization, identifying opportunities to improve infrastructure utilization without compromising performance or resiliency.
- Implement secure infrastructure patterns including IAM, least-privilege access, secrets management, network controls, and environment isolation.
- Develop automation and operational tooling using Python, Go, Bash, or similar languages.
- Create production standards, operational documentation, runbooks, escalation procedures, and deployment guidelines.
- Partner closely with AI/ML, software, data, and security engineers to ensure new applications are production-ready from the outset.
- Help engineering teams move rapidly from experimentation to production while introducing the appropriate level of operational rigor, automation, and control.
Ideal Background
- Strong experience in DevOps, SRE, platform engineering, cloud infrastructure, or production engineering.
- Proven experience operating high-availability, business-critical production systems.
- Strong cloud experience with AWS, GCP, or Azure.
- Deep experience with Terraform/IaC, Kubernetes, Docker, CI/CD, observability, monitoring, and incident management.
- Strong scripting/software engineering ability with Python, Go, Bash, or similar.
- Experience supporting data-intensive applications, distributed systems, APIs, or data pipelines.
- Strong understanding of networking, Linux, IAM, security controls, and production architecture.
- Experience within a hedge fund, trading firm, financial institution, technology company, or other high-performance engineering environment is beneficial.
- Exposure to production AI/ML or LLM applications is helpful but not required.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search