Dicetek LLC
Linkedin · Posted yesterday
AI Ops Engineer
Continue to application
Add your email once, then Caio opens the original posting.
Indexed description
Role PurposeWe are seeking an AI Ops Engineer to establish the operational backbone for enterprise AI platforms, enabling application teams to release AI products safely, repeatedly, and at scale. The role is accountable for production release discipline, LLMOps practices, deployment automation, operational governance, cost visibility, and self-service operating standards for AI-native delivery teams.
Key Responsibilities
- AI Release Engineering & CI/CD: build and design standard release pipelines and promotion controls for AI applications, agents, platform services, and configuration changes across environments, ensuring repeatable deployment, governance, and release evidence.
- LLMOps & AI Lifecycle Management: embed operating controls for models, AI assetsand release for auditable.
- Progressive Delivery: Implement deployment patterns to reduce production risk, including controlled rollout, canary release, and rollback readiness
- Evaluation, Observability & Production Readiness: Embed AI quality checks, operational telemetry, dashboards, runbooks, and readiness criteria into the delivery lifecycle so AI services are measurable and supportable.
- AI FinOps & Capacity Governance: Provide visibility and controls for AI workload consumption, including model usage, token spend, platform capacity, quota management, and optimization opportunities.
- Self-Service & Continuous Improvement: Convert proven operating patterns into reusable templates, release standards, onboarding guidance, operational playbooks, and paved-road workflows that allow teams to move quickly while maintaining enterprise control.
- Strong production engineering background operating cloud-native, AI, or high-scale API platforms in an enterprise environment.
- Hands-on experience with CI/CD, GitHub Actions or equivalent automation, and environment management.
- Working knowledge of LLMOps, telemetry, release governance, and production-readiness practices.
- Experience with observability, change control, service reliability, and continuous operational improvement.
- Ability to partner with platform engineering, QA/SRE, cybersecurity, architecture, product, and delivery teams to standardize safe and scalable AI operations.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search
Want help applying to roles like this?
Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search