AI Platform Operations Engineer (Teradyne, Costa Rica)
Indexed description
We attract, develop, and retain a high-performance workforce, comprised of people with diverse backgrounds and a shared drive for excellence. We strive to foster a positive and inclusive work environment that helps employees, and communities, thrive.
Our Purpose
TERADYNE, where experience meets innovation and driving excellence in every connection. We are fueled by creativity and diversity of thought and in our workforce. Our employees are supported to innovate and learn something new every day.
We cultivate a culture of inclusion for all employees that respects their individual strengths, views, and experiences. We believe that our differences enable us to be a better team – one that makes better decisions, drives innovation and delivers better business results.
Opportunity Overview
Teradyne is seeking a Senior AI Platform Operations Engineer to operate, secure, automate, and continuously improve our enterprise AI platforms and services. Working closely with AI Platform Architects, AI Developers, Security, Infrastructure, and Collaboration teams, this role is accountable for the operational success of Teradyne's AI ecosystem. The engineer will translate architectural designs into reliable production services, ensuring AI platforms are secure, available, observable, scalable, cost-effective, and supportable.
This role is intentionally distinct from the AI Platform Architect. The Architect defines enterprise AI strategy, architecture, standards, reusable patterns, and governance frameworks. The Operations Engineer implements, runs, monitors, automates, and continuously improves the environments and services that bring those designs into production.
AI Platform Operations & Service Reliability
- Operate enterprise AI platforms and supporting Azure cloud infrastructure across development, test, and production environments.
- Administer platform configurations, service quotas, deployment resources, environment controls, and lifecycle activities.
- Own and author run-books, on-call rotations, and the support model for production AI services, maintaining service health procedures, knowledge articles, and support handoff materials as the single accountable owner.
- Participate in incident response, root cause analysis, post-incident reviews, and continuous service improvement activities.
Azure AI Foundry & Cloud Platform Engineering
- Support Azure AI Foundry capabilities including hubs, projects, model catalogs, agents, deployments, and capacity / quota management.
- Support Azure AI Services, Azure OpenAI, Azure AI Search, vector and RAG-oriented platform patterns, and related cloud services.
- Operate core Azure dependencies including VNets, Private Endpoints, VPN connectivity, Private DNS, Key Vault, storage, data services, and monitoring integrations.
- Partner with cloud, security, and architecture teams to ensure production AI workloads are secure, performant, resilient, and supportable.
Production AI Service Enablement
- Operationalize AI capabilities designed by the AI Platform Architect and transition them into supported production services.
- Implement operational controls, monitoring coverage, capacity planning practices, scaling processes, and release readiness checkpoints.
- Accept production handoff of internally developed ML/AI solutions from the Senior ML/AI Engineer once they meet release-readiness criteria, and support AI applications, agents, service integrations, and model deployment environments going forward.
- Coordinate support across I&O, Security, Collaboration, Data, and application teams for AI platform issues and escalations.
Observability, Automation & Operational Excellence
- Implement monitoring, alerting, logging, operational dashboards, and service health reporting using Azure Monitor, Log Analytics, Application Insights, and related tooling.
- Automate recurring operational activities using PowerShell, Python, Terraform, Bicep, Azure DevOps, GitHub Actions, and GitOps practices.
- Identify reliability risks, recurring incidents, capacity constraints, cost inefficiencies, and opportunities for automation.
- Create operational KPIs for availability, incident response, deployment quality, platform adoption, and service performance.
MLOps, LLMOps & Agentic AI Operations
- Execute model and pipeline lifecycle operations—registry promotion, retraining execution, and drift remediation—against criteria set by the Senior ML/AI Engineer, and operationalize MLOps / LLMOps patterns including environment management, deployment consistency, and monitoring.
- Own production operations for agentic AI services, governed platform integrations, API-based connectivity, and operational patterns for AI agents and workflows, consulting the Senior ML/AI Engineer and AI Tools Ops Lead as needed.
- Assist with production readiness for AI builders using Azure AI Foundry, Copilot Studio, and approved enterprise AI platforms.
- Ensure AI workloads can be monitored, supported, secured, and improved throughout their lifecycle.
Security, Governance & Compliance Operations
- Implement security controls defined by Architecture, Security, and Governance teams, including identity, access provisioning, RBAC, conditional access alignment, and audit support.
- Support platform compliance reviews, access reviews, operational audits, and governance evidence collection.
- Monitor AI platform security signals and support response activities for platform-related incidents or control gaps.
- Ensure production AI services comply with enterprise operational, security, privacy, and responsible AI standards.
AI Developer & Builder Enablement
- Provide advanced operational support to AI developers, solution builders, and platform consumers.
- Assist teams with onboarding, troubleshooting, environment setup, deployment readiness, operational handoff, and support escalation.
- Partner with architecture teams to improve usability, supportability, documentation, and self-service enablement.
- Help mature enterprise AI operating practices through knowledge sharing, documentation, and practical engineering guidance.
All About You
We seek individuals who share our passion and determination. Our commitment to customer success drives us to go the extra mile. If you’re ready to join us in this mission, take a closer look at the minimum criteria for the position.
- 5-8 years of experience in Cloud Infrastructure, Platform Engineering, DevOps, Site Reliability Engineering, Infrastructure Operations, or related disciplines.
- Hands-on experience operating Microsoft Azure enterprise environments and supporting production cloud applications or platforms.
- Experience with monitoring, incident response, automation, infrastructure lifecycle management, and operational support processes.
- Working knowledge of Azure networking, identity, security, storage, monitoring, and data platform services.
- Ability to partner effectively with architects, developers, security teams, infrastructure teams, and business stakeholders.
- Strong documentation, troubleshooting, communication, and operational ownership skills.
Preferred Qualifications
- Experience supporting generative AI, LLM platforms, AI agents, RAG patterns, model deployment environments, or enterprise AI services in production.
- Experience supporting MLOps or LLMOps operating models, including deployment, monitoring, lifecycle support, and release management.
- Experience with Microsoft Copilot, Copilot Studio, Azure AI Foundry, Anthropic Claude, Google Vertex AI, Snowflake Cortex AI, or similar enterprise AI platforms.
- Azure certifications in AI, Cloud Infrastructure, Security, DevOps, or related disciplines.
- ITIL, SRE, or operational excellence experience in a global enterprise environment.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search