AI Lab Tech Engineer
Indexed description
About GSPANN:
Headquartered in California, U.S.A., GSPANN is a leading provider of consulting and IT services to global clients. We specialize in helping clients transform their IT capabilities, optimize business practices, and drive operational efficiency across industries such as retail, high-technology, and manufacturing. With five global delivery centers and over 1,900 employees, we combine the personalized approach of a boutique consultancy with the extensive capabilities of a large IT services firm.
Detailed Job Description:
Job Title: AI Lab Tech Engineer
Job Type & Duration: Long-Term Contract
Job Location: Fremont, CA (Onsite/Hybrid Work)
Job Summary:
We are looking for a highly skilled AI Lab Tech Engineer / Infrastructure Engineer to design, build, and operate the execution infrastructure that powers advanced Reinforcement Learning (RL) environments and AI agent workloads.
This role focuses on solving complex systems and infrastructure challenges that enable AI agents to operate coherently across realistic, multi-tool environments for extended periods. The ideal candidate will have strong expertise in distributed systems, containerization, sandboxing, execution infrastructure, performance optimization, observability, and large-scale computing.
You will build the platform that enables environments to be executed reliably at training scale, with capabilities such as snapshotting, restoration, inspection, branching, failure recovery, scheduling, and resource optimization. You will work closely with research and data teams, as well as frontier AI labs and enterprise customers, to transform environment designs into reliable production infrastructure.
Key Responsibilities:
- Design and own the sandboxing and execution layer that AI/RL environments run within.
- Build systems to snapshot and restore environment state, including disk, processes, and where applicable memory and accelerator state.
- Enable environments to be paused, resumed, inspected, and branched rather than treating each rollout as a one-time execution.
- Develop mechanisms to detect rollout failure modes early, including infrastructure faults, reward hacks, and fairness-related issues, and enable recovery from known-good states.
- Extend execution infrastructure to support long-horizon and multi-node environments where AI agents operate across multiple tools and services for hours or days.
- Own platform performance across throughput, latency, reliability, and cost-per-rollout at scale.
- Optimize resource utilization, scheduling, and execution efficiency to maximize the number of environment rollouts per dollar.
- Profile and eliminate bottlenecks across the infrastructure stack, including container startup, execution, scheduling, and environment teardown.
- Build comprehensive observability for monitoring and troubleshooting thousands of concurrent, long-running rollouts.
- Build and maintain frameworks for specifying, packaging, deploying, and managing RL environments.
- Develop tools that enable researchers and environment authors to debug specific failures across large volumes of long-running agent traces.
- Deploy and manage large and small AI/ML models on on-premises hardware.
- Scale research prototypes into production-grade infrastructure with reproducible workflows, strong testing, and high engineering standards.
- Develop documentation, tooling, and engineering practices that enable internal teams and external users to build effectively on the platform.
- Collaborate closely with research, data, engineering, and enterprise customer teams to translate AI research requirements into scalable infrastructure solutions.
Qualifications:
- Strong track record building production systems or research infrastructure at scale, including distributed systems, execution engines, container infrastructure, sandboxing platforms, or similar technologies.
- Deep understanding of systems-level technologies including containers, isolation, namespaces, cgroups, virtual machines, filesystems, process management, and state management.
- Experience with sandboxing technologies and architectures such as gVisor, Firecracker, or similar isolation technologies.
- Strong experience optimizing systems for performance, scalability, resource utilization, scheduling, and cost efficiency.
- Proficiency with cloud platforms such as GCP and AWS and experience working with distributed computing environments.
- Experience deploying large and small AI/ML models on on-premises hardware.
- Strong engineering fundamentals with a systematic approach to testing, validation, reliability, and production operations.
- Strong Python programming skills; experience with systems programming languages such as Rust, Go, or C++ is a plus.
- Ability to work effectively in ambiguous and rapidly evolving technical environments.
- Experience using modern AI-assisted development tools such as Claude Code effectively is a plus.
- Excellent communication and collaboration skills, with the ability to work closely with AI research teams, infrastructure engineers, and enterprise customers.
- Ability to translate complex research requirements into practical infrastructure and engineering solutions.
- Comfortable presenting technical designs, solutions, and engineering work to diverse technical and non-technical audiences.
Working at GSPANN
GSPANN is a diverse, prosperous, and rewarding place to work. We provide competitive benefits, educational assistance, and career growth opportunities to our employees. Every employee is valued for their talent and contribution. Working with us will give you an opportunity to work globally with some of the best brands in the industry.
The company does and will take affirmative action to employ and advance in employment of individuals with disabilities and protected veterans and to treat qualified individuals without discrimination based on their physical or mental disability status. GSPANN is an equal opportunity employer for minorities/females/veterans/disabled.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search