Data Center Business Technical SME
Indexed description
Position- Data Center Business Technical SME
Location- Texas/California
Interview- Phone and Skype
Duration- Full-Time Opportunity
Role Overview
Our Client is an AI datacenter and cloud services company operating a large fleet of high-performance GPU infrastructure designed for AI/ML training, inference, and Generative AI workloads. The business spans hyperscale colocation, enterprise infrastructure, AI-ready hosting, and network-rich connectivity services.
We are looking for a senior technology leader reporting directly to the CTO, responsible for the architecture, scalability, reliability, and evolution of Client's AI infrastructure and platform portfolio.
The role combines deep technical leadership, platform strategy, infrastructure economics, ecosystem alignment, and executive-level decision-making, with a strong focus on GPU utilization, platform differentiation, and long-term competitiveness.
Key Responsibilities
1. AI Infrastructure Architecture & Strategy
- Own end-to-end architecture for GPU compute, AI clusters, training/inference platforms, and managed AI datacenter services.
- Define reference architectures for large-scale GenAI training, high-throughput inference, and enterprise/sovereign AI deployments.
- Drive technical decisions across NVIDIA/AMD and emerging accelerators, networking, storage, interconnects, and hybrid/on-prem AI environments.
2. Platform Engineering & Delivery Leadership
- Lead AI infrastructure engineering, SRE, and platform teams.
- Drive high availability, performance, fault tolerance, SLA adherence, security, and operational excellence.
- Establish best practices for GPU scheduling, capacity management, cluster orchestration, automation, benchmarking, and performance optimization.
3. AI Platform & Services Evolution
- Own the roadmap for managed AI platforms, AI factories/private AI clouds, and industry-specific AI stacks.
- Define standardized blueprints for training and inference workloads across startup, enterprise, and GCC use cases.
- Drive platform differentiation through automation, tooling, MLOps integrations, developer experience, and self-service capabilities.
4. Enterprise, Startup & Ecosystem Enablement
- Serve as the senior technical sponsor for strategic enterprise/GCC customers, AI-native startups, system integrators, and ISV partners.
- Partner with BD and GTM teams to develop technically sound and commercially viable solutions.
- Support executive-level technical discussions with CTOs and AI leaders and enable partners through reference architectures, documentation, and joint solution designs.
5. Capacity Planning, Economics & Governance
- Partner with Finance and Operations on GPU capacity planning, expansion strategy, CapEx/OpEx tradeoffs, ROI, utilization, and cost efficiency.
- Establish governance around resource allocation, platform consumption models, and margin-aligned service design.
6. Technology Leadership & Internal Collaboration
- Act as a core member of the CTO leadership team.
- Drive alignment across datacenter operations, AI platform engineering, security, compliance, and risk teams.
- Mentor senior engineering leaders and help build the next generation of AI infrastructure talent.
Key Success Metrics
- Platform stability, performance, and scalability
- GPU capacity utilization and predictability
- Time to deploy AI platforms and clusters
- Platform adoption across startups, enterprises, and partners
- Infrastructure cost efficiency and margin contribution
Required Qualifications
- 15+ years of experience in AI infrastructure engineering, cloud platforms, datacenter-scale systems, or HPC/GPU environments.
- Proven experience building and scaling large-scale GPU/accelerator platforms and AI training/inference environments.
- Deep expertise in AI/ML workloads, GenAI economics, GPU architectures, interconnects, and training vs. inference performance tradeoffs.
- Experience working with startups/scaleups, system integrators, ISVs, large enterprises, and regulated customers.
- Track record operating at CTO-1, VP, Head of Engineering, or equivalent senior technology leadership level.
Preferred Qualifications
- Experience with sovereign cloud or regulated AI environments.
- Exposure to MLOps platforms, AI frameworks, and developer tooling.
- Experience scaling a new AI infrastructure or platform business.
- Participation in industry or open-source AI initiatives.
Ideal Leadership Profile
The ideal candidate will bring strong technical judgment and system-level thinking, with the ability to translate business objectives into scalable technology platforms. The individual should be comfortable operating in a high-growth and ambiguous environment and be credible when engaging with CTOs, AI leaders, customers, partners, and senior executives.
This is a strategic, hands-on technology leadership role with direct visibility to the CTO, offering the opportunity to shape the architecture and evolution of a growing AI infrastructure business.
Thanks,
Akshat Gupta
Senior Executive Search Consultant
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search