AI Infra Advisory Researcher
Indexed description
Lenovo is a US$83 billion revenue global technology powerhouse, ranked #153 in the Fortune Global 500, and serving millions of customers every day in 180 markets. Focused on a bold vision to deliver Smarter Technology for All, Lenovo has built on its success as the world’s largest PC company with a full-stack portfolio of AI-enabled, AI-ready, and AI-optimized devices (PCs, workstations, smartphones, tablets), infrastructure (server, storage, edge, high performance computing and software defined infrastructure), software, solutions, and services. Lenovo’s continued investment in world-changing innovation is building a more equitable, trustworthy, and smarter future for everyone, everywhere. Lenovo is listed on the Hong Kong stock exchange under Lenovo Group Limited (HKSE: 992) (ADR: LNVGY).
This transformation together with Lenovo’s world-changing innovation is building a more inclusive, trustworthy, and smarter future for everyone, everywhere. To find out more visit www.lenovo.com, and read about the latest news via our StoryHub.
- Please Note* This is a hybrid role in Morrisville, NC. This candidate will be required to work onsite three days a week.
The successful candidate will define technical approaches, lead complex research initiatives, and make architecture-level decisions across GPUs and other accelerators, CPUs, memory, storage, networking, distributed systems, data platforms, AI frameworks, and application workloads.
Key Responsibilities
- Define technical directions and lead major research and development initiatives in AI compute and data infrastructure, distributed AI systems, and intelligent infrastructure management.
- Identify high-impact technical opportunities based on infrastructure challenges, emerging technologies, product requirements, and business value.
- Architect end-to-end AI infrastructure solutions spanning hardware, system software, data platforms, distributed training and inference, and application workloads.
- Lead hardware/software co-analysis and co-optimization across GPUs, accelerators, CPUs, memory hierarchy, storage, networking, runtimes, frameworks, and AI applications.
- Define optimization strategies for GPU utilization, workload placement, resource orchestration, memory and cache efficiency, communication, data movement, storage access, and model execution.
- Lead the architecture and optimization of large-scale data pipelines for data ingestion, preprocessing, transformation, storage, retrieval, and delivery to AI workloads.
- Define intelligent observability and diagnostic technologies for anomaly detection, root-cause analysis, performance regression, capacity planning, system health assessment, and predictive maintenance.
- Develop resilient and fault-tolerant infrastructure architectures, including failure isolation, checkpointing and recovery, redundancy, retry, failover, graceful degradation, and automated remediation.
- Apply machine learning and deep learning to system modeling, infrastructure control, workload forecasting, resource optimization, failure prediction, and operational intelligence.
- Apply time-series modeling and signal processing to telemetry analytics, event detection, change-point detection, capacity forecasting, and system health monitoring.
- Apply causal inference and causal discovery to root-cause analysis, performance attribution, intervention evaluation, and automated decision-making.
- Define knowledge graph architectures for modeling infrastructure topology, hardware/software dependencies, workloads, operational events, and failure relationships.
- Make architecture-level trade-offs involving performance, scalability, reliability, availability, energy consumption, cost, security, and maintainability.
- Lead technical design reviews, architecture reviews, performance investigations, and resolution of complex cross-layer system issues.
- Provide hands-on technical guidance in algorithm design, software implementation, system optimization, experimental validation, and production deployment.
- Establish reusable frameworks, engineering practices, evaluation methodologies, and technical standards for AI infrastructure development.
- Partner with research, engineering, architecture, product, and business teams to transition research into Enterprise AI and Personal AI platforms and products.
- Mentor Staff Researchers and engineers and contribute to patents, publications, technical standards, and differentiated intellectual property.
- Bachelor's degree in computer science, computer engineering, artificial intelligence, electrical engineering, applied mathematics, or a related field, or equivalent practical experience.
- Five or more years of relevant experience in AI compute and data infrastructure, machine learning systems, distributed systems, data platforms, performance engineering, reliability engineering, or advanced software development.
- Demonstrated leadership of technically complex research, architecture, or system development initiatives.
- Strong hands-on programming and system architecture capabilities.
- Deep expertise in at least three of the following areas:
- Machine learning or deep learning
- GPU or accelerator optimization
- Distributed AI training and inference
- Large-scale and streaming data processing
- Hardware/software co-optimization
- Time-series modeling or signal processing
- Infrastructure diagnostics, reliability, and fault tolerance
- Causal inference
- Knowledge graphs or graph-based reasoning
- Proven record of delivering advanced software, infrastructure, or research technologies into products or production environments.
- Ability to resolve ambiguous, cross-layer technical problems and make informed architecture-level trade-offs.
- Strong technical communication, mentorship, and cross-organizational influence.
- Experience architecting enterprise-scale cloud, edge, on-premises, or hybrid AI infrastructure.
- Experience with PyTorch, TensorFlow, JAX, CUDA, ROCm, Spark, Flink, Ray, Kafka, Kubernetes, or related technologies.
- Experience with MLOps, observability platforms, distributed computing, data lakehouse architectures, or cloud-native systems.
- Strong publication record, patent portfolio, open-source contributions, or demonstrated commercial product impact.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search