AI Infra Staff Researcher
Indexed description
Lenovo is a US$83 billion revenue global technology powerhouse, ranked #153 in the Fortune Global 500, and serving millions of customers every day in 180 markets. Focused on a bold vision to deliver Smarter Technology for All, Lenovo has built on its success as the world’s largest PC company with a full-stack portfolio of AI-enabled, AI-ready, and AI-optimized devices (PCs, workstations, smartphones, tablets), infrastructure (server, storage, edge, high performance computing and software defined infrastructure), software, solutions, and services. Lenovo’s continued investment in world-changing innovation is building a more equitable, trustworthy, and smarter future for everyone, everywhere. Lenovo is listed on the Hong Kong stock exchange under Lenovo Group Limited (HKSE: 992) (ADR: LNVGY).
This transformation together with Lenovo’s world-changing innovation is building a more inclusive, trustworthy, and smarter future for everyone, everywhere. To find out more visit www.lenovo.com, and read about the latest news via our StoryHub.
- Please Note* This is a hybrid role in Morrisville, NC. This candidate will be required to work onsite three days a week.
The successful candidate will independently own substantial research and development workstreams, build production-quality software, characterize AI workloads, diagnose infrastructure issues, and develop cross-layer optimization technologies spanning GPUs and other accelerators, CPUs, memory, storage, networking, system software, data pipelines, and AI frameworks.
Key Responsibilities
- Research and develop technologies for AI compute and data infrastructure, distributed AI systems, and intelligent infrastructure management.
- Design and implement production-quality software, system components, services, APIs, diagnostic tools, and scalable data-processing pipelines.
- Characterize AI training, inference, and data-processing workloads using profiling, tracing, benchmarking, telemetry, logs, and hardware performance counters.
- Diagnose performance bottlenecks and reliability issues across GPUs, accelerators, CPUs, memory hierarchy, storage, networking, operating systems, runtimes, and AI frameworks.
- Develop hardware/software co-optimization solutions for GPU utilization, workload scheduling, resource allocation, memory and cache management, communication, data movement, storage access, and model execution.
- Optimize large-scale data ingestion, preprocessing, transformation, storage, retrieval, and delivery for AI training, inference, and analytics workloads.
- Build intelligent infrastructure diagnostics for anomaly detection, root-cause analysis, performance regression detection, system health assessment, capacity forecasting, and predictive maintenance.
- Develop fault-tolerance and resilience mechanisms, including fault detection and isolation, checkpointing, recovery, retry, failover, graceful degradation, and automated remediation.
- Apply machine learning and deep learning to workload modeling, performance prediction, resource optimization, failure prediction, and operational decision-making.
- Apply time-series analysis and signal processing to infrastructure telemetry, event detection, change-point detection, workload forecasting, and system health monitoring.
- Apply causal inference to performance attribution, root-cause analysis, intervention evaluation, and infrastructure optimization.
- Develop knowledge graphs to model infrastructure topology, hardware/software dependencies, workloads, operational events, and failure relationships.
- Optimize systems for throughput, latency, scalability, availability, resource utilization, energy consumption, and total cost of ownership.
- Collaborate with hardware, systems, software, architecture, and product teams to transition research technologies into Enterprise AI and Personal AI products.
- Contribute to patents, invention disclosures, technical publications, internal reports, and reusable software assets.
- Provide technical guidance and mentorship to junior researchers and engineers.
- Bachelor's degree in computer science, computer engineering, artificial intelligence, electrical engineering, applied mathematics, or a related field, or equivalent practical experience.
- Three or more years of relevant experience in AI compute and data infrastructure, machine learning systems, distributed systems, data platforms, performance engineering, reliability engineering, or advanced software development.
- Strong programming skills in Python, C++, Java, Go, Rust, Scala, or a comparable language.
- Demonstrated ability to design, implement, test, debug, profile, and optimize reliable software systems.
- Experience with system profiling, telemetry analytics, observability, performance diagnosis, or failure analysis.
- Technical expertise in at least two of the following areas:
- Machine learning or deep learning
- GPU or accelerator optimization
- Distributed training or inference systems
- Large-scale data processing
- Hardware/software co-optimization
- Time-series modeling or signal processing
- Infrastructure reliability and fault tolerance
- Causal inference
- Knowledge graphs or graph machine learning
- Strong analytical, experimental, communication, and cross-functional collaboration skills.
- Experience with PyTorch, TensorFlow, JAX, CUDA, ROCm, Spark, Flink, Ray, Kafka, Kubernetes, or related technologies.
- Experience with cloud, edge, on-premises, or hybrid AI infrastructure.
- Experience delivering research prototypes or advanced software into production environments.
- Publications, patents, open-source contributions, or demonstrated product impact.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search