SoftServe
Linkedin · Posted 12d ago
Lead Nvidia AI Consultant
Continue to application
Add your email once, then Caio opens the original posting.
Indexed description
About The RoleIn this role, you will combine deep hands-on engineering expertise in GPU inference, LLM deployment, and modern AI infrastructure with the consulting gravitas to lead client engagements, shape enterprise architecture decisions, and drive pre-sales. This is not a pure advisory role — you will own technical outcomes, drive AI adoption, and shape go-to-market solutions working closely with clients, Nvidia stakeholders, and internal engineering teams.
Responsibilities
- Own and lead AI engagements end-to-end — from discovery, assessment, and strategy definition through architecture design and production implementation, including clearly scoped implementation plans and delivery outcomes
- Translate complex business and operational challenges into AI use-case definitions, solution roadmaps, and corresponding reference architectures aligned with client scope and strategic objectives
- Design and validate production-grade AI architectures leveraging NVIDIA technologies including NeMo, NIM, Triton, Riva, DeepStream, Metropolis, and Omniverse across cloud and on-premises environments
- Define reference architectures for GenAI and Agentic AI solutions, advising on GPU-accelerated infrastructure design, Kubernetes-based workload orchestration, and MIG/vGPU partitioning strategies
- Profile and benchmark AI inference deployments to validate GPU utilization, memory footprint, and cost-efficiency against target SLAs — not just functional correctness
- Continuously tune deployment configurations — batch size, concurrency, tensor and pipeline parallelism, quantization level, and KV-cache settings — based on profiling data to hit optimal latency, throughput, and cost tradeoffs
- Benchmark LLM serving stacks across frameworks such as vLLM, TGI, and Triton+TensorRT-LLM, measuring throughput, TTFT, TPOT, and tokens/sec/GPU to drive informed infrastructure decisions
- Lead pre-sales activities including proposals, solution positioning, technical discovery sessions, and customer workshops, and drive GenAI proof-of-concept initiatives that demonstrate clear and measurable business value
- Contribute to NVIDIA alliance go-to-market strategy by shaping industry offerings, reusable accelerators, demos, and solution blueprints
- Drive thought leadership through whitepapers, technical blogs, conference presentations, and industry events while actively mentoring engineering and consulting teams on NVIDIA ecosystem technologies and modern AI/ML practices
- 6+ years of experience in AI consulting, GenAI/Agentic AI development, Machine Learning, or Deep Learning with a demonstrable track record of leading client-facing engagements end-to-end
- Strong hands-on expertise in Generative AI, Agentic AI, multimodal AI, transformers, LLMs, and VLMs with practical experience across the full AI lifecycle including experimentation, fine-tuning, optimization, deployment, inference, and monitoring
- Hands-on experience with Python and modern AI/ML frameworks including PyTorch, TensorFlow, Pandas, NumPy, and Hugging Face
- Substantial hands-on experience with at least three NVIDIA ecosystem platforms including NeMo, NIM, Riva, Metropolis, Omniverse, Triton, DeepStream, or TensorRT-LLM
- Hands-on experience deploying AI/ML workloads on Kubernetes including Helm, Operators, GPU device plugins, NVIDIA GPU Operator, and MIG/vGPU partitioning, alongside practical experience with self-hosted inference stacks such as vLLM or Ollama
- Working knowledge of model quantization techniques and inference optimization with hands-on experience using GPU profiling tools including NVIDIA Nsight Systems, Nsight Compute, nvidia-smi, DCGM metrics, PyTorch Profiler, and vLLM/Triton metrics endpoints
- Demonstrated ability to read GPU utilization signals, diagnose compute-bound, memory-bound, or I/O-bound bottlenecks, and translate findings into actionable architecture or configuration changes
- Experience designing and deploying AI solutions in at least one major cloud environment — AWS, Azure, or GCP — with strong understanding of TensorRT, Triton Inference Server, CUDA, DeepStream, and ONNX
- Strong advisory and stakeholder management capabilities with experience presenting to executive and leadership audiences and leading pre-sales engagements including proposals, technical discovery sessions, workshops, and solution positioning
- Deep understanding of enterprise architecture, distributed systems, Big Data, SDLC, MLOps, and AI governance practices combined with strong communication, analytical, and problem-solving skills
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search
Want help applying to roles like this?
Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search