Back to search
BairesDev Linkedin · Posted 13d ago

ML Engineer (GPU) - Remote Work | REF#299894

Mexico

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

At BairesDev®, we've been leading the way in technology projects for over 15 years. We deliver cutting-edge solutions to giants like Google and the most innovative startups in Silicon Valley.

Our diverse 4,000+ team, composed of the world's Top 1% of tech talent, works remotely on roles that drive significant impact worldwide.

When you apply for this position, you're taking the first step in a process that goes beyond the ordinary. We aim to align your passions and skills with our vacancies, setting you on a path to exceptional career development and success.

As an ML Engineer (GPU), you will specialize in multi-GPU and multi-node training to scale complex models efficiently. You will act as the link between machine learning research and high-performance hardware, ensuring that models are optimized for the Nvidia ecosystem and mixed-precision execution.

What You'll Do

  • Design and implement distributed training strategies using PyTorch, DDP, and FSDP to scale model training across multiple nodes.
  • Optimize model performance and memory usage through mixed-precision training and DeepSpeed integration.
  • Profile and debug GPU kernels and workloads using Nvidia Nsight to identify and resolve performance bottlenecks.
  • Develop high-performance inference pipelines leveraging TensorRT or Triton Inference Server for production deployment.
  • Implement low-level optimizations and maintain familiarity with CUDA and GPU architectures to maximize hardware utilization.

What We Are Looking For

  • 4+ years of experience in Machine Learning Engineering, Software Engineering, or Systems Programming with a focus on GPU acceleration.
  • Proven expertise in distributed training using PyTorch, specifically DDP, FSDP, and DeepSpeed.
  • Proficiency in Nvidia GPU profiling with Nsight and model optimization using Triton or TensorRT.
  • Deep understanding of multi-GPU and multi-node training strategies, including mixed-precision optimization.
  • Advanced proficiency in English.

How we do make your work (and your life) easier:

  • 100% remote work (from anywhere).
  • Excellent compensation in USD or your local currency if preferred
  • Hardware and software setup for you to work from home.
  • Flexible hours: create your own schedule.
  • Paid parental leaves, vacations, and national holidays.
  • Innovative and multicultural work environment: collaborate and learn from the global Top 1% of talent.
  • Supportive environment with mentorship, promotions, skill development, and diverse growth opportunities.

Join a global team where your unique talents can truly thrive and make a significant impact!

Apply now!

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search