Research Engineer – AI Retrieval & LLM Fine-Tuning
Indexed description
About the RoleVinSmart Future (VFS) is building next-generation Vietnamese-centric foundation models and intelligent search systems to power the future of AI in Vietnam. We are looking for highly capable Research Engineers to drive the development of our advanced retrieval architectures and LLM post-training pipelines.
You will work at the intersection of information retrieval, applied LLM training, and scalable systems engineering. Your focus will be on building robust Search/RAG (Retrieval-Augmented Generation) pipelines and fine-tuning models to excel in complex, multi-step reasoning and retrieval tasks for Vietnamese and multilingual contexts.
Working location: Ho Chi Minh, Vietnam.
Key Responsibilities
1. Information Retrieval
- Design and implement advanced retrieval architectures, including dense/sparse hybrid search configurations and knowledge graph integrations.
- Research and train custom embedding models and cross-encoder rerankers tailored to Vietnamese semantics.
- Explore and implement state-of-the-art retrieval methodologies to tackle hard retrieval scenarios, including multi-hop reasoning, complex query understanding, and long-tail edge cases.
2. LLM Fine-Tuning (Post-Training)
- Perform Supervised Fine-Tuning (SFT) and instruction tuning to adapt foundation models for specific retrieval, extraction, and generation tasks.
- Apply standard alignment methods (e.g., DPO or basic RLHF) to steer model outputs according to product and prompt requirements.
- Manage end-to-end training runs, monitor loss metrics, evaluate model checkpoints, and troubleshoot common training stability issues (e.g., overfitting, loss spikes).
3. Large-Scale Retrieval & Systems Implementation
- Build and scale robust retrieval pipelines and vector database infrastructure (e.g., Milvus, Qdrant, Faiss) capable of handling millions of documents efficiently.
- Deploy LLMs and retrieval models (embeddings, rerankers) into production environments using standard serving frameworks like vLLM, TGI, or Triton.
- Optimize end-to-end latency, caching strategies, and throughput for high-concurrency RAG applications.
4. Evaluation & Benchmarking
- Prepare and maintain comprehensive Vietnamese and multilingual benchmark suites for both LLM generation and retrieval quality (e.g., MRR, NDCG, Recall@K).
- Implement automated evaluation pipelines to monitor the "see-saw" effects of multi-task learning and alignment interventions.
Minimum Requirements
- Bachelor’s/Master’s/PhD degree in CS/AI/ML or related fields.
- Strong Python programming and deep PyTorch experience.
- Solid understanding of transformer architectures, tokenization, and modern Information Retrieval (IR) systems.
- Hands-on experience with LLM fine-tuning techniques (LoRA, QLoRA, full-parameter tuning).
- Experience with building and scaling data pipelines or search systems.
Preferred Qualifications
- Experience with alignment techniques (RLHF, DPO) or model distillation.
- Familiarity with high-performance inference serving (vLLM, TensorRT) and production-grade vector databases.
- Deep knowledge of Vietnamese NLP nuances and enterprise search challenges.
- Familiarity with tracking tools, experiment logging, and multi-GPU distributed training frameworks (FSDP, DeepSpeed).
- Contributions to open-source AI/ML or retrieval projects.
What VFS Offers
- Opportunity to build next-generation Vietnamese LLMs and search engines.
- Access to large-scale multi-GPU clusters.
- High-growth environment bridging cutting-edge research and enterprise product deployment.
- Collaboration with strong, fast-paced AI teams.
- Competitive compensation.
How to Apply: Please send your CV and relevant links (GitHub, Google Scholar, Portfolio) to us.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search