Senior LLM/VLM Training Data Engineer
Indexed description
Salary
$140,000–$185,000 annually
Akoncagua AI builds domain-specialized AI systems for technically demanding industries. Our work spans large language models, vision-language models, agentic AI, model adaptation, synthetic data, evaluation, RAG, and production AI infrastructure.
We are looking for a Senior AI Training Data Engineer with 7+ years of experience to build the datasets and data pipelines used to train domain-specialized LLMs and VLMs.
This role is focused on transforming large collections of proprietary documents, images, technical drawings, structured data, and domain knowledge into high-quality datasets for continued pretraining (CPT), supervised fine-tuning (SFT), multimodal post-training, synthetic data, and model evaluation.
What You’ll Do• Build scalable pipelines for preparing CPT, SFT, multimodal, preference, and evaluation datasets.
• Transform PDFs, technical documents, images, OCR, tables, drawings, and structured data into model-ready training examples.
• Design domain-specific SFT task taxonomies, schemas, instruction-response formats, and dataset quality standards.
• Use frontier LLMs/VLMs as teacher models to generate synthetic training examples from proprietary domain data.
• Build automated pipelines for grounding, validating, filtering, deduplicating, and scoring synthetic and human-generated training data.
• Develop multimodal datasets combining images, OCR text, layout regions, bounding boxes, tables, semantic descriptions, and document context.
• Build text-only, image-text, and interleaved datasets for domain-specific continued pretraining and long-context model training.
• Design human-in-the-loop workflows where domain experts review, correct, and validate high-value training examples.
• Establish Gold/Silver dataset tiers, provenance, lineage, versioning, reproducibility, and auditability.
• Build held-out benchmarks and evaluation datasets while preventing train/test contamination and data leakage.
• Analyze model failures and convert them into new SFT examples, hard cases, synthetic datasets, and evaluation scenarios.
• Work closely with ML researchers and AI engineers to continuously improve model performance through better training data.
Required Qualifications• 7+ years of experience in Data Engineering, ML Data Infrastructure, ML Engineering, AI Training Data, or a related field.
• Expert-level Python with strong software-engineering and data-processing fundamentals.
• Hands-on experience preparing datasets for LLM/VLM training, fine-tuning, post-training, or evaluation.
• Experience creating or curating SFT/instruction-response datasets.
• Strong experience with synthetic-data generation using LLMs or VLMs and automated quality validation.
• Experience processing large-scale PDF, OCR, image, table, document-layout, and multimodal datasets.
• Strong understanding of tokenization, context-length management, dataset mixing, deduplication, sampling, and train/validation/test separation.
• Experience with dataset provenance, lineage, versioning, quality metrics, and reproducible data pipelines.
• Experience with distributed data processing and cloud infrastructure, preferably AWS.
• Strong understanding of model evaluation, benchmark construction, and failure-driven dataset improvement.
Preferred Qualifications• Experience with continued pretraining (CPT/DAPT), SFT, LoRA/QLoRA, DPO/RLHF, or other LLM post-training techniques.
• Experience preparing training data for multimodal models such as Qwen-VL/Qwen3-VL, InternVL, LLaVA, or similar VLMs.
• Experience constructing long-context datasets from 32K to 256K+ tokens.
• Experience with multimodal document understanding, OCR, visual grounding, and technical-document intelligence.
• Experience building human/expert annotation, active-learning, and model-in-the-loop data-curation systems.
• Experience creating golden datasets and domain-specific evaluation benchmarks.
• Experience working with highly technical datasets such as engineering, scientific, geospatial, CAD, financial, legal, or medical documents.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search