Back to search
Akoncagua AI Linkedin · Posted 17d ago

Senior LLM/VLM Training Data Engineer

Lakeland, Florida, United States

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Salary


$140,000–$185,000 annually


About Akoncagua AI

Akoncagua AI builds domain-specialized AI systems for technically demanding industries. Our work spans large language models, vision-language models, agentic AI, model adaptation, synthetic data, evaluation, RAG, and production AI infrastructure.

We are looking for a Senior AI Training Data Engineer with 7+ years of experience to build the datasets and data pipelines used to train domain-specialized LLMs and VLMs.

This role is focused on transforming large collections of proprietary documents, images, technical drawings, structured data, and domain knowledge into high-quality datasets for continued pretraining (CPT), supervised fine-tuning (SFT), multimodal post-training, synthetic data, and model evaluation.

What You’ll Do

• Build scalable pipelines for preparing CPT, SFT, multimodal, preference, and evaluation datasets.

• Transform PDFs, technical documents, images, OCR, tables, drawings, and structured data into model-ready training examples.

• Design domain-specific SFT task taxonomies, schemas, instruction-response formats, and dataset quality standards.

• Use frontier LLMs/VLMs as teacher models to generate synthetic training examples from proprietary domain data.

• Build automated pipelines for grounding, validating, filtering, deduplicating, and scoring synthetic and human-generated training data.

• Develop multimodal datasets combining images, OCR text, layout regions, bounding boxes, tables, semantic descriptions, and document context.

• Build text-only, image-text, and interleaved datasets for domain-specific continued pretraining and long-context model training.

• Design human-in-the-loop workflows where domain experts review, correct, and validate high-value training examples.

• Establish Gold/Silver dataset tiers, provenance, lineage, versioning, reproducibility, and auditability.

• Build held-out benchmarks and evaluation datasets while preventing train/test contamination and data leakage.

• Analyze model failures and convert them into new SFT examples, hard cases, synthetic datasets, and evaluation scenarios.

• Work closely with ML researchers and AI engineers to continuously improve model performance through better training data.

Required Qualifications

• 7+ years of experience in Data Engineering, ML Data Infrastructure, ML Engineering, AI Training Data, or a related field.

• Expert-level Python with strong software-engineering and data-processing fundamentals.

• Hands-on experience preparing datasets for LLM/VLM training, fine-tuning, post-training, or evaluation.

• Experience creating or curating SFT/instruction-response datasets.

• Strong experience with synthetic-data generation using LLMs or VLMs and automated quality validation.

• Experience processing large-scale PDF, OCR, image, table, document-layout, and multimodal datasets.

• Strong understanding of tokenization, context-length management, dataset mixing, deduplication, sampling, and train/validation/test separation.

• Experience with dataset provenance, lineage, versioning, quality metrics, and reproducible data pipelines.

• Experience with distributed data processing and cloud infrastructure, preferably AWS.

• Strong understanding of model evaluation, benchmark construction, and failure-driven dataset improvement.

Preferred Qualifications

• Experience with continued pretraining (CPT/DAPT), SFT, LoRA/QLoRA, DPO/RLHF, or other LLM post-training techniques.

• Experience preparing training data for multimodal models such as Qwen-VL/Qwen3-VL, InternVL, LLaVA, or similar VLMs.

• Experience constructing long-context datasets from 32K to 256K+ tokens.

• Experience with multimodal document understanding, OCR, visual grounding, and technical-document intelligence.

• Experience building human/expert annotation, active-learning, and model-in-the-loop data-curation systems.

• Experience creating golden datasets and domain-specific evaluation benchmarks.

• Experience working with highly technical datasets such as engineering, scientific, geospatial, CAD, financial, legal, or medical documents.

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search