Back to search
Fleet AI, Inc. Linkedin · Posted 12d ago

Synthetic Data

Buffalo-Niagara Falls

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

About Fleet

Fleet studies how environments produce intelligence. We believe intelligence is an emergent property of environmental pressures — the environment determines what capabilities develop, what behaviors survive, and what "good" looks like. We work with frontier labs on post-training across modalities — building benchmarks that expose where frontier models break, training recipes that close those gaps, and scalable oversight for long-horizon agents. Backed by Sequoia Capital, Menlo Ventures, BCV, and SV Angel.

About The Role

We're looking for a Member of Technical Staff to own the synthetic data that powers our training environments. Each environment needs a purpose-built dataset — deeply understood, carefully constructed, and continually refined. This is not a data platform or data pipeline role. You'll go deep on individual datasets, treating each one as a product with its own quality bar, iteration cycle, and domain nuances.

What You'll Work On

  • Build and own synthetic data products for training environments
  • Design pipelines and tooling for generating rich, structured synthetic documents with deep context across multiple artifacts
  • Design agentic hillclimbing mechanisms to improve data realism. Use systems-level thinking to improve consistency of data.
  • Define and enforce data quality standards for each dataset — understand what "good" looks like for each domain and iterate until it's right
  • Collaborate with environment and task-writing teams to ensure data products serve the exact needs of each training environment
  • Develop scalable approaches combining code generation for structure with LLM research agents for domain-specific detail (medical, insurance, financial, etc.)

How We Work

  • You own your data seeds end-to-end — from understanding the domain to shipping a production-quality data product ready for use in an environment
  • Direct collaboration with frontier labs and internal research teams

What We're Looking For

  • Experience building a dataset that was the core product — not serving other teams' data, but owning and refining a dataset as the thing customers or systems depend on (e.g., Mapbox places/geometry, knowledge graphs, domain-specific corpora)
  • Comfort going deep on a single domain and iterating through many refinement cycles to get it right
  • Applied research background in synthetic data generation, agentic data construction, or domain-specific data pipelines

Location

On-site in San Francisco or New York. Highly competitive salary + equity.

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search