Back to search
Mercor Builtin · Indexed 2026-08-18

Research Engineer – Benchmarking

San Francisco, CA, USA

Mid level Builtin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Mercor Research Engineer – Benchmarking 10 Hours AgoSaved In-Office San Francisco, CA, USA Mid level Mid levelArtificial Intelligence • SoftwareDesign, implement, and operate large-scale evaluation and benchmarking pipelines for LLMs. Build rubrics, automated scorers, dashboards, and failure-analysis workflows to quantify model errors, guide data curation, and inform training and reward-design decisions. Collaborate with researchers, applied-AI teams, and data producers to prioritize and scale evaluations aligned with product and research goals.Top Skills: APIsCloud PlatformsLlmsNoSQLSQL

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search