Research Engineer – Benchmarking
Indexed description
Mercor Research Engineer – Benchmarking 10 Hours AgoSaved In-Office San Francisco, CA, USA Mid level Mid levelArtificial Intelligence • SoftwareDesign, implement, and operate large-scale evaluation and benchmarking pipelines for LLMs. Build rubrics, automated scorers, dashboards, and failure-analysis workflows to quantify model errors, guide data curation, and inform training and reward-design decisions. Collaborate with researchers, applied-AI teams, and data producers to prioritize and scale evaluations aligned with product and research goals.Top Skills: APIsCloud PlatformsLlmsNoSQLSQL
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search