Gramian Consulting Group Himalayas · Posted 2mo ago

AI Evaluation Engineer (Data Analysis & Multi-Agent Systems)

USD Full time Remote

Continue to application Add your email once, then Caio opens the original posting.

Indexed description

We are looking for an AI Evaluation Engineer specialized in data analysis to design benchmark tasks that simulate real-world analytical workflows. The ideal candidate will have 5+ years of experience in data analysis or analytics-heavy roles, strong proficiency in Python and SQL, and experience working with real-world, messy datasets.

Requirements

Design and develop multi-agent benchmark tasks focused on complex data analysis workflows
Create or curate realistic datasets
Implement evaluation pipelines using Python and SQL
Create reproducible environments using Docker
Analyze task performance and refine for clarity, difficulty, and scoring accuracy

Benefits

Competitive salary

Originally posted on Himalayas

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search

Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.

View Managed Job Search

Gramian Consulting Group Company profile preview

Source: Himalayas
Location
Compensation: USD
Open on Caio: 7 roles

Salary insight

USD

Caio highlights salary ranges whenever the original posting exposes them. Compare similar roles as the index fills in.

Similar role details

Full time roles Remote matches Himalayas postings

Company stats

Current index details for Gramian Consulting Group, based on roles Caio has indexed from public sources.

7open roles 1sources 0markets Posted 2mo agolatest role