← all roles

Outlier · via Braintrustnew this week

AI Benchmark Researcher

Pay
$1,350/task
Type
Per-task
Spots left
100
Location
Remote
Posted
September 18, 2026

Good fit for: software engineers, developers, and programmers.

Apply via referral →

Listing verified on the platform’s official board. Last checked September 25, 2026.

Bring your expertise to the problems today’s most advanced AI agents still struggle to solve.

We’re looking for experienced professionals to develop challenging tasks for an AI benchmark. You’ll draw on problems from your own field and turn them into clearly defined challenges that an AI agent must solve from scratch in a Linux terminal.

What you’ll do

  • Identify difficult, meaningful problems in your area of expertise.
  • Translate those problems into tasks that can be solved in a Linux terminal.
  • Design challenges that expose limitations in advanced AI agents. Each task is tested against three leading models, with three attempts per model, and must prove difficult enough that the models mostly fail.

Who we’re looking for

  • A completed master’s degree or higher. Candidates with a bachelor’s degree, at least 10 years of relevant professional experience and a publication may also be considered.
  • At least 10 years of professional experience in a technical domain.
  • Expertise in coding, machine learning, systems, cybersecurity or hardware.
  • At least one academic or professional publication.
  • Practical comfort working in a Linux terminal.
  • Currently based in the United States.

Project details

  • Task-based project expected to run for at least seven weeks.
  • Onboarding includes two courses and a graded assessment, taking approximately 90 minutes.
  • A live onboarding webinar is also required.

Compensation

  • Earn up to $1,350 per task, with opportunities for higher rates based on task quality and submission volume.
  • Tasks take six to seven hours on average, although initial tasks may take longer as you get familiar with the process.
Applying via Braintrust: Braintrust is a zero-fee talent network — you keep 100% of your rate. This role is posted by the listed company; apply through Braintrust.

Apply via referral →

Similar live roles

Outlier
PhD in Quantitative Finance, Remote AI Research Evaluator
$50–$150/hr
Outlier
STEM PhD Expert for AI Reasoning & Evaluation
$50–$150/hr
Outlier
Expressive Voice AI Trainer - Simplified Chinese
$20/hr
Outlier
Italian Expressive Voice AI Trainer
$20/hr