Job Description

Role Overview

We're building a dataset to evaluate AI coding agents by creating challenging tasks and evaluation criteria within realistic simulated environments. You'll design tasks, write tests, and iterate on tasks and tests based on QA feedback.

What You Will Do

Create realistic developer environments, design tasks, write tests, and iterate on tasks and tests to evaluate AI agent solutions.

Why It Might Be a Fit

You need 5+ years in software development, experience writing tests, and English proficiency (B2+). You'll work on challenging tasks that require understanding where models fail and what scenarios reveal the difference between good and bad solutions.

Requirements

  • 5+ years in software development
  • Experience writing tests (functional, integration)
  • English proficiency - B2+

Benefits

  • Up to $50/hr equivalent
  • Flexible schedule
  • 20 hours per task
Apply now
Report job

More job openings