Job Description

Role Overview

We're building a dataset to evaluate AI coding agents by creating challenging tasks and evaluation criteria within realistic simulated environments. You'll design tasks, write tests, and iterate on tasks and tests based on QA feedback.

What You Will Do

Create tasks and evaluation criteria, write tests, and iterate on tasks and tests to ensure the evaluation is fair and robust. You'll also review agent solutions, analyze failures, and refine the evaluation process.

Why It Might Be a Fit

You need to have 5+ years of software development experience, with a strong understanding of Python, JavaScript/TypeScript, Docker, Postgres, Kafka, and Redis. You should also have experience writing tests and be proficient in English (B2+).

Requirements

  • 5+ years in software development
  • Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
  • Experience writing tests (functional, integration)
  • English proficiency - B2+

Benefits

  • Up to $40/hr equivalent
  • Flexible schedule
  • Tasks are estimated at ~20 hours each
Apply now
Report job

More job openings