Job Description

Role Overview

We're building a dataset to evaluate AI coding agents by creating challenging tasks and evaluation criteria within realistic simulated environments. You'll design tasks, write tests, and iterate on tasks and tests based on QA feedback.

What You Will Do

Create challenging tasks and evaluation criteria, design tasks from intermediate states of environments, write tests that verify agent solutions, and iterate on tasks and tests based on QA feedback.

Why It Might Be a Fit

You'll need 5+ years of software development experience, core stack experience in Python, JavaScript/TypeScript, Docker, Postgres, Kafka, and Redis, and experience writing tests (functional, integration).

Requirements

  • 5+ years in software development
  • Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
  • Experience writing tests (functional, integration)
  • English proficiency - B2+

Benefits

  • Up to $50/hr equivalent
  • Flexible schedule
  • 20 hours per task, set your own pace
Apply now
Report job

More job openings