Please submit your CV in English and indicate your level of English proficiency. Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.
What this opportunity involves
We are building a dataset to evaluate AI coding agents—how well a model handles real-world tasks performed by developers. You will create challenging tasks and evaluation criteria in realistic simulated environments:
– Build realistic developer environments: a virtual company including a codebase, infrastructure, and context (tickets, documentation, conversations) that reflects a believable development history
– Create tasks from intermediate states of these environments: craft the prompt, define what “solved” means, and ensure the task can be completed by an AI agent
– Write tests to confirm agent solutions: accept all valid approaches and reject incorrect ones, without being too strict or too lenient
– Iterate on tasks and tests based on QA feedback: review agent outputs, analyze failures, and refine the evaluation until it is fair and reliable
What this is NOT
– Not data labeling
– Not prompt engineering
– Not writing code from scratch: the agent writes most of the code; you guide and evaluate
What we look for
– 5+ years of software development experience
– Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
– Experience writing tests (functional and integration)
– English proficiency (B2+)
Why this is hard
Frontier models are already strong at coding. However, creating tasks that truly challenge the best models is difficult. You need a deep understanding of where models fail and which scenarios distinguish a strong solution from a weak one. Since tasks can have many correct solutions, writing tests that accept all valid answers while rejecting incorrect ones is harder than it may seem.
How it works
Apply → Complete qualification(s) → Join a project → Finish tasks → Get paid
Effort estimate
Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate rather than a fixed schedule: you decide when and how to work. You must submit tasks by the deadline and meet the acceptance criteria to be approved.
Compensation
Up to a $50/hour equivalent, depending on your level and pace. Each task is estimated at about 20 hours; you set your own schedule.
To apply for this job, please visit the application page

