Skip to main content
Loading
loadingbar
Loading, Please wait..!!

Senior Python Engineer - AI Coding Agent Evaluation (Freelance)

  • Job type Posted on: Jul 19, 2026
  • Experience level Mind Rift
  • Employment type New York, New York
  • Remote status Salary: $416,000 per year
  • Employment type Onsite
  • Salary Full-time

Point Apply Here APPLY LATER

Curious about compensation?

Explore the historical salary trends, average pay, and estimated compensation for Senior Python Engineer - AI Coding Agent Evaluation (Freelance) roles in New York.

View Salary Guide →

Job Title :

Senior Python Engineer - AI Coding Agent Evaluation (Freelance)

Job Type :

Full-time

Job Location :

New York New York United States

Remote :

No

Jobcon Logo Job Description :

Overview Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment. This opportunity involves building a dataset to evaluate AI coding agents and how well a model handles real-world developer tasks. Responsibilities Build realistic developer environments: a virtual company with a codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks from intermediate states of these environments: craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent Write tests that verify agent solutions: accept all valid approaches and reject incorrect ones, neither too strict nor too lenient Iterate on tasks and tests based on QA feedback: review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT Not data labeling Not prompt engineering Not writing code from scratch - the agent writes most of the code; you guide and evaluate What we look for 8+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Why this is hard Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds. How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid. Effort estimate Timeline & Expectations Tasks for this project are estimated to take 30 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted. Compensation Up to $200/hr equivalent , depending on level and pace. Tasks are estimated at ~30 hours each; you set your own schedule. #J-18808-Ljbffr

Jobcon Logo Position Details

Posted:

Jul 19, 2026

Reference Number:

14660_ADA0D2C65C943FBB60C6CDC7596E09E5

Employment:

Full-time

Salary:

Not Available

City:

New York

Job Origin:

APPCAST_CPC

Share this job:

  • linkedin

Jobcon Logo
A job sourcing event
In Dallas Fort Worth
Aug 19, 2017 9am-6pm
All job seekers welcome!

Senior Python Engineer - AI Coding Agent Evaluation (Freelance)    Apply

Click on the below icons to share this job to Linkedin, Twitter!

Overview Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment. This opportunity involves building a dataset to evaluate AI coding agents and how well a model handles real-world developer tasks. Responsibilities Build realistic developer environments: a virtual company with a codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks from intermediate states of these environments: craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent Write tests that verify agent solutions: accept all valid approaches and reject incorrect ones, neither too strict nor too lenient Iterate on tasks and tests based on QA feedback: review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT Not data labeling Not prompt engineering Not writing code from scratch - the agent writes most of the code; you guide and evaluate What we look for 8+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Why this is hard Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds. How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid. Effort estimate Timeline & Expectations Tasks for this project are estimated to take 30 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted. Compensation Up to $200/hr equivalent , depending on level and pace. Tasks are estimated at ~30 hours each; you set your own schedule. #J-18808-Ljbffr

Loading
Please wait..!!