SWE-Bench AI Task Auditor - Freelance AI Trainer Project
Agency · World Wide - Remote
About this role
Project Overview We are sourcing experienced technical specialists in Software Engineering (SWE-Bench) to audit tasks used to train and evaluate AI systems. The objective of this project is to ensure that AI training workflows and tasks are technically rigorous, practical, and highly accurate within your specific area of expertise.
Project Deliverables & Scope
Task Evaluation: Assess whether software engineering tasks are technically accurate, realistic, solvable, reproducible, and supported by reliable tests and evaluation criteria.
Technical Auditing & Feedback: Provide clear, actionable feedback on any identified codebase integration issues, test failures, or logic errors within the tasks.
Required Expertise
Demonstrable professional experience and deep technical knowledge in software engineering (including complex codebase navigation, real-world application development, and SWE-Bench standards). (Note: Candidates will be considered for one specialty based on their experience; expertise across other domains is not required).…
Summary from Agency's official Greenhouse career feed — read the full description on the original posting ↗
The full job description lives on Agency's official careers page. AI Stack Jobs links you straight to it — we never sit between you and the employer.
More AI/ML roles at Agency
- 2026-09-04 World Wide - Remote Remote
- 2026-09-04 World Wide - Remote Remote
- 2026-09-04 World Wide - Remote Remote
- 2026-09-04 World Wide - Remote Remote
- 2026-09-04 World Wide - Remote Remote
- 2026-09-04 World Wide - Remote Remote