- 経験
- 4–5 yrs
- 給料
- USD 75 – USD 75 / hour
- 求人情報
- 1
- 投稿済み
- 3時間前
- 作業モード
- 在任中
- 資格
- Open to professionals with suitable experience in software development and English proficiency at B2 level or higher.
- 再開する
- 応募必須
勤務地
仕事内容
About the Role
Mindrift offers project-based AI opportunities where specialists contribute to advancing AI systems for leading technology firms. As a Python Engineer Freelance AI Trainer, you'll engage in designing rigorous evaluation tasks that test AI coding agents not just for task completion but for safe and ethical coding practices.
Key Responsibilities
- Create realistic simulated developer environments that mimic a functioning company, complete with codebases, infrastructure, and contextual artifacts like tickets, documentation, and conversations.
- Develop task scenarios combining legitimate development goals with tempting unsafe shortcuts, such as policy breaches or scope manipulations.
- Construct comprehensive tests assessing whether AI agents execute tasks properly, recognizing unsafe or corner-cutting behavior beyond simple output correctness.
- Iterate on task designs and testing protocols by analyzing agent outputs and quality assurance feedback to ensure evaluations are accurate and robust.
What This Role is Not
- This is not a data labeling role.
- It does not involve prompt engineering activities.
- The position is not focused on cybersecurity or red-teaming; no attacker scenarios are involved. Though cybersecurity experience is beneficial, the role emphasizes software engineering over security testing.
- You will not be responsible for writing majority code; AI agents generate the code, while you design scenarios and evaluate results.
Required Qualifications
- Minimum of 4 to 5 years professional software development experience.
- Proficient in Python and JavaScript/TypeScript.
- Solid expertise in designing functional and integration tests that differentiate between safe and unsafe code completions.
- Hands-on familiarity with AI coding assistants like Claude Code, GitHub Copilot CLI, Codex, or comparable tools.
- Experience as a user with GitHub pull requests and continuous integration workflows.
- Broad technology stack knowledge, including backend systems, databases, CI pipelines, and deployment scripts is advantageous but not mandatory.
- English language proficiency at B2 level or higher.
Challenges of the Role
Developing evaluation tasks that challenge cutting-edge AI coding models is complex. You must design situations where the easier route tempts unsafe coding decisions, then devise tests that accurately identify those unsafe shortcuts, recognizing multiple valid solutions.
Work Expectations and Compensation
Task workload during active project phases is roughly 20-25 hours per week, though this is an estimate and not guaranteed. Completion deadlines and acceptance criteria must be met for task approval.
Compensation can reach up to $75 per hour depending on contribution level and pace. Pay rates may vary across different projects.
Application Instructions
Applicants should submit their CV in English and clearly state their English proficiency level.