- Experience
- 5+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 4 hours ago
- Work mode
- Work from home
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Role
This is a full-time remote position hiring a Senior AI Engineer focused on benchmarking AI capabilities in software engineering tasks. The role entails designing and implementing multi-agent benchmark scenarios that mimic real-world open-source code adjustments such as bug fixes, migrations, and refactoring. The objective is to assess how well AI agents comprehend extensive codebases, apply accurate changes, and deliver verifiable, functional outcomes.
Key Responsibilities
- Create and construct multi-agent benchmark tasks reflecting authentic open-source code modifications including fixing bugs, migrating code, and refactoring changes.
- Utilize the Harbor evaluation framework to execute and verify benchmarks within isolated Docker containers.
- Produce well-defined task instructions clarifying file paths, function signatures, expected behaviors, and applicable constraints.
- Develop Python scripts to automatically verify the correctness of code alterations generated by AI agents.
- Break down complex engineering problems across several specialized AI agents and optimize benchmark tasks within containerized environments.
Required Qualifications
- Minimum of five years experience developing software in Python and JavaScript.
- Proven expertise with AI-driven coding benchmarks such as SWE-bench or Terminal-bench.
- Strong proficiency in reading and maneuvering through expansive open-source codebases, particularly frameworks like Django, Flask, FastAPI, or Node.js.
- Experience with Git version control systems, including managing pull requests, diffs, cherry-picking, and specific commits.
- Hands-on knowledge of Docker, including authoring Dockerfiles and managing container builds.
Additional Information
This position provides a distinctive chance to collaborate with an international leader in technology and information sectors, working on forefront AI systems that address critical mission challenges. The tasks support the evaluation of state-of-the-art AI agents in realistic software engineering contexts.
Equal Opportunity
We maintain a policy of unbiased recruitment, valuing skills and expertise over background or prior work history. Applications are assessed solely based on technical competencies and qualifications.
Level
Senior