Website:
lh2.ai
Job details:
About LH2 AI Labs
Built by second-time founders who have built and sold companies before, LH2 AI Labs is building the post-training infrastructure for frontier AI models.
We bring private, high-quality institutional datasets and vetted domain experts into frontier AI pipelines across verticals such as coding, computer use, agentic workflows, medical, audio, and more.
For AI to keep progressing, it needs high-quality training data drawn from real production use cases. The public web has already been crawled and trained on—there is limited new signal left there. That is where we come in.
Our vision is to create a world where frontier models can access high-quality data on tap, the same way they access compute today.
About the roleYou will create high-quality, SWE-bench-style benchmark tasks from real-world software repositories, and build the pipelines and tooling that make task creation faster, more reliable, and more scalable. This is a dual-mandate role: roughly half hands-on task authoring and review, and half engineering the automation that the whole team depends on.
Key responsibilities- Evaluate real-world software repositories for reproducibility, test quality and suitability for benchmark task development.
- Understand unfamiliar architectures, dependencies and services, and establish reliable local or containerized execution environments.
- Identify self-contained bugs and feature changes from issues, pull requests, commits and repository history.
- Construct complete benchmark tasks from selected bugs and feature changes, including the task setup, specification, tests and containerized execution environment.
- Write precise, implementation-neutral task specifications that define the expected behavior without revealing the solution.
- Create and review tests that measure observable behavior, cover important edge cases and allow valid alternative implementations.
- Verify that the original revision demonstrates the expected failure, the reference solution passes and unrelated behavior remains stable.
- Independently review tasks created by other engineers and document clear, evidence-based release decisions.
- Improve authoring standards, validation automation and review practices while mentoring less-experienced engineers.
- Design and maintain the task-development pipeline: automated environment builds, validation harnesses (fail-before/pass-after checks, determinism and flakiness runs) and task packaging for evaluation at scale.
- Build tooling that reduces per-task authoring effort, including repository ingestion, candidate mining from issues and pull requests, containerized execution templates and CI integration for task validation.
- Automate quality gates so that validation evidence is generated, checked and archived without manual steps.
- Improve authoring standards, validation automation and review practices while mentoring less-experienced engineers.
Required qualifications- 4+ years of professional software-engineering experience with strong debugging and code-review skills.
- Proficiency in at least one relevant ecosystem, such as Python, JavaScript/TypeScript, Java or Go.
- Hands-on experience with Git, Linux, Docker, CI/CD, dependency management and automated testing.
- Experience building internal tooling or CI/CD automation that other engineers depend on, such as test harnesses, build pipelines or containerized development environments.
- Ability to understand unfamiliar codebases, multi-service dependencies and complex build or test failures.
- Strong technical writing and sound judgment about whether requirements and tests evaluate behavior fairly.
Preferred experience- Experience with agent-evaluation harnesses and benchmark tooling such as Harbor (agent eval framework), SWE-bench / SWE-gym style task pipelines, Terminal-Bench, or equivalent internal systems.
- Platform or infrastructure engineering: orchestration of containerized workloads, artifact registries, build caching or large-scale batch execution.
- Monorepos, microservices, open-source projects or containerized test infrastructure.
What success looks like- Released tasks are reproducible, deterministic and supported by complete validation evidence.
- Task specifications and tests remain aligned without leaking or over-constraining the implementation.
- Reviews identify defects early, maintain a low rework rate and improve team quality and throughput over time.
- Validation and environment setup are increasingly automated, measurably reducing time-per-released-task and eliminating manual verification steps.
- Tooling you build is adopted by the team and reduces onboarding time for new task authors.
Click on Apply to know more.