Website:
lh2.ai
Job details:
Research Engineer — Coding Evaluation & Training Data
About LH2 AI Labs
Built by second-time founders who have built and sold companies before, LH2 AI Labs is building the post-training infrastructure for frontier AI models.
We bring private, high-quality institutional datasets and vetted domain experts into frontier AI pipelines across verticals such as coding, computer use, agentic workflows, medical, audio, and more.
For AI to keep progressing, it needs high-quality training data drawn from real production use cases. The public web has already been crawled and trained on — there is limited new signal left there. That is where we come in.
Our vision is to create a world where frontier models can access high-quality data on tap, the same way they access compute today.
The Role
We are looking for a research engineer — coding evaluation & training data — to help us build and run coding data projects for AI systems.
In this role, you will work at the intersection of software engineering, AI evaluation, and data quality. You will design realistic coding tasks, set up repo-based evaluation environments, define rubrics and validation workflows, analyse task quality, and help build the pipelines that turn raw engineering work into reliable training and evaluation datasets.
This is a strong fit for a software engineer who enjoys working with real codebases, debugging environments, understanding how engineers actually solve problems, and improving the quality of data used to train and evaluate AI coding agents.
What You’ll Do
•Own coding data projects end-to-end, from scoping and pilot design to execution, quality review, iteration, and customer delivery.
•Design coding tasks that reflect real software engineering work, such as debugging, refactoring,
feature implementation, code review, large repo navigation, and test-based problem-solving.
•Source and analyse tasks from repositories, pull requests, issues, test suites, diffs, execution logs, and customer-provided codebases.
•Set up technical environments for coding tasks, including repositories, containers, dependencies, test harnesses, sandboxes, and code execution workflows.
•Define task instructions, grading criteria, rubrics, golden answers, validation checks, and pass/
fail expectations.
•Evaluate coding task outputs from humans, models, or coding agents and decide whether they
meet the required quality bar.
•Identify broken, ambiguous, low-signal, or non-reproducible tasks before they reach customers.
•Build and improve data pipelines for collecting, structuring, validating, analyzing, and packaging
coding evaluation data.
•Partner with customer teams to translate high-level training or evaluation goals into concrete
task designs and technical workflows.
•Work with internal operations, engineering, and review teams to improve task quality, throughput, tooling, and delivery processes.
•Analyse failure modes across tasks, models, workers, and environments, and recommend improvements to the dataset or workflow.
What We’re Looking For
• Graduate from a Tier 1 engineering institution such as IIT, BITS, NIT, IIIT, or equivalent. •
3–6+ years of professional software engineering experience building or maintaining real systems.
•Strong coding ability in at least one mainstream language such as Python, JavaScript/TypeScript, Java, Go, C++, or Rust.
•Comfort reading, understanding, and debugging unfamiliar production codebases.
•Strong understanding of Git, GitHub workflows, pull requests, issues, commits, diffs, branches, and code review.
•Experience working with tests, test runners, CI workflows, failing builds, dependency issues, and
debugging logs.
•Ability to set up and debug technical environments using Docker, shell scripts, package managers, and local or cloud-based execution setups.
•Strong engineering judgement: you care about correctness, reproducibility, code quality, edge cases, and whether a task reflects real-world software engineering.
• Ability to convert fuzzy customer requirements into structured workflows, task specs, quality checks, and deliverables.
• Strong written and verbal communication skills, especially for documenting task design, quality issues, failure analysis, and customer-facing outputs.
• Interest in AI coding agents, LLM evaluation, training data quality, post-training workflows, and benchmark design.
Nice to Have
• Experience with SWE-bench, coding benchmarks, AI evaluations, coding agents, or repo-based
task environments.
• Experience building training or evaluation datasets for LLMs, coding models, RLHF, RLVR, or
post-training workflows.
• Familiarity with agents or developer tools such as Cursor, Claude Code, OpenAI Codex-style
agents, SWE-agent, Devin-style systems, or similar tools.
• Experience designing rubrics, golden sets, reward signals, or verifier-based evaluation
workflows.
• Experience with sandboxing, automated grading, unit-test generation, or secure code execution.
• Experience building internal tools, dashboards, pipelines, or QA systems for large-scale data
operations.
• Prior experience working with AI labs, data vendors, eval teams, or applied AI teams.
Success in This Role Looks Like
• LH2 AI Labs can reliably turn real-world coding work into high-quality training and evaluation
datasets.
• Coding tasks are realistic, reproducible, well-scoped, and verifiable.
• Customer requirements are translated into clear task workflows, technical environments, and quality standards.
• Broken or low-quality tasks are identified early through strong review and validation processes.
•Internal pipelines improve the speed, quality, and consistency of coding data delivery.
•Customers receive clean datasets, clear analysis, and trustworthy evaluation outputs.
•The team develops strong reusable standards for coding task design, evaluation, review, and delivery.
Why This Role Matters
Training and evaluating AI coding systems requires more than collecting code. The best datasets need realistic engineering tasks, reliable environments, strong rubrics, reproducible tests, and careful human judgement.
This role will help LH2 AI Labs build the technical workflows and quality standards needed to deliver high-value coding data for customers building the next generation of AI systems.
Click on Apply to know more.