Website:
zstate.ai
Job details:
← All roles
Research Engineer
Founding Team
- Full-time
- Gurugram, India Apply for this role → Or write to careers@zstate.ai
About Zstate
As AI models become increasingly capable, the industry needs better ways to measure reasoning, reliability, and real-world performance.
Zstate AI is building the evaluation and post-training infrastructure powering the next generation of AI systems. We create expert-driven benchmarks, reward signals, RL environments, and high-quality datasets that help frontier AI models become more capable, reliable, and trustworthy across diverse real-world domains.
Zstate AI is led by a founding team from IITs, IIMs, and DCE, bringing together deep technical and strategic expertise to build category-defining AI infrastructure. We are backed by Giraffe Studios, a venture studio focused on deep-tech startups founded by Himanshu Aggarwal (Co-founder & CEO, Aspiring Minds) and Mohit Tandon (Co-founder, Delhivery).
As one of our earliest hires, you'll work directly with the founders to solve challenging engineering problems and help shape both the product and the company from day one.
What this role looks like
You are one of the first research engineers at Zstate AI. AI labs come to us because they need to know where their models break in the real world. You will work with domain experts to design a benchmark, shape the tasks, write the scoring rubrics, and run models against it. Then you will find the failure modes, refine the benchmark, and turn the whole thing into a signal AI labs can trust and compare over time.
A day in the life
- Design benchmark suites across diverse real-world domains, then build the evaluation framework and scoring methodology that makes the results defensible.
- Build pipelines that run models at scale, score outputs, and generate performance breakdowns a lab can act on.
- Read model outputs carefully to find where and how they fail, and use those failures to improve both the benchmark and the reward design.
- Develop expert grading rubrics, LLM-as-a-Judge workflows, and evaluation methodologies for post-training.
- Build internal tools and workflows that make benchmark creation, expert authoring, and evaluation faster and cleaner.
- Collaborate with domain experts to translate messy real-world workflows into production-ready evaluation datasets.
Who you are
- 3-8 years of experience in software or ML engineering, with a real track record of shipping production systems.
- Experience in AI/ML engineering, model evaluation, benchmarking, or post-training, backed by strong software engineering fundamentals.
- Strong programming skills and experience building ML systems, backend applications, or scalable data pipelines.
- Solid understanding of APIs, databases, cloud infrastructure, and modern engineering best practices.
- Analytical thinking: you can read model behaviour and design robust evaluation methodologies from what you see.
- Excellent problem-solving ability, high ownership, and the curiosity to tackle open-ended technical challenges.
- The ability to thrive in a fast-moving startup environment with significant ownership and responsibility.
Extra credit
- Experience building LLM evaluation datasets, benchmarks, RL environments, or post-training pipelines.
- Experience with reinforcement learning, synthetic data generation, or reward modelling.
- Research publications, open-source contributions, or public work related to AI evaluation or machine learning.
- Experience working with frontier language models or AI systems in production environments.
Why this matters
- Competitive compensation for the right candidate.
- Build the evaluation infrastructure that shapes how frontier AI models are tested and improved.
- Work directly with the founders and help shape both the product and the company from day one.
- Solve challenging engineering problems that influence how state-of-the-art AI models are evaluated.
- Take ownership of meaningful technical work with rapid learning, visible impact, and growth opportunities as we scale.
If structured thinking, messy problems, and building the evals that hold AI systems to a higher bar sounds interesting, send your resume and a short note on why this role interests you to vijay@zstate.ai or careers@zstate.ai.
Click on Apply to know more.