Fabric
Website:
fabrichq.ai
Job details:
About the role
Fabric is building voice-native AI agents for hiring, across audio, interviews, resume screening, JD parsing, and sourcing. The quality of every one of those agents comes down to one thing: how well we understand our data and how rigorously we evaluate our systems.
This role owns exactly that. You’ll spend a lot of your time in the data: reading it, understanding it, and turning it into high-quality datasets and eval sets. Just as importantly, you’ll build the tooling and practices that let everyone on the team use data to improve their pipelines and models. This is not primarily a model-training role; the leverage here is in data and evaluation, and in the tooling and standards you bring to the whole team.
What you’ll do
- Spend meaningful time reviewing data (transcripts, model outputs, agent traces) to deeply understand what’s actually happening across our systems.
- Build high-quality datasets and eval sets for each of our agents: audio, interviews, resume screening, JD parsing, and sourcing.
- Design and build tooling that lets everyone on the team access, read, and make use of data to improve their pipelines and models.
- Establish the best practices, workflows, and standards for data and evaluation across the company, and bring in the approaches you already know work well.
- Partner with engineers across the stack; your eval sets and tooling become the backbone of how every agent gets better.
Must-haves
- 3–5 years of experience.
- Proven experience in a similar data/evals-focused role where you’ve spent real time looking at data and building eval sets and datasets, not just training models.
- Strong Python.
- Experience with NLP-related or text-related problems.
- Deep familiarity with the tooling and workflows that make data and evaluation work effective. You know what works best and can build the same practices here.
- Comfort working across a range of ML / LLM problems and moving between different agents and pipelines.
Nice-to-have
- Experience evaluating LLM, agentic, or voice systems.
- Experience with human-in-the-loop review, annotation, or labeling pipelines.
- Comfortable using AI-native tools (e.g. Cursor, Claude, Copilot) as part of your daily workflow.
Who you are
- Rigorous and detail-oriented. You actually enjoy staring at data until it makes sense.
- A systems thinker who wants to build reusable tooling and standards, not one-off scripts.
- Comfortable with the pace and ambiguity of an early-stage startup.
Click on Apply to know more.