Uplevyl
Website:
uplevyl.com
Job details:
About Uplevyl
Uplevyl builds AI-powered knowledge and community infrastructure for organizations serving women. Our products include UpGenie (our domain-specific AI assistant), WeHub (our community platform), and UpSocial (a social platform for women). We work with mission-driven partners to turn complex, high-stakes information into clear, trustworthy guidance, and to build systems that create lasting impact.
We hold a simple conviction: in high-stakes domains like rights, law, and financial security, a generic AI is not enough. The answers people stake their lives and livelihoods on need a purpose-built system with verified, native data. That is what we build, and it is why the work is urgent. AI is reshaping how the world learns, works, and earns, and the women we serve cannot afford to be left off that train.
We have also made a deliberate choice about how we build: a small team of exceptional people, paid well above market, each doing work that would normally take several. We would rather be ten people who move the world than thirty who move paper. That choice sets the bar for every hire, including this one.
About the role
This internship exists to answer one question, over and over: when UpGenie tells someone what US law says about their rights at work, is it actually right? You will study the underlying US statutes and case law, and check whether our AI's answers match them, jurisdiction by jurisdiction.
You will sit between our legal and engineering teams. You will test and evaluate the model's answers against real US employment law, verify the legal sources behind each answer, and identify exactly where and why it gets things wrong. When you find a gap or an error, you will work directly with the AI team to help close it.
This is US law only, federal and state, and it is not a litigation or advisory practice: you will not be giving legal advice or appearing for clients. What you need is a practitioner's eye for US statutes, and the discipline to prove, with evidence, whether an answer is right or wrong.
What you'll own
- Legal accuracy testing. Test UpGenie's answers against real US employment and workplace-rights statutes and case law, jurisdiction by jurisdiction, and flag every place it gets the law wrong.
- Source verification. Check the legal sources behind our AI's answers: are they current, correctly scoped, and actually saying what we cite them for.
- Error identification and triage. When an answer is wrong, work out why, a bad source, a stale statute, a jurisdiction mismatch, and document it clearly enough for engineering to act on.
- The improvement loop with engineering. Sit alongside our AI engineers, evaluate model or retrieval changes before and after they ship, and tell them, with evidence, whether the answer got better or worse.
- Benchmark and dataset support. Help build and maintain the test sets and scoring criteria we use to measure legal accuracy over time.
- Plain-language documentation. Write up findings so both lawyers and engineers can act on them without a translation layer.
Who we're looking for
This is an internship, not a formality. We are looking for someone who can already read a US statute like a practitioner and is comfortable interrogating an AI system's output like an engineer would.
- Legal training in progress or complete. Pursuing or holding a JD or equivalent, with real exposure to US employment, labor, or workplace-rights law, not just coursework.
- A practitioner's eye for US statutes. You can read a US law closely enough to tell when a source is wrong, stale, or wrongly scoped for the jurisdiction.
- Evaluation instinct. Given a model's answer, you can say precisely what is wrong with it, why, and what would have caught it earlier.
- Comfort with ambiguity in the law. Much of this domain has no clean answer. You can make a defensible call, write down your reasoning, and revisit it when the evidence changes.
- Some comfort with AI or data tools. You don't need to be an engineer, but you should be comfortable working inside spreadsheets, structured datasets, or a scoring rubric, and picking up new tools fast.
- Sharp writing. You can explain a jurisdictional subtlety to an engineer and a technical finding to a non-technical teammate, each in a few clear lines.
- High agency. You unblock yourself. Nobody has done this exact job before you, so there is no playbook waiting; you help write it.
Even better if
- You have experience with US employment law specifically, or multi-jurisdictional legal research at scale.
- You have worked on legal AI, legal tech, or an AI product in another regulated domain.
- You have built or used an evaluation harness, scoring rubric, or structured dataset, even a simple one
- You have published research, case notes, or writing on AI, law, or legal evaluation.
- You have worked in mission-driven, social-impact, or women-focused ecosystems.
- You have founder, startup operator, or early-employee experience.
Before you apply
We want to be direct about what this is. You will spend as much time in datasets, rubrics and evaluation runs as you will in statutes, and if that combination does not appeal to you, this is not the right fit. It is also a role of one, on a small senior team, at a company where the accuracy of what we tell people matters more than almost anything else we do. The work is demanding, the expectations are real, and the rewards, in compensation, in ownership, and in what you will learn, match them. If you want to define how a legal AI system proves that it can be trusted, we want to hear from you. Our interview process is thorough, and we will walk you through every step of it in our first conversation.
How we work at Uplevyl
- We finish what we start, and we do it well. We follow through on what we say we'll do, and we care about the difference our work makes.
- We listen before we decide. Before we act, we ask who it actually affects: a customer, a partner, a teammate, or the communities we serve. Trust here is built the plain way, by consistently showing up for people.
- We move with ownership and urgency. We don't wait for perfect information or for someone else to raise their hand. If something looks like it could go wrong, we say so early. We make the call, we move fast, and we hold a high bar for quality.
- We are better together than alone. We work across teams, say what we actually think, and go out of our way to help each other succeed. Leadership here is measured by the impact you create.
- We stay curious and keep learning, including how to use AI well. We hold ourselves to the bar we're building toward: questioning our own assumptions and using AI to think better and move faster.
Click on Apply to know more.