Website:
patentvc.com
Job details:
Founding Engineer — Voice & Transcription AI PatentVC · India (Remote) · Full-time
Built by a founder who’s spent the last two years shipping voice AI with real customers — now tackling its hardest problem: deciding what a meeting should never record.
The product
We’re building a real-time meeting system whose governance engine decides, as people speak, what is ever recorded at all — for enterprises handling M&A and deal-making, regulated data, and privileged matters, where the wrong sentence sitting in a searchable transcript is a liability, not a convenience. Think Otter or Granola, but the interesting half isn’t the transcript — it’s the engine on top that classifies live speech against plain-English policy and decides, per moment, what happens before anything is written down.
Transcription is the substrate. Governance is the product.
The role
You are employee #1 under the founder, on a real path to Head of Engineering. You build the product with your own hands now — you own the architecture, the models, and the shipping. As pilots convert, you define how we build, then hire and lead the engineers behind you and own the technical org. We’re hiring someone who wants that arc, not just a senior IC seat.
You’ll work directly with the founder and with our early pilot customers — sitting in with the GCs, deal teams, and compliance leads we’re piloting with, hearing what they actually need, and turning messy real-world requirements into clean architecture. You can explain a hard tradeoff to a GC or a CISO without dumbing it down.
The hard problems you’ll own
- A real-time active-inference engine. Concurrently with transcription, over a rolling window of speech (a span of time or N complete sentences, re-evaluated every interval), classify each moment against natural-language policy and detect sensitive subject matter — contemplated deals, named projects, contractual and litigation exposure — and choose an action before the words are committed.
- A natural-language policy layer. An admin interface where compliance teams write governance rules in plain English, compiled into something the inference engine evaluates against live speech inside a tight latency budget, applied org-wide across every meeting.
- The action set the inference drives: transcribe and commit; drop — don’t record, no copy kept (the default for sensitive speech); redact a single value; withhold as privileged (counsel present → kept out of the produced transcript); alert a reviewer in real time or after the meeting; route to a sanctioned service; invite a further party (identify counsel from a directory and send a join request); or surface a relevant document, regulation, or prior transcript via retrieval — sometimes through a rate-limited in-meeting agent that helps without taking over the room.
- A provably-ephemeral memory architecture. Audio lives in a volatile store only transiently; nothing reaches durable storage unless policy permits; and when a moment is dropped, no copy is retained and no audit entry is written — provable absence, not a sealed record. If “we never wrote it down” can’t be proven, the product doesn’t work. This is the deepest systems problem here and the credibility anchor for selling to regulated buyers.
- Per-speaker consent enforced by voiceprint. Use speaker models / acoustic signatures to identify who’s talking and selectively transcribe only consenting participants while declining the rest, in real time.
- The real-time audio spine: low-latency streaming transcription, speaker diarization, and high-accuracy captioning across accents and noisy audio, captured across Zoom, Google Meet, Microsoft Teams, phone/SIP, and desktop audio.
How we work — and what AI-native means here
AI-native: you ship faster with AI than most teams do without it — because you can tell good output from plausible-looking output, and you don’t merge what you can’t stand behind. You own the judgment; AI is your leverage. The bar is your taste, not the model’s.
This is governance software for privileged and regulated data, so speed and care aren’t in tension — you’re fast where it’s cheap to be wrong and careful where it isn’t. You ship real products, and you own every line.
What we need from you
- Deep real-time audio. You’ve shipped production low-latency / streaming audio — ideally STT and diarization with one of Whisper, Deepgram, AssemblyAI, Google/Azure Speech, or self-hosted/fine-tuned models over WebRTC/WebSockets. This is our strongest signal. (Adjacent path: if your deepest shipping is real-time voice agents, live media, or streaming infra and you can clearly own ASR from here, we want to talk — we’re hiring an architect, not a checklist.)
- An architect who designs for scale. You think through complex problems and design systems that are expensive to get wrong — not just code that works today.
- Real-time inference under a latency budget. You can build the engine that classifies live speech against policy every interval and chooses an action before words are committed — the system at the center of this product.
- LLM-driven systems in production. Real-time classification, structured outputs, latency/cost control, evals. You’ve built the tooling around models, not just called them.
- Strong full-stack engineering (Python and/or Node/TypeScript) to build the product around the model. Today we lean Python/TS; you set the stack.
- Security-sensitive instinct. You can design a system where “we never kept a copy” is provably true: encryption, key/secret lifecycle, retention/deletion guarantees, residency, SSO, and auditability are things you reason about, not afterthoughts.
- Clear communication. You can run a technical conversation with a non-technical pilot customer and leave with the right requirements.
Strong pluses
- Model fine-tuning or domain-adapted ASR; diarization and speaker-ID at scale.
- Agent and LLM tooling — the eval harnesses, tool/function layer, retrieval/RAG, and guardrails that make a real-time agent trustworthy, including retrieval that can surface a statute, a prior admission, or an internal doc mid-meeting.
- Meeting-platform bots/SDKs (Recall.ai-style, Zoom RTMS, Teams media bots, Meet Media API).
- You’ve built for regulated or legal/compliance/M&A buyers and understand why one wrong sentence in a searchable transcript is a liability, not a feature.
- Customer-facing experience translating messy enterprise requirements into architecture.
- You’ve been an early or founding engineer before.
Forward-looking, not day-one: over time we expect to push inference toward the network boundary — routing audio to sanctioned vs. unsanctioned services at the edge. There’s no hardware product today.
What you’ll have shipped in your first 6–12 months
- ~6 months: our pipeline extended from prototype to a live first pilot, with policy-driven drop/redact working end-to-end.
- ~12 months: consent-gated, multi-channel capture in production, the ephemerality guarantee demonstrable under audit — and, as pilots convert, the first engineers hiring in behind you.
Comp & ownership
- Competitive upper half for the staff-level / Head-of-Eng track we’re hiring toward; for a standout we go beyond the band.
- Meaningful founding equity — real ownership as employee #1, with standard vesting. We’ll walk through the specifics together early in the process.
- Remote-first across India, working directly with the founder.
- A true ground-floor seat: @Dhruv Channa — founding engineer #2 at CollectWise (YC F24) and founding FDE at Cekura (YC F24), with deep voice-AI domain expertise (both voice-AI companies — debt collection and voice-agent testing, respectively) — is building this inside PatentVC. The wedge is a patent-pending approach to deciding what gets recorded before it’s written down.
To apply
Send your CV and links to what you’ve built — voice/transcription products, real-time audio or agent systems, GitHub, demos, shipped apps. We care more about evidence than cover letters.
Click on Apply to know more.