PatentVC
Company:
https://www.linkedin.com/company/patentvc
Seniority: Entry level
Industries: Venture Capital and Private Equity Principals
Job details:
Voice Diarization Engineer (Top Pay) — Startup Rebuilding the AI Notetaker Around Rules of Privilege
AI meetings that know what not to record. The technology is being built — by you.
Location: Bengaluru, India (In Person only) Experience: Open to all levels — from exceptional new graduates to senior speech engineers. We hire for raw ability and slope, not years.
What we're building
What we're building An AI notetaker for the enterprise that captures every meeting like the tools you already know, then records only what you've agreed to. A policy you write in plain words decides, in real time as people speak, what gets transcribed and what gets sealed. Privileged conversations, deal terms, customer data, and off-hand admissions never reach the record. It can pull in counsel the moment a matter turns legal, surface context from past meetings, and keep a full audit trail of every decision, all running inside your own environment across Zoom, Meet, Teams, phone, and email.
How it's different Every notetaker today (Otter, Granola, Fireflies, Copilot) makes the same bet: record the whole conversation, then produce a transcript. They only compete on where the audio goes. None of them decide what should be written down in the first place. We do. We're the only one that governs capture in the moment: privilege-aware, running in your own tenant, with every decision logged.
Why it matters For law firms, finance, healthcare, and government, anyone who handles privilege, deals, or regulated data, an ungoverned transcript is a liability. Every word is discoverable, leakable, and a compliance risk, and one stray line can sink a case. These teams want AI notetaking, but they cannot accept indiscriminate capture. That is a large, urgent, underserved market, and we have a patent-pending wedge into it.
Who we're looking for You have lived this problem. You have worked on transcription or speech systems at a company like Google, Otter, or Granola, or you have built a comparable voice-transcription product yourself, and you are genuinely passionate about voice, meetings, and getting transcription right.
The team is already building
Our founding ML engineer was early hire #10 at Otter.ai, where he spent eight years building the company's core AI from 0 to 1 — speaker diarization and recognition, streaming pipelines, and the LLM fine-tuning and evaluation frameworks that took AI transcription from demo to daily habit for millions. He holds multiple issued U.S. patents in real-time transcription, and he leads our engineering team today.
The technology has its builder. What the company needs is additional engineering talent to realize its potential — the person who owns what gets built, who it's for, and how it gets sold.
Track 1 — Speech & ML
The real-time active-inference engine: concurrently with transcription, over a rolling window of speech, classify each moment against natural-language policy — detecting contemplated deals, named projects, contractual and litigation exposure — and choose an action before the words are committed.
Streaming speech-to-text, decoding, VAD, and speaker diarization — learning directly from the engineer who built Otter's diarization from scratch.
Per-speaker consent enforced by voiceprint: speaker models and acoustic signatures that selectively transcribe only consenting participants, in real time.
Model fine-tuning and distillation for line-rate policy classification; evaluation frameworks measuring false-seal and false-capture rates, diarization error, and latency budgets. Governance accuracy is the product.
Track 2 — Real-Time Audio & Systems
The real-time audio spine: low-latency streaming across Zoom, Meet, Teams, phone/SIP, and desktop audio — WebRTC, WebSockets, meeting-platform bots and SDKs, jitter and latency engineering.
A provably-ephemeral memory architecture: audio lives in a volatile store only transiently; nothing reaches durable storage unless policy permits; when a moment is dropped, no copy is retained — provable absence, not a sealed record. If "we never wrote it down" can't be proven, the product doesn't work. This is the deepest systems problem here and the credibility anchor for regulated buyers.
Encryption, key/secret lifecycle, retention and deletion guarantees, immutable audit logging, and deployment inside customer environments.
Track 3 — Product & Agent Engineering
The natural-language policy layer: an admin interface where compliance teams write governance rules in plain English, compiled into something the inference engine evaluates against live speech inside a tight latency budget, applied org-wide.
The action set the inference drives: transcribe and commit; drop; redact a single value; withhold as privileged; alert a reviewer; invite counsel from a directory; or surface a relevant document, regulation, or prior transcript via retrieval — sometimes through a rate-limited in-meeting agent that helps without taking over the room.
Live meeting UX: real-time transcript views, sealed-moment indicators, and audit-trail and compliance views for CFOs, GCs, and security teams. TypeScript/React front end, Python/Node services, real-time state throughout.
Who we're looking for
Ideally, you have lived this problem: you've worked on transcription or speech systems at a company like Google, Otter, or Granola — or Deepgram, AssemblyAI, Fireflies, Rev, or a Whisper-based product — or you've built a comparable voice product yourself, and you're genuinely passionate about voice, meetings, and getting transcription right. You want to own a product end to end, not be one engineer of fifty.
But no prior industry experience is required. We would rather hire the best new graduate in India — and pay top of market for them — than an average engineer with five years of experience. Signals we care about:
Exceptional CS fundamentals. Data structures, algorithms, systems. Top of your class at a top program (IIT, BITS, NIT, IIIT, or equivalent — or you outperformed that bar from anywhere).
Evidence you build. A GitHub with real projects, open-source contributions, something you shipped, research you published, or competitive results (ICPC, Codeforces, Kaggle). This is our strongest signal.
Depth in at least one relevant area: ML/deep learning (PyTorch, Transformers), speech/audio, distributed systems, real-time networking, or full-stack product engineering.
Python and one of C++/Java/TypeScript. Clean, fast code; you profile before you guess.
Curiosity about speech and language AI. You've trained a model, read the Whisper paper, or built something with an LLM because you couldn't help yourself.
Founder energy. You move fast, own outcomes, and want to be an early engineer on something that matters.
In-person, in Bengaluru. This team works together in one room. That's how 0-to-1 gets built.
Bonus points: on-device/edge inference, model fine-tuning, meeting-platform bots and SDKs, LLM systems, privacy- and security-sensitive systems, or prior early/founding engineer experience.
How we work — and what AI-native means here
AI-native: you ship faster with AI than most teams do without it — because you can tell good output from plausible-looking output, and you don't merge what you can't stand behind. You own the judgment; AI is your leverage. The bar is your taste, not the model's.
This is governance software for privileged and regulated data, so speed and care aren't in tension — you're fast where it's cheap to be wrong and careful where it isn't. You ship real products, and you own every line.
What you'll have shipped in your first year
~6 months: production code in a live pilot — policy-driven drop/redact working end-to-end, with your components in the critical path.
~12 months: consent-gated, multi-channel capture in production and the ephemerality guarantee demonstrable under audit — and you've grown from apprentice to owner of a core subsystem.
What you'll get
Top-of-market pay. We benchmark against the best-paying Bengaluru AI startups and pay at or above it. For a standout, we go beyond the band. Cash compensation leads; this is not an equity-for-salary trade.
Potential for options as an early employee after 3 months, with standard vesting — meaningful upside as the company grows.
An apprenticeship you cannot buy: daily, in-person work with an Otter.ai founding-era engineer — diarization, streaming ASR, LLM systems, and evals, learned at the source.
Silicon Valley backing: PatentVC funding, patent-pending IP drafted by the attorneys behind Trademarkia, and direct exposure to Sand Hill Road fundraising as it happens.
Scope that compounds. The systems you build in year one are the company's core IP. No legacy code. A short path from apprentice to owner of a core subsystem.
How to apply
Apply through LinkedIn with your relevant experience — and include links to what you've built (GitHub, shipped products, papers, competition profiles). What you've built is our strongest signal.
Click on Apply to know more.