Transcription and on-device AI

Speaker diarization: how transcripts work out who spoke

Speaker diarization splits a recording by who spoke when. How it works, how it differs from recognising voices, and what two tracks give you for free.

By the Notey team at AInject · · · 6 min read

In short

Speaker diarization is the part of transcription that answers "who spoke when". It splits a recording into stretches of speech and groups the stretches that sound like the same person, so a transcript can say "Speaker 1" and "Speaker 2" instead of one unbroken block.

It does not know anyone's name. Putting names to those speakers is a separate step, done by a person or by comparing voices with ones stored earlier.

This guide explains how diarization works, where it goes wrong, how it differs from recognising voices across meetings, and why recording two tracks settles the most important split without any of it.

Diarization and identification are different things

The two are easy to confuse because both end with names on lines.

DiarizationIdentification
Question it answersWhich parts were said by the same person?Who is this person?
Needs stored voicesNoYes, one per person it can name
OutputAnonymous labels: Speaker 1, Speaker 2Names
ScopeOne recordingAcross recordings

Diarization works inside a single recording. Identification compares a voice with voices it has kept from before, which means keeping something about each person's voice. That stored representation, a voiceprint, is what raises privacy and legal questions, covered in recognising voices across meetings.

How diarization works

Systems differ, but most follow the same broad steps.

  1. Find the speech. Voice activity detection marks which parts of the audio contain speech and which are silence or noise.
  2. Cut it into short pieces. The speech is split into segments short enough that each probably has one speaker.
  3. Describe each piece's voice. A model turns each segment into a set of numbers, an embedding, that captures what the voice sounds like rather than what it says.
  4. Group similar pieces. Segments with similar embeddings are clustered together. Each cluster becomes one speaker.
  5. Line it up with the words. The speaker labels are matched to the transcript's timestamps so each line gets a speaker.

Step 4 is where most of the difficulty sits. The system usually does not know how many people were in the meeting, so it has to decide how many clusters there are as well as who belongs in each.

Where it goes wrong

  • Similar voices. Two people of similar age and accent on the same kind of microphone can be merged into one speaker.
  • One voice, two speakers. The same person can be split in two if their voice changes, for example when they move away from the microphone or start speaking louder.
  • Short turns. "Yes", "right" and "mm" are too short to describe reliably, so they are often attributed to whoever spoke before.
  • Overlap. When two people talk at once, most systems give the stretch to one of them.
  • Shared microphones. A meeting room with one conference microphone puts everyone at different distances from it, with the room's echo mixed in.

The first two show up as a transcript with the wrong number of people in it; too many or too few speakers in a transcript covers fixing that after the meeting.

Diarization is usually measured with diarization error rate, which counts the share of time that is attributed to the wrong speaker, missed, or marked as speech when it was not. Like word error rate for words, it depends heavily on the audio it is measured on.

Why two tracks give you "you and them" for free

In a video call there is one split that matters more than any other: what you said and what the other side said. A notetaker can get it in two ways.

  • From one mixed recording, by diarizing it and working out which cluster is you.
  • From two recordings, one of your microphone and one of what the computer plays. Your side is the microphone track and their side is the other. No voice needs to be analysed.

The second way does not make the mistakes listed above for the you-and-them split, because it is not guessing from sound. It knows which input the audio came in on. Why a meeting recorder should keep you and them on separate tracks goes into the approach in more detail.

It has one weak point. If you use speakers rather than headphones, your microphone also hears the other side, faintly, and some of their words can land on your track. Headphones stop that. If it happens anyway, the fix is to move the affected lines, which is quicker than correcting a diarization error because you know which side they belong on.

Two tracks do not help with a different case: several people on the other side of a call all arrive on the same track. Telling them apart still needs diarization.

How Notey does it

Notey records your microphone and what your Mac plays as two separate tracks and transcribes both on the Mac. On-device transcription explains how the transcription side works.

  • You and Them come from the tracks. Give either side a name for that meeting. If your speakers put someone on your side, move the line.
  • Several people on the other side are told apart after the recording stops, by a pass that runs on the Mac. You name the voices it finds. If it got a change wrong, mark the line where somebody starts speaking and they are credited from there. How to label speakers in a meeting transcript has the steps.
  • Recognising people across meetings is off by default. When you switch it on, naming someone saves their voice from that meeting on your Mac, and in later meetings a similar voice is offered as a suggestion, which you accept or reject. Notey never applies a name by itself, and "Forget this voice" removes someone.

The first two need no stored voices at all. Only the third keeps anything about a person's voice beyond the meeting, and it does nothing until you turn it on.

Frequently asked questions

What is the difference between speaker diarization and speaker identification?

Diarization groups a recording into anonymous speakers, "Speaker 1" and "Speaker 2". Identification says who a speaker is by comparing their voice with voices stored earlier. Diarization needs no stored voices; identification does.

Why does my transcript merge two people into one speaker?

Similar voices, short turns and people talking over each other are hard to separate from sound alone. Correct the lines where the change happens; a tool that lets you mark where someone starts speaking makes this quick.

Does speaker diarization store voiceprints?

It does not have to. Diarization compares voices within one recording and can discard what it computed when it is done. Storing voices to recognise people in later meetings is a separate step with its own privacy questions.

How does Notey know which lines are mine?

Your microphone and what the Mac plays are recorded as two separate tracks. Lines from the microphone track are yours and lines from the other are theirs, so no voice has to be analysed to tell the two sides apart.