Transcription and on-device AI

Naming the speakers in a meeting transcript

Why transcripts say Speaker 1 and Speaker 2, how to label speakers in a transcript with real names, and how to fix lines credited to the wrong person.

By the Notey team at AInject · · 6 min read

In short

Transcripts say "Speaker 1" and "Speaker 2" because software can tell voices apart but not who they belong to. Listen to each speaker's first lines, name that speaker, and check where the voices change. Name speakers before asking an AI for notes, so tasks go to the right people.

Transcripts say "Speaker 1" and "Speaker 2" because the software can tell voices apart without knowing whose they are. To label speakers in a transcript, listen to the first few lines from each unnamed speaker, give that speaker a name, and then check the places where the voices change hands. Do it before you ask an AI for a summary or action items, so the notes credit the right people.

This guide explains where the labels come from, the steps to name them, and how to fix lines credited to the wrong person.

Name the speakers in a transcript

  1. Find each unnamed speaker's first lines. Scroll to where "Speaker 1", "Speaker 2" and so on first appear.
  2. Listen to them. Play the recording from those lines. Most people recognise a colleague from a sentence.
  3. Give each speaker a name. In most tools, click the label and type. The name applies to every line credited to that speaker in this meeting.
  4. Check the changes of speaker. Read the transcript around places where two people traded short turns. That is where grouping goes wrong.
  5. Fix wrong lines. Move a line to the right person, or mark where someone actually started speaking.
  6. Then generate notes. Summaries and action items use the names you gave.

Where "Speaker 2" comes from

Working out who spoke is a separate job from working out what was said. There are two ways a transcript can do it:

Neither method knows anyone's name. Names come from a person, or from a voice saved from an earlier meeting and matched to a new one.

When the grouping is wrong

Voice separation makes predictable mistakes:

ProblemWhyFix
One person split into two speakersTheir voice changed: a cough, laughter, moving away from the microphoneGive both labels the same name
Two people merged into oneSimilar voices, or both on one microphoneMark where the second person starts
Short interjections given to the wrong person"Yes" and "right" are too short to identifyUsually harmless; fix the ones that matter
CrosstalkTwo people speaking at onceListen and decide who said the words that matter

A voice needs some speech before it can be told apart. Someone who says only one short sentence may not get a separate label at all. Transcript shows too many (or too few) speakers goes into these cases.

Saving voices for next time

Some tools remember a voice once you name it, and suggest that name in later meetings. That saves naming the same people every week. It also means storing a voiceprint, a mathematical representation of someone's voice, which may be biometric data under laws such as the GDPR and Illinois's BIPA, for people who never installed the tool. Recognising voices across meetings: the privacy questions covers what to consider before turning it on.

Why naming comes before notes

A language model writing notes from a transcript sees only text. If a line reads "Speaker 2: I'll send the contract on Friday", the model either writes "Speaker 2 will send the contract" or guesses who Speaker 2 is from context. Either way, the action item is less useful or wrong. Naming speakers first is the cheapest way to make AI notes more accurate.

Names in notes you share

Once speakers have names, those names travel: into summaries, action items, exports and anything you paste elsewhere. Two habits help:

  • Use the names people use. "Priya", not a nickname from the calendar invite, so readers know who is meant.
  • Check the names before you share. A wrong name on a line is a small error in a transcript and a bigger one in an action item sent to a client.

Naming speakers in Notey

Notey records your microphone and what your Mac plays as two tracks, and transcribes both on your Mac. How speakers are named follows from that.

  • You and Them come from the tracks. Your microphone is You; what the Mac plays is Them. Give either side a name for that meeting.
  • Several people on the other side are told apart after the recording stops, by a pass that runs on your Mac. They appear as numbered speakers; name each one.
  • Mark where somebody starts speaking. If a change of speaker was missed, mark the line where the other person starts, and they are credited from there.
  • A voice needs a few seconds of speech to count as a separate person.
  • People in the room with you are not separated. Voices on your microphone track stay on your side, as one speaker. If your speakers put someone from the call on your side, move the line to theirs.
  • Recognising people across meetings is off by default. When you turn it on, naming someone saves their voice from that meeting on your Mac, and a similar voice in a later meeting is offered as a suggestion, labelled as one, which you accept or reject. Notey never applies a name by itself.
  • Names stay on your Mac, and are sent with the transcript text only when you ask for AI notes, so the notes can use them.

One trade-off to know: naming or marking a speaker counts as changing the meeting. After an English meeting, Notey reads the recording again on your Mac and keeps the words the recognisers agree on, but it leaves a meeting you have changed as it was recorded. Nothing on a meeting says whether it has been improved yet, so if the better transcript matters more to you than naming people straight away, leave the naming for later. On-device transcription explained covers the transcription itself.

Frequently asked questions

Why does my transcript say Speaker 1 and Speaker 2?

The software grouped the speech by voice and does not know whose voices they are. Telling voices apart is one step; knowing who they belong to needs you, or a voice saved earlier.

How do I rename a speaker in a transcript?

Most tools let you click the speaker's label and type a name, which changes every line credited to that speaker. Check a few lines where the voice changes to make sure the grouping was right.

What if one person's lines are split between two speakers?

Voice separation sometimes splits one person into two, or merges two into one. Give both labels the same name, or mark where the real change happens. The guide on too many speakers covers this.

Can a transcript name speakers automatically?

Only if the tool has a stored voice for each person, which is voiceprint recognition and raises its own privacy questions. Without that, the names come from you.

Why should I name speakers before generating AI notes?

An AI reading "Speaker 2 will send the deck" cannot know who that is, and may guess from context. Named speakers mean action items go to the right person.

Can I name people who were in the room with me?

In Notey, voices on your microphone track are not separated, so people in the room with you share your side. Note who said what in your own notes if it matters.