In short
A transcript shows too many speakers when voice separation splits one person, often after a change of microphone or connection, and too few when similar voices or crosstalk are merged. Fix it by crediting the split lines to one person and marking the line where each other person starts speaking.
A transcript shows too many speakers when the software that tells voices apart decides one person is two, and too few when it decides two people are one. Both come from the same guesswork: the software groups stretches of speech by how alike they sound, and some voices, microphones and conversations make that hard. The fix is the same in both directions: tell the transcript where each person starts speaking.
This guide explains why transcripts split and merge voices, how to spot which one happened, and how to correct it. The steps use Notey; the causes apply to any notetaker that separates speakers.
In Notey, voices on the other side of a call are told apart after the recording stops, and each gets a label such as Speaker 2. A line that starts a new person's turn is marked from here; the lines after it are credited to that person until the next mark.
Fix split or merged speakers in the transcript
- Open the meeting and the Transcript tab. The Speakers row above the transcript lists your side, the other side and each voice found, such as Speaker 2 and Speaker 3.
- Read the first few lines of each speaker. Decide whether two labels are the same person (a split) or one label covers two people (a merge).
- For a split, find the first line credited to the extra speaker. It is marked from here.
- Click Who is this? on that line and choose the right person. The line, and the lines after it up to the next mark, are credited to them.
- Repeat for every line where the extra speaker starts again. When the extra speaker has no lines left, click their name in the Speakers row and choose Remove. Only a speaker with no lines can be removed.
- For a merge, find the first line the second person said. Click Who is this? on it and choose them from the list, or choose Someone new… and type their name.
- Mark the line where the first person speaks again. Click Who is this? and choose them, so the second person is not credited with everything after their first line.
- Name the speakers. Click a label in the Speakers row, such as Speaker 2, type the name and choose Save. The name replaces the label on every line credited to that person.
A line that follows a mark, rather than carrying one, says who it is credited to and why: "Credited to whoever was marked as speaking before this. If someone else is talking here, say who." That is the hint to look at when a single line is wrong.
Why a transcript splits one person into two
Speaker separation, also called diarization, works by cutting the audio into short stretches, describing how each stretch sounds, and grouping stretches that sound alike. Speaker diarization explained covers the method. It does not know who anyone is; it knows only that some stretches are closer to each other than to the rest. Anything that changes how one person sounds partway through a meeting can push their later speech into a new group.
Common causes:
- They changed microphone or device. Someone who switches from a headset to their laptop's microphone, or drops off and rejoins from a phone, can sound like a different person to the software.
- Their distance to the microphone changed. Leaning back, walking around the room or turning away changes the room sound in their voice.
- Their connection changed the sound. A call app that lowers audio quality on a poor connection, or switches a Bluetooth headset into a lower-quality mode, changes how the voice is encoded.
- One person on two microphones. If a participant's voice also reaches the call through someone else's microphone in the same room, their voice arrives twice, with two different sounds.
- They spoke in two very different ways. Presenting to the group and chatting afterwards can differ enough in pitch and pace.
Why a transcript merges two people into one
The opposite mistake is usually one of these:
- Similar voices. Two people of similar age, accent and pitch, especially on the same kind of microphone, can sit close enough to be grouped together.
- Crosstalk. When people talk over each other, the stretch of audio holds both voices, and it gets credited to one of them.
- Too little speech. Separation needs a few seconds of speech from a person before it treats them as someone separate. Someone who said "Agreed" and "Thanks, bye" may be folded into whoever spoke around them.
- Several people on one microphone. A meeting room on the other side of the call, with four people around one speakerphone, can sound like one or two voices rather than four.
Where Notey does and does not separate voices
Notey records your microphone and the sound your Mac plays as two separate tracks. You and Them come from which track the audio arrived on, so your side and theirs are never confused by voice, only by sound leaking between tracks.
- The other side is separated. After a recording stops, voices on the system-audio track (the people on the call) are told apart. You name them and mark where they start speaking.
- Your side is not. Voices your microphone hears, including people in the room with you, all show as You. The transcript cannot split them, and marking a speaker skips your own lines.
- One voice on the other side stays as Them. In a one-to-one call there is nothing to separate, so the transcript keeps the one label rather than numbering a single person.
When the wrong side gets a line
A different problem looks similar: a line the other person said shows on your side. That is not separation getting it wrong; it is their voice coming out of your speakers and back into your microphone. Move the line to the other side, and use headphones next time. Echo in a meeting recording explains how that happens and how to prevent it.
Prevent it in the next meeting
You cannot make voice separation perfect, but you can give it better audio:
- Ask people to stay on one device. A participant who switches from laptop to phone mid-call is the most common cause of a split.
- Wear headphones. It keeps the other side's voices out of your microphone, which keeps your side and theirs clean.
- Mute when not speaking in a large call. Background noise and crosstalk are what merge voices.
- Say names early. "Over to you, Priya" gives you a line to mark and name later. Voices in a transcript are labels until someone names them; naming the speakers in a transcript covers that step.
Your corrections are per meeting. Naming Speaker 2 "Priya" in one meeting does not name anyone in the next, unless you switch on voice recognition, which is off by default and only ever suggests a name.
If the problem is not the number of speakers but a missing side, for example a transcript with no lines from the other people at all, that is a recording problem rather than a separation one. Start from the troubleshooting guide for microphone permission and the checks that follow it.
Frequently asked questions
Why does my transcript show Speaker 2 and Speaker 3 when only one other person was on the call?
The software that tells voices apart decided the same person sounded like two. It happens when someone changes microphone or device mid-call, moves away from the microphone, or their connection changes the sound of their voice. Credit the second speaker's lines to the first and remove the empty one.
Why are two people shown as one speaker?
Voices that sound alike, people talking over each other, or a speaker who said only a few words can all be merged. Mark the line where the second person starts speaking, and they are credited from there until the next mark.
Why are the people in the room with me not told apart?
In Notey, only the other side of the call (what your Mac plays) is separated into voices. Everyone your own microphone hears is on your side and shows as You, because the microphone track is not separated.
Will correcting speakers be undone when the transcript is improved?
No. Once you have changed anything in a meeting, including naming or marking a speaker, Notey leaves that meeting's transcript as it is rather than replacing it with an improved one.
Does fixing speakers teach Notey the voices for next time?
Only if you switch voice recognition on, which is off by default. With it on, naming someone saves their voice from that meeting, and a later meeting offers the name as a suggestion you accept or reject.