In short
AI meeting notes are made in four steps. The meeting's audio is captured, a speech recogniser turns it into a transcript, a large language model reads the transcript text, and it writes a summary, action items or other notes. Where each step runs decides who holds your audio and text.
AI meeting notes are made in four steps. The meeting's audio is captured; a speech recogniser turns it into a transcript; a large language model reads the transcript text; and the model writes notes such as a summary, action items or minutes. Every AI notetaker does some version of this. They differ in where each step runs, and that decides who ends up holding your audio and your words.
This guide explains how AI meeting notes work at each step, where the work can happen, what each step gets wrong, and what that means for privacy. It is the starting point for this section; the articles on prompts, labels, transcripts and clients go deeper on one part each.
The four steps at a glance
| Step | What happens | Where it can run | What goes wrong |
|---|---|---|---|
| 1. Capture | The meeting's sound is recorded | A bot in the call, a browser extension, or an app on your computer | Missing audio, one side silent, echo |
| 2. Transcription | Speech becomes text | A cloud speech service, or on your computer | Wrong names and numbers, words misheard |
| 3. Reading | A language model reads the transcript | Almost always a cloud model | Things stated that were not said |
| 4. Notes | The model's output is shown and shared | In the app, then wherever you paste it | Unlabelled drafts treated as a record |
Errors travel downwards. A misheard name in step 2 becomes a wrong owner in step 4, and the model cannot know it was misheard.
Step 1: capture
Before anything can be written, the tool has to hear the meeting. There are three common ways.
- A bot joins the call as a participant and records what the meeting platform sends it. Everyone sees it in the list. How meeting bots work covers what that means.
- A browser extension reads the meeting's own captions or audio inside the browser tab. It works only for meetings in the browser.
- An app on your computer records the microphone and what the computer plays. Nothing joins the call, and it works with any app.
Capture decides what the rest of the chain has to work with. A recording where one side dropped out for ten minutes produces a summary of half a meeting, and a summary rarely says which half is missing.
Step 2: transcription
A speech recogniser turns the audio into text. This is where most of the privacy question sits, because it is the step that needs the audio itself.
- Cloud transcription uploads the audio to a speech service run by the vendor or a provider it uses. That service holds the audio for as long as its terms say.
- On-device transcription runs the recogniser on your computer. The audio stays on your disk. On-device transcription explained covers how that works on a Mac.
Either way, recognisers make predictable mistakes: names, product names, numbers, overlapping speech and quiet speakers. A transcript can also say who spoke, either from separate tracks (your microphone and the other side) or by working out voices afterwards, and that can be wrong too.
Step 3: a language model reads the text
The transcript is sent to a large language model with instructions: summarise this, list the action items, draft a follow-up email. The model is almost always a large cloud service, because models good enough to write reliable notes are large.
What the model receives is text, not sound. It does not hear tone, hesitation or who sounded unsure, only what the transcript wrote down. It also cannot tell a misheard word from a real one.
A language model writes by producing text that is likely given what it has read and what it was asked. Most of the time that follows the transcript closely. Sometimes it produces a sentence that sounds right and is not supported by the transcript: a deadline nobody set, an owner for a task nobody took, a decision where there was only a discussion. This is usually called hallucination. It is not a rare bug that a better tool has removed; it is how these models can fail, and the defence is checking. Can you trust an AI meeting summary? goes through the specific patterns and how to check a summary against the transcript.
What the model is asked matters too. An instruction such as "list decisions, and say if none were made" gets different output from "summarise this meeting". Prompts for summarising a meeting transcript has examples if you write your own.
Step 4: the notes
The output is shown in the app, and from there it gets copied into email, documents and tickets. Two things decide whether that goes well.
- Whether it is labelled. A summary that says it was written by AI tells the reader to check it. One pasted without a label reads like something a person vouched for. How to label AI-written notes you share has wording that works.
- Whether someone checked it. The person who was in the meeting is the best reviewer, and soon after the meeting is the best time.
Some tools also answer questions about a meeting. That is the same step 3 with a question instead of an instruction; asking questions about a meeting transcript covers how to get useful answers.
Where each step runs, and why it matters
For privacy, the two questions are where the audio goes and where the text goes.
| Design | Audio goes to | Text goes to |
|---|---|---|
| Bot or cloud recorder | The vendor's servers | The vendor, and its model provider |
| Browser extension reading captions | The meeting platform (for its captions) | The vendor, and its model provider |
| Desktop app, cloud transcription | The vendor's speech service | The vendor, and its model provider |
| Desktop app, on-device transcription | Nowhere; it stays on your disk | The model provider, when you ask for notes |
Every design sends text somewhere if you want AI notes, because the model is not on your computer. What changes is whether the audio travels as well, when the text is sent, and whether it is kept or used to train models. Where does your meeting audio go? is a full checklist for any tool, and does your AI notetaker train on your meetings? takes the training question on its own.
Transcript or summary: which one is the record
The transcript is the closest thing to evidence of what was said: imperfect, but made directly from the audio. The summary is a reading of the transcript. When they disagree, the transcript and the recording decide. That shapes what you keep, what you share and what you quote; meeting transcript vs summary covers it in detail.
Two related questions come up once notes are routine:
- Automatic notes. Some tools write notes as soon as a meeting ends. That is convenient and it means text leaves your computer without a click each time. Getting a summary automatically when a meeting ends covers the trade.
- Notes for clients. Sending AI notes outside your organisation raises the stakes on checking and labelling. Sending AI meeting notes to clients has a checklist.
How Notey does it
Notey is a meeting notetaker for Macs with Apple silicon running macOS 26 or later. Its four steps run like this.
- Capture on your Mac. It records your microphone and what your Mac plays as two separate tracks. Nothing joins the call.
- Transcription on your Mac. Apple's on-device speech recognition transcribes both tracks live. After an English meeting, Notey reads the recording again on your Mac and keeps the words the recognisers agree on. Recording and transcribing need no account and no network.
- Only text goes to the AI, and only when you ask — or when a meeting ends, if you have chosen to have meetings written up automatically. What is sent is the transcript text and the names you gave people, never audio. It goes to Notey's service and on to OpenAI as a processor; it is not used for training and not retained by Notey.
- Labelled notes. A summary with action items (it says so when nobody committed to anything), minutes, tough questions, where you differed and a follow-up email. Each is labelled AI-generated, and your edits are marked. Ask Notey answers questions about one meeting and says when the transcript does not cover the question.
The kinds of note are fixed; there is no custom prompt and no choice of model. AI notes need an account and a paid plan; see pricing. The privacy policy sets out exactly what is sent and when.
Frequently asked questions
Does the AI listen to the meeting audio?
In most notetakers, no. A speech recogniser turns the audio into a transcript first, and the language model reads that text. Some tools send the audio to a cloud service for transcription, which is a separate question from where the summary is written.
Why do AI meeting notes contain things nobody said?
A language model writes text that is likely given what it has read. Usually that matches the transcript; sometimes it produces a plausible sentence the transcript does not support, such as a deadline or an owner. Check anything you will act on against the transcript.
Can AI meeting notes be made without uploading the recording?
Yes, if the transcript is made on your computer. Then only the text needs to go to the language model, and only when you ask for notes. Recording and transcribing can happen with no network at all.
Are AI meeting notes used to train AI models?
It depends on the vendor and on the model provider behind it. Read both privacy policies and look for a plain statement about training, retention and opt-out defaults.
Which is more accurate, the transcript or the summary?
The transcript is closer to what was said, errors included. The summary is a reading of the transcript and can add its own mistakes on top. When they disagree, the transcript and the recording decide.