Transcription and on-device AI

Can you trust an AI meeting summary?

Where AI meeting summaries go wrong — invented commitments, missing dissent, wrong names — and a quick way to check one against the transcript before you send.

By the Notey team at AInject · · · 7 min read

In short

An AI meeting summary is a draft written by a model that read the transcript. It is usually right about what the meeting covered and often wrong in the details people act on: who agreed to do what, by when, and what was left undecided.

You can trust one after you have checked those details against the transcript, which takes a few minutes, and not before.

This guide covers where summaries go wrong, why, how to check one, and why "the transcript does not say" is a good answer from an AI rather than a failure.

Two layers of error

A summary can be wrong for two separate reasons, and it helps to know which you are looking at.

The transcript was wrong. Speech recognition mishears names, numbers and jargon, and it can put words in the wrong person's mouth if audio from both sides of a call reaches one microphone. A model summarising a wrong transcript will faithfully summarise the mistake. Why meeting transcripts get names and numbers wrong covers this layer.

The summary was wrong. The transcript was fine and the model misread it, left something out, or added something that was never said. This is the layer this guide is about.

Where summaries go wrong

Commitments nobody made

The most damaging error. A meeting ends with "we should look into the pricing" and the summary says "Tom to review pricing by Friday". Nobody said Tom, nobody said Friday. A language model is built to produce plausible text, and a summary with an owner and a date looks more like a summary than one without.

If a follow-up email then goes out with that line in it, Tom has been assigned work in writing that he never agreed to.

Missing dissent

Summaries compress, and disagreement is often the first thing lost. A meeting where one person objected three times and was overruled can come out as "the team agreed to proceed". For a decision that may be revisited, who disagreed and why is exactly what the record needs.

Hedges turned into facts

"I think we can probably get it done by March, but don't hold me to it" becomes "delivery in March". Conditions ("if legal signs off"), hypotheticals ("say we did drop the free tier") and jokes are all at risk of being written up as plans.

Wrong people

On a call with several people on the other side, the model only knows who said what if the transcript does. If speakers are labelled "Speaker 2" or wrongly named, the summary will attribute decisions accordingly.

Wrong numbers

Figures, dates and amounts are hard for speech recognition and easy for a model to round, combine or move. "Fifteen" and "fifty" is a classic transcription error; a model then carrying "50" into a budget line is a summary error on top.

Things that were not in this meeting

A model asked for "context" may bring in general knowledge. A summary that explains what a term means, or adds background nobody discussed, is mixing the meeting with the model's own text.

How to check a summary

Checking every sentence is rarely worth it. Check the parts people will act on.

  1. Action items. For each one, find the moment in the transcript where someone agreed to it. Confirm the owner said yes, and that any date was actually said. If you cannot find it, delete it or mark it "to confirm".
  2. Decisions. Find where each was made. Check nobody objected afterwards, and that it was a decision and not a proposal.
  3. Numbers and dates. Check each one against the transcript, and play the audio where the transcript itself looks doubtful.
  4. Names. Check who is credited with each commitment and statement.
  5. What is missing. Skim the transcript for objections, open questions and anything someone asked to be written down.

A summary that points to where each item came from in the recording makes this much faster. If yours does not, search the transcript for a keyword from the item.

Writing meeting action items that get done explains what a good action item needs, which is also a checklist for spotting a bad one. Before a summary becomes a follow-up email, it should have passed step 1.

"The transcript does not say" is a good answer

A meeting without a clear outcome should produce a summary without one. A tool that writes "no action items were agreed" or "the transcript does not say who will do this" is reporting the meeting accurately. It is more useful than a confident invention, because it tells you something you need to know: the meeting ended without an owner, and someone should fix that.

The same goes for questions. If you ask an AI about a meeting and the answer was never discussed, the right reply is that the transcript does not cover it, not a plausible guess.

When you compare tools, test for this. Record a meeting that deliberately ends without decisions and see what each summary says.

Why labelling matters

A note labelled as AI-written is read differently from one a colleague wrote after the meeting, and it should be. The label tells the reader the note has not necessarily been checked by anyone who was there. Once you have checked and edited it, it is useful for the record to show that too.

There is a legal side in some places. The EU AI Act includes transparency obligations for certain AI-generated content; labelling AI-written meeting notes explains what they cover. Labelling is good practice whether or not a law requires it in your case.

How Notey handles this

  • Summaries say when nobody committed to anything. Notey's Summarise produces a title, a summary and action items; if nobody committed to anything, it says so rather than inventing items.
  • Ask Notey says when it doesn't know. Answers to questions about a meeting are drawn from that meeting's transcript, and if the transcript does not cover the question, the answer says so.
  • Minutes point back to the recording. Each topic in the minutes is marked with where it began in the recording, and every line of the transcript plays from its timestamp, so you can check a point by listening to it.
  • Everything the AI wrote is labelled, wherever it appears, including in Markdown exports. When you edit it, it is marked as edited by you.
  • Your side and theirs come from separate tracks, so the transcript knows which side said a line without guessing from voices. On speakers, the microphone can pick up the other side; you can move a line to the right side before asking for notes.
  • Notes are written only when you ask, or when a meeting ends if you have turned automatic write-ups on. Only the transcript text is sent to produce them, never the audio.

None of this makes a summary correct. It makes checking one quicker. On-device transcription explains how the transcript underneath is produced.

Frequently asked questions

Are AI meeting summaries accurate?

Often good enough to start from, and sometimes wrong in ways that matter. They are most reliable on what was discussed and least reliable on who agreed to what, numbers and dates, and what was left undecided. Check those before anyone acts on them.

Why does an AI summary include things nobody said?

A language model writes the most likely text given the transcript and its instructions. When a meeting did not produce a clear decision or owner, it can fill the gap with one that sounds right. A summary that says "no action items were agreed" is doing its job.

How do I check an AI summary quickly?

Take each action item, decision and number, find where it was said in the transcript, and play that moment if the wording is unclear. Anything you cannot find, delete or mark as unconfirmed.

Should AI-written meeting notes be labelled?

Yes. Readers weigh a note differently when they know a model wrote it, and in the EU the AI Act adds transparency rules for some AI-generated content. Keep the label until a person has checked the note.