Transcription and on-device AI

Apple's SpeechAnalyzer: speech recognition in macOS 26

What Apple shipped for speech-to-text in macOS 26: SpeechAnalyzer and its transcribers, their languages, and how they differ from SFSpeechRecognizer.

By the Notey team at AInject · · 7 min read

In short

SpeechAnalyzer is the speech-to-text API Apple added in macOS 26 and the matching releases of iOS, iPadOS, tvOS and visionOS. It comes with a new model that Apple built for long-form audio, including meetings, and it runs on the device.

It sits alongside the older SFSpeechRecognizer, which has been in macOS since 10.15.

This guide is for people choosing or evaluating a transcription tool on a Mac who want to know what the platform provides underneath. It is based on Apple's developer documentation and its WWDC25 session "Bring advanced speech-to-text to your app with SpeechAnalyzer", both read on 25 September 2026.

What Apple shipped

Apple's documentation describes SpeechAnalyzer as a class that "analyzes spoken audio content in various ways and manages the analysis session". It does not transcribe anything by itself. An app adds modules to it, and each module does one kind of work:

  • SpeechTranscriber: speech-to-text "appropriate for normal conversation and general purposes", using Apple's new model.
  • DictationTranscriber: speech-to-text "similar to system dictation features and compatible with older devices", using the same models as system dictation.
  • SpeechDetector: "voice activity detection", which notices when someone is speaking and can be used alongside a transcriber.

The app gives the analyser audio, from a file or from a live source such as a microphone, as a stream, and reads results back as another stream. Apple's framework takes care of converting the audio into the format the model wants.

The model behind SpeechTranscriber

In the WWDC25 session, Apple says the model is "good for long-form and distant audio, such as lectures, meetings, and conversations", and that it is "both faster and more flexible than the one previously available through SFSpeechRecognizer". Apple uses it in its own apps, including Notes.

Three details about where the model lives matter for privacy and for disk space:

  • It "operates outside of your application's memory space", so an app does not carry the model inside itself.
  • It is "retained in system storage and does not increase the download or storage size of your application".
  • The system updates it: Apple says it "constantly improve[s] the model", and new versions are installed automatically.

Where the model comes from

The models are not part of each app. Apple's AssetInventory documentation says they are "machine-learning models downloaded from Apple's servers and managed by the system". Once one is downloaded, the system "retains and updates it automatically, and shares it with other apps".

In practice:

  • The first time an app asks for a language that is not yet on the Mac, the system downloads that language's model. This needs a connection.
  • If the model is already there (preinstalled, or fetched earlier by another app), the download finishes at once.
  • Each app can hold a limited number of languages at a time, and the system may drop a language nobody has used for a while.

The download is the model coming to your Mac. Transcription itself does not send audio anywhere; Apple describes the processing as happening on the device. On-device transcription explained goes through what that does and does not change for privacy.

Live results: volatile and finalized

A transcriber working while someone talks has to commit to words before the sentence is over. SpeechTranscriber handles this by sending two kinds of result:

  • Volatile results arrive "almost as soon as they're spoken but they are less accurate guesses".
  • Finalized results replace them once the model is sure. After a finalized result for a stretch of audio, "the transcriber won't deliver any more results for this range of audio".

This is why a live transcript built on it shows words changing for a moment and then settling. An app can show volatile text differently so you can tell which words are still provisional. The same API also transcribes whole files, which is the basis for going over a recording again after a meeting; see live versus post-meeting transcription.

Languages and hardware

SpeechTranscriber covers a short list of languages. On a Mac running macOS 26.6 in September 2026, it reported English, French, German, Italian, Spanish, Portuguese, Japanese, Korean, Chinese and Cantonese, most with several regional variants. Apple says "more to come". Which languages a Mac can transcribe on the device lists the variants.

On hardware, Apple's session says SpeechTranscriber is available on every platform but watchOS "with certain hardware requirements", and the API gives apps a check for whether the current device supports it. Apple does not give a device list in that documentation. For anything outside the supported languages or devices, Apple points to DictationTranscriber.

How it differs from SFSpeechRecognizer

SFSpeechRecognizer has been the Speech framework's main class since iOS 10 and macOS 10.15. According to Apple's documentation and the WWDC25 session:

SFSpeechRecognizerSpeechTranscriber
IntroducediOS 10, macOS 10.15iOS 26, macOS 26
Designed forShort-form dictationLong-form and distant audio, such as meetings
Runs whereOn Apple's servers, or on the device if the app requires it and the language supports itOn the device
ModelShared with Siri and dictationApple's new model, kept in system storage
LanguagesMany, some only over the networkA short list (ten as of September 2026)

Two points from Apple's SFSpeechRecognizer documentation are worth knowing if a tool still uses it:

  • For server recognition, Apple tells developers to "plan for a one-minute limit on audio duration", and says daily limits may apply per device and per app.
  • Apple says on-device requests with the older recogniser "won't be as accurate" as server ones.

DictationTranscriber is Apple's bridge between the two. It supports "the same languages, speech-to-text model, and devices as iOS 10's on-device SFSpeechRecognizer", but through the new API, and without users having to turn on Siri or keyboard dictation for a language first. Its documentation adds that it does not support languages the older recogniser handles only over the network.

What this means for choosing a tool

If a Mac transcription tool says it uses Apple's speech recognition, it is worth asking which one:

  • SpeechTranscriber means Apple's newer model, on the device, in the supported languages, on macOS 26 or later.
  • DictationTranscriber or on-device SFSpeechRecognizer means the older dictation models, on the device, in more languages.
  • SFSpeechRecognizer without the on-device setting may send audio to Apple for recognition.

None of these is a whole notetaker. Apple provides the recogniser; working out who said what, keeping the recording and turning a transcript into notes are left to the app.

How Notey uses it

Notey's live transcription uses Apple's on-device speech recognition on Macs with Apple silicon running macOS 26 or later. It transcribes your microphone and your Mac's audio as separate tracks, so each line is attributed to your side or theirs by the track it came from. A phrase still being revised is shown lighter until it settles. After an English meeting, a second on-device recogniser goes over the recording; why Notey re-transcribes a meeting after it ends explains when and why. No audio is sent anywhere for either.

Frequently asked questions

What is Apple's SpeechAnalyzer?

A speech-to-text API Apple introduced in macOS 26, iOS 26 and its other platforms except watchOS. An app adds modules to it, most often SpeechTranscriber, and feeds it audio from a file or a live source; the transcription runs on the device.

What is the difference between SpeechTranscriber and DictationTranscriber?

SpeechTranscriber uses Apple's newer model, built for long-form and distant audio such as meetings, in fewer languages and with hardware requirements. DictationTranscriber uses the same models as system dictation and the older on-device recogniser, and covers more languages and older devices.

Does SpeechAnalyzer send audio to Apple?

Apple describes the transcription as running on the device. What comes from Apple's servers is the model itself, which the system downloads once per language and shares between apps.

Is SFSpeechRecognizer deprecated?

Apple's documentation still lists it. Apple presents SpeechAnalyzer as the newer, faster and more flexible option, and DictationTranscriber as the route for apps that need what the older recogniser covered.