Thursday, 4:15 p.m. You are three paragraphs into a reply about a contract, your wrists hurt, and you would rather say the rest than type it. You hold a shortcut, talk for forty seconds, and the words appear in the mail window you were already in. Nothing opened. Nothing uploaded.
That last sentence is the whole point of this article, and it is worth being precise about it.
The trade most dictation asks you to make
Dictation is usually sold on accuracy and speed. The part that goes unmentioned is where the audio goes. In a lot of tools, your voice is streamed to a server, converted there, and sent back as text. It works well. It also means the sentence you were nervous about saying out loud has left your machine.
You notice this the moment the content gets real. People happily dictate a shopping list and then type out the paragraph about a salary, a diagnosis, or a colleague. The tool has quietly taught them where its limits are.
What runs on your Mac
Zen Whisper is a native Apple Silicon app for macOS 14 and later. Core speech recognition runs on your Mac, using speech packs you download and keep on the machine. Optional network-backed features exist and are separate, and the privacy policy sets out the exact data flow rather than leaving you to infer it.
The practical result is not a feeling. It is that the list of things you are willing to dictate stops being shorter than the list of things you write.
It types where you are already typing
Text is inserted at the active cursor, so it works in most Mac apps where you can already type or paste. There is no separate editor to dictate into and then copy out of. The copy step is where dictation tools usually lose the argument, because two extra actions turn a fast tool into a slow one.
Formatting takes its cue from where you are. A chat app stays casual, mail stays formal, and a terminal gets light technical cleanup. Replying in a thread, it reads the last message and matches it, so a casual conversation does not suddenly receive a paragraph of business prose.
110 languages, including the ones usually left out
The app exposes 110 spoken language options, including Hindi, Hinglish, English regional variants, and many Indian, regional, and Western languages, depending on the speech model you choose. If you switch languages mid-day, that matters more than another point of accuracy in English.
Files, media, and searchable history
- Transcribe uploaded audio and video, and supported public links, in the same place you dictate.
- Come back to a searchable transcript history rather than hunting for the window you dictated into.
- Keep voice memos, snippets, and a dictionary that remembers the words you actually use.
What it costs, plainly
A 14-day free trial, and after that it stays usable at 3,000 words a week on the free tier. Paid seat packs cover 1, 2, 5, or 10 Macs, sold on Gumroad. Each purchase is a lifetime license for that bundle plus one year of updates, and existing customers can renew update access later at a reduced price.
That is a deliberate shape. Dictation is a utility you use every day, and a monthly subscription for a utility is the thing people resent paying twice.
What it does not do
- It is Mac only, and Apple Silicon only. There is no iPhone or iPad version.
- On-device recognition means the speech packs live on your disk. Larger, more accurate models take more space, and that is a real trade rather than a free lunch.
- Optional network-backed features are exactly that. If you want everything local, leave them off.
Zen Whisper is available now with a free trial on Gumroad. If you want the longer feature list first, the Zen Whisper page has it.
How to tell where your voice is actually going
Almost every dictation product describes itself as private, and the word covers several different arrangements. The useful question is not whether it is private, it is where the audio is turned into text.
Four arrangements that all get called private
| Arrangement | Audio leaves the machine | What a breach exposes |
|---|---|---|
| Fully on device | No | Nothing. There is nothing to breach |
| On device, cloud fallback for hard cases | Sometimes, often silently | Whatever fell back |
| Cloud, deleted after processing | Yes | Whatever had not been deleted yet |
| Cloud, retained for model training | Yes | Everything you ever said |
The second row is the one worth knowing about, because it is common and it is usually invisible. A product can be truthful in saying it runs on device and still send the sentences it found difficult, which are disproportionately the unusual ones, which are disproportionately the ones you would not have sent.
The aeroplane test
There is a two second version of this check. Turn the network off entirely and dictate a paragraph. If the text appears, the recognition is genuinely local. If it stalls, hangs or quietly degrades, it was not, whatever the settings screen says.
It is worth doing once per product and once per major update, because the answer changes and nobody announces it when it does.
What running locally costs, honestly
- Battery. Real time speech recognition is sustained work for the machine, and on a laptop away from power you will notice it.
- Disk. Models are large, and a better model is a larger one.
- A first run that is slower than you expect, while the model loads.
- Occasionally worse accuracy on hard audio than the largest cloud models, which have no size or power budget to respect.
For most desk work on modern Apple Silicon those costs are small enough to ignore. They are not zero, and a product that claims they are is not being straight with you.
What you get in return
The gain is not really a privacy feeling, it is a permission change. Work that you are not allowed to send to a third party becomes work you are allowed to dictate: client material under an NDA, medical or legal notes, anything under an employer policy about where data may be processed.
That is the practical difference. Not that your grocery list is safer, but that the category of things you can use dictation for at all gets considerably larger.

