You recorded a forty minute conversation. Turning it into text means choosing a service, creating an account, and uploading a file that contains another person talking candidly because they trusted the room, not the room plus a vendor.
That consent question is the part transcription tools rarely put on the pricing page. The person you interviewed agreed to talk to you.
What changes when the file never leaves
Zen Whisper transcribes uploaded audio and video, and supported public links, on your Mac. Core speech recognition uses speech packs you download and keep. There is no account to create for the recognition itself, and no upload step to explain to anybody.
For anybody handling interviews, medical notes, legal calls, or research, that removes a conversation with a compliance team as well as a technical dependency.
The part that matters a week later
A transcript you cannot find is a file you paid to create twice. Transcripts land in a searchable history, alongside voice memos and snippets, so the useful quote is retrievable without remembering which folder you were in when you ran the job.
The workflow
- Bring in the audio or video file, or a supported public link.
- Choose the speech pack. Larger models are more accurate and slower; that trade is yours to make per job.
- Let it run on your Mac, then read the transcript from searchable history.
- Dictate your own notes into the same app, at the cursor, in whatever you are writing in.
What to expect, honestly
- Local transcription is bounded by your Mac. A long recording on a larger model takes real time, and that is the cost of not uploading it.
- Accuracy varies with audio quality. Two people over a bad connection is hard for every system, local or not.
- Mac only, Apple Silicon, macOS 14 or later.
Free for 14 days, then 3,000 words a week on the free tier, with lifetime seat packs on Gumroad. The Zen Whisper page lists the rest.
What "uploading an interview" actually commits you to
An interview recording is rarely only yours. It contains a second person who agreed to talk to you, not to a transcription vendor, and who was never shown that vendor’s retention policy. In a lot of professional contexts that is not a preference, it is the reason the recording exists under terms at all.
- Journalism, where a source agreed to speak to a named person and nobody else.
- Research interviews, where an ethics approval usually names where the data may be stored.
- HR and legal conversations, where the recording is evidence and its chain of custody matters.
- Medical and therapeutic work, where the bar is set by regulation rather than by taste.
None of that means cloud transcription is wrong. It means the decision belongs to you rather than to whichever tool was quickest to install.
Local transcription is genuinely practical now, with caveats
Speech models that run on an Apple Silicon Mac are good enough for real work, and they were not a few years ago. The honest caveats are about speed and about hard audio.
Where each approach wins
| Situation | Local | Cloud |
|---|---|---|
| A one hour interview, clear audio | Fine, and nothing leaves the Mac | Faster, and the file leaves the Mac |
| Heavy accents or overlapping speakers | Workable, expect editing | Usually better, at a cost |
| Sensitive or embargoed material | The obvious answer | A decision you have to justify |
| A laptop on battery in a hurry | Costs battery and time | Costs bandwidth |
Getting a usable transcript rather than a rough one
- Record the room, not the laptop. A cheap external microphone placed between two people beats an expensive one pointed at the ceiling.
- Say who is speaking at the start. "This is Priya, interviewing Sam" gives you an anchor no diarisation setting can invent.
- Avoid recording over air conditioning or a fridge. Steady background noise is worse for accuracy than a single loud interruption.
- Transcribe the first two minutes before the interview ends if you can, as a check that the audio is actually being captured.
- Read the transcript against the audio for the parts you intend to quote. Every model gets numbers, names and negations wrong occasionally, and those are exactly the words that carry meaning.
That last one is not a criticism of any particular tool. It is the difference between a transcript as a search index, which is what it is good at, and a transcript as a quotation source, which needs a human ear on the sentences that matter.

