Turn a recording, an imported file or live dictation into accurate text on your iPhone, iPad or Mac - with speakers labelled, timestamps you can jump to, and nothing uploaded.
No. Recognition runs on your iPhone, iPad or Mac. There is no account and no upload step, and the app works with the network switched off entirely - which is also the simplest way to prove to yourself that nothing is leaving.
What audio and video formats can I import?
m4a, mp3, wav, aiff and caf for audio, and mp4 and mov for video. Importing exists because the recording you most need a transcript of is usually one somebody else made and sent you.
How does the dictation shortcut work?
Press Control-Option-D and dictation starts wherever the cursor already is, in whatever app you are using. Text lands directly in that document instead of in a separate window you then copy from.
How accurate is speaker identification?
It separates the voices in a recording and lets you name them, which is what makes a transcript quotable. It works best when speakers do not talk over each other and the microphone is roughly equidistant - a phone in the middle of a table beats a phone in front of one person.
Can I get subtitles for a video out of it?
Yes. Export as SRT or WebVTT and the timings come from the same pass that produced the transcript, so the captions line up with the audio they were made from.
Which export format should I pick?
Plain text for pasting somewhere quickly, Markdown when you want speaker names and structure preserved in a notes app or document, SRT or WebVTT for video captions, and JSON when another tool is going to read the transcript rather than a person.
Why does the transcript appear as I speak on one machine but not another?
Real-time transcription needs a newer system; on older ones the text arrives after the recording rather than during it. The result is the same either way - only the timing of when you see it differs.
Which languages are supported?
Well over ninety depending on the engine, including German, Spanish, French, Italian, Portuguese, Dutch, Polish, Czech, Ukrainian, Swedish, Danish, Finnish, Greek, Turkish, Arabic, Hebrew, Hindi, Chinese, Japanese and Korean.
Can I find a moment again without listening through the whole file?
Yes. Every transcript goes into a searchable library, and selecting a line moves the audio to that timestamp - which is the normal way to check a quote before you use it.
What do I need to run it?
iPhone on iOS 17, iPad on iPadOS 17, or a Mac on macOS 14, and later versions of each.
Get Voice to Text: Offline AI Note
Record, import or dictate and get accurate text in seconds - entirely on device, exporting to notes, documents or subtitles.
This is the mode people underuse, because it does not require opening the app at all.
Put the cursor where the text should appear - a note, an email, a document.
Press Control-Option-D to start dictation.
Speak normally; punctuation is easier to add afterwards than to dictate.
Press the shortcut again to stop.
Tidy the result in place, or send it to TextLab if it needs reshaping.
Dictation is the fastest route from a thought to a written line, and it is worth learning the shortcut properly - it removes the step where you open an app first and lose the sentence you had.
The aim is captions whose timings came from the real audio rather than being nudged by hand.
Import the video file directly - there is no need to extract the audio first.
Let the transcript finish, then read it once for names, jargon and numbers, which is where recognition slips.
Fix those words in the transcript before exporting, so the correction carries into the captions.
Export as SRT for most editors and platforms, or WebVTT for the web.
Attach the file to the video in your editor and check two or three moments against the audio.
Proper nouns are the reliable failure case in every transcription engine. Reading for names first, before anything else, catches most of what a viewer would notice.