Voice to Text Help & FAQ

Turn a recording, an imported file or live dictation into accurate text on your iPhone, iPad or Mac - with speakers labelled, timestamps you can jump to, and nothing uploaded.

Voice to Text: Offline AI Note app icon Voice to Text: Offline AI Note View product page →

Frequently asked questions

Does my audio get uploaded anywhere?

No. Recognition runs on your iPhone, iPad or Mac. There is no account and no upload step, and the app works with the network switched off entirely - which is also the simplest way to prove to yourself that nothing is leaving.

What audio and video formats can I import?

m4a, mp3, wav, aiff and caf for audio, and mp4 and mov for video. Importing exists because the recording you most need a transcript of is usually one somebody else made and sent you.

How does the dictation shortcut work?

Press Control-Option-D and dictation starts wherever the cursor already is, in whatever app you are using. Text lands directly in that document instead of in a separate window you then copy from.

How accurate is speaker identification?

It separates the voices in a recording and lets you name them, which is what makes a transcript quotable. It works best when speakers do not talk over each other and the microphone is roughly equidistant - a phone in the middle of a table beats a phone in front of one person.

Can I get subtitles for a video out of it?

Yes. Export as SRT or WebVTT and the timings come from the same pass that produced the transcript, so the captions line up with the audio they were made from.

Which export format should I pick?

Plain text for pasting somewhere quickly, Markdown when you want speaker names and structure preserved in a notes app or document, SRT or WebVTT for video captions, and JSON when another tool is going to read the transcript rather than a person.

Why does the transcript appear as I speak on one machine but not another?

Real-time transcription needs a newer system; on older ones the text arrives after the recording rather than during it. The result is the same either way - only the timing of when you see it differs.

Which languages are supported?

Well over ninety depending on the engine, including German, Spanish, French, Italian, Portuguese, Dutch, Polish, Czech, Ukrainian, Swedish, Danish, Finnish, Greek, Turkish, Arabic, Hebrew, Hindi, Chinese, Japanese and Korean.

Can I find a moment again without listening through the whole file?

Yes. Every transcript goes into a searchable library, and selecting a line moves the audio to that timestamp - which is the normal way to check a quote before you use it.

What do I need to run it?

iPhone on iOS 17, iPad on iPadOS 17, or a Mac on macOS 14, and later versions of each.

Voice to Text: Offline AI Note

Get Voice to Text: Offline AI Note

Record, import or dictate and get accurate text in seconds - entirely on device, exporting to notes, documents or subtitles.

How-to guides

How to transcribe a meeting you are about to attend

The decisions that matter are made before the meeting starts, not afterwards.

  1. Put the device where it can hear everyone - the middle of the table, not next to you.
  2. Start the recording and check the live waveform actually moves when someone speaks.
  3. Let it run; do not stop and restart between agenda items, since one file keeps the speaker labels consistent.
  4. When the meeting ends, open the transcript and name each speaker once.
  5. Export as Markdown so the names and structure survive into your notes.

Tell the room it is being recorded. That is a courtesy in most places and a legal requirement in some.

Try this in Voice to Text: Offline AI Note →

How to turn a recording somebody sent you into a transcript

Works the same for an audio file and for the audio track of a video.

  1. Import the file - m4a, mp3, wav, aiff, caf, mp4 or mov.
  2. Wait for the transcript; nothing is uploaded, so this runs at the speed of your own machine.
  3. Name the speakers if there is more than one voice.
  4. Search the transcript for the part you actually need and select the line to jump to that moment in the audio.
  5. Export in the format the next step expects.

If the source is a video you are going to publish, export SRT rather than text - you get the transcript and the captions from one pass.

Try this in Voice to Text: Offline AI Note →

How to dictate notes straight into another app

This is the mode people underuse, because it does not require opening the app at all.

  1. Put the cursor where the text should appear - a note, an email, a document.
  2. Press Control-Option-D to start dictation.
  3. Speak normally; punctuation is easier to add afterwards than to dictate.
  4. Press the shortcut again to stop.
  5. Tidy the result in place, or send it to TextLab if it needs reshaping.

Dictation is the fastest route from a thought to a written line, and it is worth learning the shortcut properly - it removes the step where you open an app first and lose the sentence you had.

Try this in Voice to Text: Offline AI Note →

How to subtitle a video you made

The aim is captions whose timings came from the real audio rather than being nudged by hand.

  1. Import the video file directly - there is no need to extract the audio first.
  2. Let the transcript finish, then read it once for names, jargon and numbers, which is where recognition slips.
  3. Fix those words in the transcript before exporting, so the correction carries into the captions.
  4. Export as SRT for most editors and platforms, or WebVTT for the web.
  5. Attach the file to the video in your editor and check two or three moments against the audio.

Proper nouns are the reliable failure case in every transcription engine. Reading for names first, before anything else, catches most of what a viewer would notice.

Try this in Voice to Text: Offline AI Note →

Related guides

Still need help?

Our support team usually replies within one business day.

Download Voice to Text: Offline AI Note Contact support