Skip to content
AI transcription

Audio to text converter

Turn voice memos, calls, interviews and podcasts into text you can read, search and quote. Free and without an account.

  • No account, no signup
  • TXT, SRT and VTT are generated in your browser
  • Size and length limits shown before you upload

How it works

  1. Add your recording

    MP3, WAV, M4A, AAC, OGG and FLAC all work. Most phone recordings are M4A; most downloads are MP3.

  2. Tell us the language

    Setting the language explicitly avoids the most common cause of a bad transcript: the wrong language being assumed.

  3. Clean it up and export

    Correct names and jargon in the browser, then take the text as TXT or the timed version as SRT or VTT.

What actually decides audio accuracy

The single biggest factor is the distance between the speaker and the microphone. A phone lying on a table two metres away picks up far more room reverberation than speech, and reverberation is much harder to interpret than quiet. A recording made with the phone held near the speaker will usually transcribe better than one made with a more expensive microphone placed badly.

The second factor is overlapping speech. When two people talk at the same time, the audio contains both voices mixed together and there is no reliable way to pull them apart. In practice, sections where people interrupt each other are where you will find most of the errors, and they are worth checking first when you review a transcript.

Background noise matters, but not in the way people expect. Steady noise — a fan, an air conditioner, traffic hum — is handled surprisingly well because it is predictable. Sudden noise is worse: a door slamming, cutlery, a chair scraping. These can swallow an entire word, and no amount of processing recovers what was never captured.

Recording habits that pay off

If you have any control over how the audio is captured, a few small habits improve the result more than any setting on this page. Put the microphone closer than feels necessary. Record in a room with soft furnishings rather than a bare meeting room, because hard walls create the reverberation mentioned above. Ask people to avoid talking over each other, which helps the transcript and the meeting.

For interviews specifically, it is worth starting the recording with each person saying their name. It gives you an anchor when you read the transcript back and cannot remember who was speaking, since the transcript itself does not identify speakers.

Reviewing the transcript

Proper nouns are where errors cluster. Names of people, companies, products and places are the words a speech model is least likely to get right, because they are rare and often invented. Scan for them first.

Numbers deserve a second look too, particularly amounts and dates spoken quickly. If a figure matters, check it against the recording rather than trusting the transcript.

The timestamped view makes both of these checks fast: find the moment, listen to that spot in your original file, correct the text. Everything you fix is included in the file you download.

What people use it for

Voice memos

Turn a thought you recorded while walking into something you can actually use later.

Recorded calls

Keep a written record of what was agreed, without listening to the whole call again.

Interviews

Quote accurately and find the exact moment a quote came from.

Podcast episodes

Produce show notes, pull quotes, or a full transcript for people who prefer reading.

Supported formats

Audio

mp3wavm4aaacoggflac

Video

mp4movmkvwebm

Up to 1000 MB and 360 minutes

Frequently asked questions

Which audio formats can I upload?

MP3, WAV, M4A, AAC, OGG and FLAC. If your file is one of these, it will be accepted regardless of the extension it was given, because the format is detected from the file contents.

Does the transcript say who is speaking?

No. The transcript is a single stream of text without speaker labels. If you need to tell people apart, having each person state their name at the start of the recording gives you something to anchor to.

Will a noisy recording work?

Usually, but with more errors. Steady background noise such as a fan or traffic is handled reasonably. Sudden noises and people talking over each other cause the most problems.

Can I convert audio into subtitles?

Yes. Timestamps are produced for every recording, so you can download SRT or VTT even when the source was pure audio. This is useful if you plan to pair the audio with visuals later.

Ready to try it with your own file?

No account, no watermark. Pick a file and the transcript comes back on this page.

Choose a file
Audio to Text Converter — Free & Accurate | MediaToText