AIFree · No sign-up · AI runs on your device

Speech to text.

Transcribe a recording, voice note, meeting or video into text, or record straight from your microphone. The transcribing happens in your browser, so your audio stays with you.

Audio never leaves your device Over 20 languages Text or subtitles (.srt) No account needed
or record
AI: an on-device speech model (about 40–150 MB, downloaded once and kept by your browser). Your audio is never uploaded.
Your recording is never uploaded to us. It’s transcribed right here in this browser tab. The first time, your browser downloads the speech model itself — your audio is never part of that download.
How it works

Add a recording, press Transcribe, get the text

The transcriber runs OpenAI’s open-source Whisper speech model inside your browser, so the recording is processed on your own device.

01

Add a recording

Upload audio or video, or record from your microphone.

02

Transcribe

Choose the language (or let it detect it) and press Transcribe.

03

Download

Save the transcript as text, or as .srt subtitles for a video.

Why your recording never leaves your device

Most transcription services upload your audio to their servers. Altto runs the speech model inside your browser instead, so interviews, meetings and voice notes stay private.

  • Nothing is uploadedYour audio is decoded and transcribed on your own device.
  • Subtitles includedDownload an .srt file with timings, ready for video editors and YouTube.
  • No account, no minutes limitTranscribe as much as you like, free.
0

recordings ever uploaded to our servers. The speech AI runs on your device.

Before you start

Will it run on my device?

Because the AI runs on your own device instead of on a server, your computer does the work — that’s what keeps your files private. Most computers from the last few years are fine, but a more powerful one is noticeably faster.

Your device:Checking…

Minimum

  • ComputerAny laptop or desktop from about the last 6 years
  • Memory (RAM)4 GB
  • BrowserThe latest Chrome, Edge, Firefox or Safari
  • QualityUse “Fast” (about 40 MB download)

Without graphics acceleration, a recording can take as long to transcribe as it does to play.

Recommended Fastest

  • ComputerA recent laptop or desktop (e.g. Intel Core i5 / AMD Ryzen 5 or better)
  • Memory (RAM)8 GB or more
  • BrowserChrome or Edge, which can use your graphics chip (WebGPU)
  • Quality“More accurate” (about 80–150 MB download)

Phones can transcribe short voice notes; for long recordings use a computer.

Tested on a mid-range PC with AMD Ryzen built-in graphics (no separate graphics card), in Chrome: An 11-second clip transcribed in about 3 seconds once the model was downloaded (roughly 20 minutes for an hour-long recording on the same PC).

FAQ

Frequently asked questions

Will it work on my computer or phone?
Most computers from the last few years run it fine. Because the AI runs on your own device, a more powerful computer is faster. See minimum and recommended specs, including a quick check of your device.
Is my recording uploaded?
No. The audio is decoded and transcribed inside this browser tab. It is never sent to Altto or to a transcription company.
What gets downloaded?
Only the speech model: about 40 MB for Fast, or around 80–150 MB for More accurate, depending on your device. Your browser keeps it, so next time it starts straight away.
How long does it take?
It depends on your device. On a modern computer with a recent browser a one-hour recording often takes a few minutes; on older phones it’s slower. Fast is quicker; More accurate makes fewer mistakes.
Which files work?
Anything your browser can play: MP3, WAV, M4A, OGG, WebM, FLAC and the sound from MP4 or MOV videos. Some phone formats such as AMR can’t be opened by browsers — convert them to MP3 first.
Can it translate?
Yes — choose “English translation” under Output and it will write an English transcript of speech in another language.
Does it tell who is speaking?
No, it doesn’t label speakers. It starts a new paragraph when there’s a pause, and the timestamps view shows when each line was said.