NotaFrog / Real time voice to text
Free online toolReal time voice to text
Press record and the words appear while you are still speaking. Live speech to text in your browser — free, no sign-up, nothing to install.
Free real time voice to text converter
The box is editable — click into it to fix a word. Nothing on this page is sent to NotaFrog or saved anywhere; close the tab and the text is gone.
Last updated: 21 August 2026
What is real time voice to text?
Real time voice to text is speech recognition that returns words while you are still speaking, instead of waiting until you stop. Audio is sent as a continuous stream and the recogniser emits partial guesses that get corrected as more of the sentence arrives. That is why text in the tool above appears greyed out first and then settles — you are watching the model change its mind.
The alternative is batch transcription: record a file, upload it, get the whole transcript back at once. Batch is usually more accurate, because the model can see the entire sentence before it commits. Real time is what you want when somebody has to read along now.
How to use the real time voice to text tool
- Pick your language in the dropdown. Accuracy drops sharply when the language is wrong, and it will not correct itself.
- Press "Start speaking" and allow microphone access when the browser asks. The dot turns red while it is listening.
- Talk normally. Do not slow down or over-enunciate — recognisers are trained on natural speech and get worse when you fight them.
- Edit in place, then copy or download. The transcript is a normal editable field, so names and jargon can be fixed by hand before the text goes anywhere else.
Real time vs batch transcription
| Real time (streaming) | Batch (upload a file) | |
|---|---|---|
| When text appears | While the sentence is still being spoken | After the whole recording is processed |
| Accuracy | Lower — early words are guessed from partial context | Higher — the model sees the full sentence |
| Best for | Live captions, dictation, following along in a second language | Interviews, lectures, anything you will read later |
| Can you re-run it? | No — the audio is gone once it has passed | Yes, the file is still there |
| Cost pattern | Bills for the whole session, silence included | Bills only for the audio you chose to upload |
Most people who think they need real time actually want batch. If nobody is reading the text as it appears, record first and transcribe afterwards — it is cheaper and the result is better.
How fast is "real time", really?
Fast enough that you stop noticing the delay, but not instant. Vendors of streaming speech engines publish latency in the low hundreds of milliseconds up to about a second — Speechmatics advertises final transcripts in under one second, and ElevenLabs quotes under 150 ms for its realtime model. What you feel in a browser is a partial word landing within a fraction of a second, then the sentence firming up a beat later.
Two things dominate that delay, and neither is the model itself: the network round trip, and how long the recogniser waits before it decides a phrase has ended. The second one is why a live tool can feel sluggish on short utterances and snappy on long ones.
Be sceptical of any product that says "instant" without a number. A tool that transcribes in fixed chunks — a common and much cheaper design — can be ten seconds behind the speaker and still call itself live.
How accurate is real time voice to text?
Expect a clean transcript from a decent microphone with one person speaking at a time, and expect it to degrade fast outside those conditions. Accuracy is measured as word error rate, the share of words inserted, deleted or substituted. What pushes it up is predictable:
- Distance from the microphone. The single biggest factor and the cheapest to fix. A headset beats any software setting.
- Crosstalk. Two voices at once is the hardest case; one of them usually vanishes.
- Names, jargon and acronyms. Anything specific to your company or field was not in the training data. Expect to correct the same words every time.
- The wrong language setting. A recogniser told to expect English will confidently render Vietnamese as English nonsense.
- Noise and reverb. A hard-surfaced room is worse than an open-plan office.
Treat any live transcript as a strong first draft. It is not a record, and anything that matters should be read by a person before it is circulated.
Does it work offline?
Not on this page. The tool above uses the browser's built-in Web Speech API, and in Chrome that sends your audio to Google's servers for recognition — which is exactly why it is free and why it needs a connection. Genuinely offline speech to text means running a model on your own machine, which costs you a large download and, usually, accuracy.
NotaFrog transcribes in the cloud too. We say so plainly because most free tools do not: "runs in your browser" and "your audio never leaves your device" are not the same claim, and only one of them is true here.
Can it tell speakers apart?
Not reliably. Working out who said what is called speaker diarization, and it is a different problem from transcription — it needs several seconds of each voice before it can label anyone confidently, which is awkward in a live stream. The tool on this page does not attempt it, and speaker detection inside NotaFrog is not dependable enough to build a workflow on yet.
The workaround costs nothing: say names out loud. "Sam, you own the migration by Friday" transcribes perfectly. A nod does not.
Browser dictation or a transcription app?
Use the free tool on this page for a paragraph you are about to paste somewhere. Use an app when the text has to survive the session.
| This free browser tool | NotaFrog | |
|---|---|---|
| Sign-up | None | Free account |
| Languages | Whatever your browser ships | 99 spoken languages |
| Upload a recording | No — live microphone only | Yes |
| Others can follow along | No | Live captions in a room others can open |
| Output | Raw text you copy out | Cleaned, structured, editable notes |
| Kept anywhere | No — gone when the tab closes | Saved to your account, searchable and taggable |
| Export | Copy or .txt | Markdown, HTML, PDF |
| Browser support | Chrome, Edge, Safari — not Firefox | Every modern browser |
The row that matters is the output. A raw transcript is not a note: turning ten minutes of talking into something a colleague can read takes filler removal, headings and structure, which is the work that happens after transcription finishes. More on that in what AI transcription actually is, and in what an AI note taking app does with the transcript.
What does it cost?
The tool on this page is free and stays free — it uses a capability your browser already has, so it costs nothing to offer. NotaFrog itself has four plans: Free at $0 forever, Lite at $1 for 7 days with 300 transcription minutes and 1,250 translation credits, Pro at $8 per month, and Team at $18 per month.
Count minutes, not features. Add up how much audio you actually want transcribed in a week — most conversations do not need a transcript at all. If that number is small, Free is a real plan rather than a demo. Full details are on the pricing page.
Privacy
Nothing dictated on this page is sent to NotaFrog, stored by NotaFrog, or attached to an account. Audio goes from your browser to your browser vendor's speech service, and the text stays in the tab until you close it.
Sign in and transcribe inside NotaFrog and the picture changes: audio is sent to a speech-recognition provider to produce a transcript, no recording library is kept, and the note is saved to your account. The privacy policy names who processes what.
Related guides
Talking is fasterthan typing.
Keep the transcript instead of losing it. NotaFrog turns what you said into a structured note you can edit, tag and export.