NotaFrog / Real time voice to text

Free online tool

Real time voice to text

Press record and the words appear while you are still speaking. Live speech to text in your browser — free, no sign-up, nothing to install.

Free real time voice to text converter

Idle
Your words show up here. Allow the microphone when your browser asks.
0 words · 0 chars

The box is editable — click into it to fix a word. Nothing on this page is sent to NotaFrog or saved anywhere; close the tab and the text is gone.

Last updated: 21 August 2026

What is real time voice to text?

Real time voice to text is speech recognition that returns words while you are still speaking, instead of waiting until you stop. Audio is sent as a continuous stream and the recogniser emits partial guesses that get corrected as more of the sentence arrives. That is why text in the tool above appears greyed out first and then settles — you are watching the model change its mind.

The alternative is batch transcription: record a file, upload it, get the whole transcript back at once. Batch is usually more accurate, because the model can see the entire sentence before it commits. Real time is what you want when somebody has to read along now.

How to use the real time voice to text tool

  1. Pick your language in the dropdown. Accuracy drops sharply when the language is wrong, and it will not correct itself.
  2. Press "Start speaking" and allow microphone access when the browser asks. The dot turns red while it is listening.
  3. Talk normally. Do not slow down or over-enunciate — recognisers are trained on natural speech and get worse when you fight them.
  4. Edit in place, then copy or download. The transcript is a normal editable field, so names and jargon can be fixed by hand before the text goes anywhere else.

Real time vs batch transcription

 Real time (streaming)Batch (upload a file)
When text appearsWhile the sentence is still being spokenAfter the whole recording is processed
AccuracyLower — early words are guessed from partial contextHigher — the model sees the full sentence
Best forLive captions, dictation, following along in a second languageInterviews, lectures, anything you will read later
Can you re-run it?No — the audio is gone once it has passedYes, the file is still there
Cost patternBills for the whole session, silence includedBills only for the audio you chose to upload

Most people who think they need real time actually want batch. If nobody is reading the text as it appears, record first and transcribe afterwards — it is cheaper and the result is better.

How fast is "real time", really?

Fast enough that you stop noticing the delay, but not instant. Vendors of streaming speech engines publish latency in the low hundreds of milliseconds up to about a second — Speechmatics advertises final transcripts in under one second, and ElevenLabs quotes under 150 ms for its realtime model. What you feel in a browser is a partial word landing within a fraction of a second, then the sentence firming up a beat later.

Two things dominate that delay, and neither is the model itself: the network round trip, and how long the recogniser waits before it decides a phrase has ended. The second one is why a live tool can feel sluggish on short utterances and snappy on long ones.

Be sceptical of any product that says "instant" without a number. A tool that transcribes in fixed chunks — a common and much cheaper design — can be ten seconds behind the speaker and still call itself live.

How accurate is real time voice to text?

Expect a clean transcript from a decent microphone with one person speaking at a time, and expect it to degrade fast outside those conditions. Accuracy is measured as word error rate, the share of words inserted, deleted or substituted. What pushes it up is predictable:

  • Distance from the microphone. The single biggest factor and the cheapest to fix. A headset beats any software setting.
  • Crosstalk. Two voices at once is the hardest case; one of them usually vanishes.
  • Names, jargon and acronyms. Anything specific to your company or field was not in the training data. Expect to correct the same words every time.
  • The wrong language setting. A recogniser told to expect English will confidently render Vietnamese as English nonsense.
  • Noise and reverb. A hard-surfaced room is worse than an open-plan office.

Treat any live transcript as a strong first draft. It is not a record, and anything that matters should be read by a person before it is circulated.

Does it work offline?

Not on this page. The tool above uses the browser's built-in Web Speech API, and in Chrome that sends your audio to Google's servers for recognition — which is exactly why it is free and why it needs a connection. Genuinely offline speech to text means running a model on your own machine, which costs you a large download and, usually, accuracy.

NotaFrog transcribes in the cloud too. We say so plainly because most free tools do not: "runs in your browser" and "your audio never leaves your device" are not the same claim, and only one of them is true here.

Can it tell speakers apart?

Not reliably. Working out who said what is called speaker diarization, and it is a different problem from transcription — it needs several seconds of each voice before it can label anyone confidently, which is awkward in a live stream. The tool on this page does not attempt it, and speaker detection inside NotaFrog is not dependable enough to build a workflow on yet.

The workaround costs nothing: say names out loud. "Sam, you own the migration by Friday" transcribes perfectly. A nod does not.

Browser dictation or a transcription app?

Use the free tool on this page for a paragraph you are about to paste somewhere. Use an app when the text has to survive the session.

 This free browser toolNotaFrog
Sign-upNoneFree account
LanguagesWhatever your browser ships99 spoken languages
Upload a recordingNo — live microphone onlyYes
Others can follow alongNoLive captions in a room others can open
OutputRaw text you copy outCleaned, structured, editable notes
Kept anywhereNo — gone when the tab closesSaved to your account, searchable and taggable
ExportCopy or .txtMarkdown, HTML, PDF
Browser supportChrome, Edge, Safari — not FirefoxEvery modern browser

The row that matters is the output. A raw transcript is not a note: turning ten minutes of talking into something a colleague can read takes filler removal, headings and structure, which is the work that happens after transcription finishes. More on that in what AI transcription actually is, and in what an AI note taking app does with the transcript.

What does it cost?

The tool on this page is free and stays free — it uses a capability your browser already has, so it costs nothing to offer. NotaFrog itself has four plans: Free at $0 forever, Lite at $1 for 7 days with 300 transcription minutes and 1,250 translation credits, Pro at $8 per month, and Team at $18 per month.

Count minutes, not features. Add up how much audio you actually want transcribed in a week — most conversations do not need a transcript at all. If that number is small, Free is a real plan rather than a demo. Full details are on the pricing page.

Privacy

Nothing dictated on this page is sent to NotaFrog, stored by NotaFrog, or attached to an account. Audio goes from your browser to your browser vendor's speech service, and the text stays in the tab until you close it.

Sign in and transcribe inside NotaFrog and the picture changes: audio is sent to a speech-recognition provider to produce a transcript, no recording library is kept, and the note is saved to your account. The privacy policy names who processes what.

Related guides

Talking is fasterthan typing.

Keep the transcript instead of losing it. NotaFrog turns what you said into a structured note you can edit, tag and export.