RichScriptsTools for a Smarter WebAll tools

Voice to Text

Your voice, turned into text. Transcribe on your device with no audio uploads or API fees.

Local Whisper · Free · Up to 2 minutes

No audio selected

Loading Cloudflare human verification…

Choose audio or record your microphone to begin.

First use downloads a large model and runtime. Processing may take several minutes. Desktop browsers recommended.

Convert audio to text locally

Turn a short voice recording into editable text without uploading the recording to a transcription service. This speech-to-text tool runs Whisper Tiny inside your browser. Use it for personal notes, short interviews, practice recordings, and draft captions. Choose an audio file or record with your microphone, select the spoken language, and press Transcribe. The first version accepts recordings up to two minutes long and files up to twenty-five megabytes.

The first transcription downloads the model and browser runtime from Hugging Face and a public content delivery network. This can require substantial data and time, especially on a mobile connection. These services receive normal download requests, but this page does not send them your recording. Downloaded files may remain in your browser cache. A cached model can reduce later loading time, although browser storage eviction or missing runtime files can require another download.

Microphone recording starts only after you press the recording button and grant permission. Stop recording when you finish; recording automatically stops at the two-minute limit. You can listen to the selected audio before transcription. File support depends on your browser’s audio decoder. WAV and MP3 are useful choices; some browsers also decode M4A, WebM, and Ogg. An unsupported or damaged file produces an error instead of a transcript.

Processing happens locally on your device’s processor. A desktop computer usually offers a more practical experience than a low-memory phone. Download progress refers to the current model file, not the entire transcription. During recognition, the indicator remains indeterminate because the engine cannot reliably predict completion time. Cancel stops the worker and lets you retry with the same audio.

Review every result before using it. Background noise, overlapping speakers, accents, music, and unfamiliar names can reduce accuracy. Automatic language detection can also make mistakes; choose the language manually when needed. Copy the transcript or download a text file. SRT export includes estimated segment timings, not guaranteed word-level alignment. Editing the transcript disables the original SRT export to avoid mismatched captions. Clear removes the current recording and result from the page; it does not delete model files cached by your browser.

Frequently asked questions

Is my recording uploaded?

No. Audio decoding and recognition happen in your browser. The runtime and model are downloaded from external hosts.

Is this free?

Yes. There is no paid transcription API. Your device supplies the computing resources and model downloads use your internet connection.

Does it work offline?

The initial model and runtime downloads require a connection. Cache availability varies, so offline operation is not guaranteed.

Can I transcribe long recordings?

This initial version supports up to two minutes and 25 MB per recording. Split longer recordings before using it.