AI Speech-to-Text Transcriber Beta
Convert audio files to text instantly using the NVIDIA Nemotron 3.5 ASR (0.6B) model.
Drop your audio or video file here
or click to browse from device (MP3, WAV, M4A, etc.)
Push-to-Talk Voice
Hold Space or press and hold the button while speaking. Release to process the clip.
Frequently Asked Questions
What model is running behind this tool?
We host the official NVIDIA Nemotron-3.5-ASR-Streaming-0.6b model. It uses the FastConformer architecture optimized for low-latency transcribing, native capitalization, and formatting.
Is my audio data secure?
Yes. Your uploaded files are stored temporarily in memory or a secure `/tmp` directory during processing, and are wiped immediately after the transcription is completed. We do not store or catalog your files.
Does it support long audio files?
Yes. The backend parses files of various durations. However, for large files, conversion and transcription might take a few moments. We limit uploads to a generous size.