Server-side speech-to-text powered by OpenAI Whisper. Transcribe, translate and export subtitles for any audio file.
Upload Audio
🎤Drop your audio file here
or click to browse — MP3, WAV, FLAC, OGG, AAC, M4A…
Input preview
⚠️ File exceeds the 10-second limit
Free accounts can transcribe up to 10 seconds. Please split your file with our audio splitter, or subscribe for unlimited duration and premium processing.
⭐ This feature requires a Premium account
Longer audio and additional export formats are available to Premium subscribers. Upgrade your account to unlock all features.
Vocal Isolation
Isolate vocals with htdemucs (2-stem) before transcribing — usually improves accuracy on music, but adds processing time and cost. Turn it off to transcribe the original audio directly.
1
Higher = cleaner vocal isolation, slower. 1 (fast) to 5.
Shorter segments use less memory but may introduce more boundary artefacts.
Higher overlap reduces boundary artefacts at the cost of extra processing time.
🤖 faster-whisper-small🌏 auto-detect📈 beam = 1📢 VAD on
faster-whisper Small (CTranslate2). Fast on CPU, good accuracy.
Language
1
Higher beam size explores more hypotheses — more accurate but slower. Default is 5 for large models, 1 for tiny.
Audio chunk size fed to Whisper. 30 s is the standard window.
Voice Activity Detection (VAD)
VAD filters silent segments, improving speed and accuracy on sparse speech.
Word-level Timestamps
Enables per-word timing in SRT/VTT output. Slightly slower.
Output Format
Plain text transcript — clean, easy to copy.
SubRip subtitle format (.srt) — compatible with most video players and editors.
WebVTT subtitle format (.vtt) — web standard, use with HTML5 <track>.
Full JSON with chunks, timestamps and language metadata — ideal for programmatic use.
⭐ Premium Processing
Processing threads
More threads finish faster. Cost is billed per audio-second per thread at the whisper model's rate (scaled by beam size), plus vocal isolation (10/s, scaled by shifts) when enabled. If the shared pool is busy you'll be asked to start with fewer threads or wait.
When the job finishes
Save-to-storage keeps the transcript until you delete it. Browser results stay available for 48 hours.
⏱ Est. time: –🎫 Est. cost: – tokensBalance: – tokens
Processing
queued
Position - of - in queue — your job will start shortly.
0%
Step:-ETA:-
Transcript
We use cookies to improve performance, analytics, and marketing.
Essential cookies are always active.