Catalog/ № 063
Whisper
shipOpenAI's open-weights speech recognition — the default transcription model.
github.comVerified 2026-08-15
ledger entry · observed by @whysanesanders
The reason paid transcription became a niche: large-v3 accuracy is production-grade across dozens of languages and it runs on a laptop. Every transcription pipeline starts here.
Known limitations
- Hallucinates text on silence and noisy segments.
- No built-in speaker diarization (pair with pyannote).
- large-v3 needs a decent GPU for real-time work.
- whisper-1 API is legacy; the current hosted line is the gpt-transcribe family.
Facts
- pricing
- free1
- price note
- open source (MIT), self-host free; hosted: gpt-4o-transcribe $0.006/min
- free tier
- yes
- open source
- yes
- api
- yes
- self-host
- yes
- category
- audio
Models used
Sources
If we can't show the source, we don't print the number.
Same shelf — audio
Deepgram
audio
Speech-to-text API built for developers — fast, cheap, real-time.
APIFREE TIER
freemium$200 free credits; Nova-3 STT from $0.0048/min (promo)
ElevenLabsfeatured
audio
The reference standard for AI voice — TTS, cloning, dubbing and voice agents.
APIFREE TIER
freemiumfree tier (10k credits/mo, non-commercial); Starter $6/mo, Creator $22/mo
Murf AI
audio
Business-focused voiceover studio for presentations and e-learning.
APIFREE TIER
freemiumfree tier; Creator Lite $29/mo or $23/mo billed yearly