name: whisper
description: OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
version: 1.0.0
author: Orchestra Research
license: MIT
dependencies: [openai-whisper, transformers, torch]
metadata:
hermes:
tags: [Whisper, Speech Recognition, ASR, Multimodal, Multilingual, OpenAI, Speech-To-Text, Transcription, Translation, Audio Processing]
Whisper - Robust Speech Recognition
OpenAI's multilingual speech recognition model.
When to use Whisper
Use when:
- - Speech-to-text transcription (99 languages)
- - Podcast/video transcription
- - Meeting notes automation
- - Translation to English
- - Noisy audio transcription
- - Multilingual audio processing
- - 72,900+ GitHub stars
- - 99 languages supported
- - Trained on 680,000 hours of audio
- - MIT License
- - AssemblyAI: Managed API, speaker diarization
- - Deepgram: Real-time streaming ASR
- - Google Speech-to-Text: Cloud-based
- - English (en)
- - Spanish (es)
- - French (fr)
- - German (de)
- - Italian (it)
- - Portuguese (pt)
- - Russian (ru)
- - Japanese (ja)
- - Korean (ko)
- - Chinese (zh)
- - GitHub: https://github.com/openai/whisper ⭐ 72,900+
- - Paper: https://arxiv.org/abs/2212.04356
- - Model Card: https://github.com/openai/whisper/blob/main/model-card.md
- - Colab: Available in repo
- - License: MIT
Metrics:
Use alternatives instead:
Quick start
Installation
`bash
Requires Python 3.8-3.11
pip install -U openai-whisper
Requires ffmpeg
macOS: brew install ffmpeg
Ubuntu: sudo apt install ffmpeg
Windows: choco install ffmpeg
`
Basic transcription
`python
import whisper
Load model
model = whisper.load_model("base")
Transcribe
result = model.transcribe("audio.mp3")
Print text
print(result["text"])
Access segments
for segment in result["segments"]:
print(f"[{segment['start']:.2f}s - {segment['end']:.2f}s] {segment['text']}")
`
Model sizes
`python
Available models
models = ["tiny", "base", "small", "medium", "large", "turbo"]
Load specific model
model = whisper.load_model("turbo") # Fastest, good quality
`
| Model | Parameters | English-only | Multilingual | Speed | VRAM |
| ------- | ------------ | -------------- | -------------- | ------- | ------ |
| tiny | 39M | ✓ | ✓ | ~32x | ~1 GB |
| base | 74M | ✓ | ✓ | ~16x | ~1 GB |
| small | 244M | ✓ | ✓ | ~6x | ~2 GB |
| medium | 769M | ✓ | ✓ | ~2x | ~5 GB |
| large | 1550M | ✗ | ✓ | 1x | ~10 GB |
| turbo | 809M | ✗ | ✓ | ~8x | ~6 GB |
| Model | Real-time factor (CPU) | Real-time factor (GPU) | |||
| ------- | ------------------------ | ------------------------ | |||
| tiny | ~0.32 | ~0.01 | |||
| base | ~0.16 | ~0.01 | |||
| turbo | ~0.08 | ~0.01 | |||
| large | ~1.0 | ~0.05 |
Real-time factor: 0.1 = 10× faster than real-time
Language support
Top-supported languages:
Full list: 99 languages total
Limitations
1. Hallucinations - May repeat or invent text
2. Long-form accuracy - Degrades on >30 min audio
3. Speaker identification - No diarization
4. Accents - Quality varies
5. Background noise - Can affect accuracy
6. Real-time latency - Not suitable for live captioning
Resources