Audio

Whisper

OpenAI · #6 most active of 15 in Audio

Compare →

A widely deployed multilingual transcription and translation model.

67

Trust score · Mixed

3 data points checked · 2026-09-11

Model stack: DisputedEntry price: UnverifiableUsage claim: VerifiedVersion: DisputedHow this works →

Current version

gpt-transcribe

Entry cost

Pricing not specified in the document

Changes / 30d

2

The short answers · verified September 14, 2026

What is the latest version of Whisper?
The current shipped version of Whisper is gpt-transcribe, as of September 14, 2026.
How much does Whisper cost?
Whisper has no published flat entry price. Pricing not specified in the document
What AI model does Whisper use?
Whisper runs primarily on gpt-transcribe.

Source: OpenAIopenai.com · Reusable under CC BY 4.0 — cite Tomorrow

Capabilities

  • 99 languages
  • Translation to English
  • Word timestamps
  • Open weights
  • Robust to noise
  • File transcription
  • Language detection
  • Support for mp3, mp4, mpeg, mpga, m4a, wav, and webm formats
  • Up to 25 MB file uploads
  • Custom context via prompts, keywords, and languages
  • file transcription
  • realtime transcription
  • translation
  • word timestamps
  • speaker labels
  • subtitle formats
  • Realtime transcription
  • Speech generation
  • speech-to-text
  • Subtitle formats
  • Translation into English
  • Speaker labels
  • Audio transcription
  • Audio translation
  • Support for mp3, mp4, mpeg, mpga, m4a, wav, webm
  • transcription
  • Speech-to-text
  • Translation
  • speech translation
  • speech recognition
  • multilingual transcription
  • zero-shot performance
  • timestamps
  • Speech to text
  • Subtitles
  • Translations
  • audio transcription
  • speech to text
  • file-transcription
  • audio translation
  • translation into English
  • Transcription
  • Audio to text

Pricing

  • Open Source Model

    Model weights and code available for free download

    Free

Entry price over time

7/24/2026 · $09/14/2026 · $0

Geek mode

Models under the hood

  • gpt-transcribe

    OpenAI

    disclosed
  • Whisper

    OpenAI

    disclosed
  • Whisper large-v3

    OpenAI

    disclosed
  • whisper-1

    OpenAI

    inferred

Context window

30 second windows

Public API

yes

Multi-model routing

yes

Stack signals

  • PyTorch
  • whisper.cpp
  • Transformers

Shared model stack

Other tracked products running on the same foundation models — a quick read on how much of the catalog moves when one of these models changes.

green disclosed · amber inferred · grey unattributed. Aliases are folded into one model; provider concentration counts only models with an established vendor.

Reported scale

Reference data. Each figure is whatever the source actually said — weekly users, downloads, revenue run-rate — with its own definition and date. These are not comparable between tools and are never used to rank anything.

  • GitHub stars

    106k

    stars

    Public registry

    as of Aug 2026·GitHub repository ·disclosed

  • GitHub stars

    75k

    stars

    Public registry

    Developer-side proxy only.

    as of Jun 2025·GitHub repository ·disclosed

Also in Audio

All Audio
Mureka

Kunlun TechChina

7

Music generation with stem export and a public API.

Version
Mureka AI MV Director
Cost
Free tier · from $8/mo
Model
Mureka V8
Deepgram

Deepgram

6

Real-time speech to text

Version
Nova-3
Cost
Free tier · from $4,000/min
Model
Aura-1
Suno

Suno

6

Full songs from a text prompt

Version
v6
Cost
Free tier · from $10/mo
Model
Bark
AssemblyAI

AssemblyAI

5

Speech intelligence API

Version
Universal-3.5 Pro
Cost
Free tier with $50 free credits, then pay-as-you-go starting at $0.15/hr for pre-recorded STT (Universal-2)
Model
Claude
ElevenLabs

ElevenLabs

3

Lifelike text to speech

Version
Eleven v3
Cost
Free tier · from $6/mo
Model
Dubbing v2
Cartesia

Cartesia

Ultra-low-latency voice models built on state space models.

Version
Sonic 3
Cost
Free tier · from $5/mo
Model
Sonic 3

Change history

  • capability

    New capabilities: Audio to text

    42 tracked43 tracked · +Audio to text

    source
  • capability

    New capabilities: Transcription

    41 tracked42 tracked · +Transcription

    source
  • capability

    New capabilities: translation into English

    40 tracked41 tracked · +translation into English

    source
  • capability

    New capabilities: audio translation

    39 tracked40 tracked · +audio translation

    source
  • capability

    New capabilities: file-transcription

    38 tracked39 tracked · +file-transcription

    source
  • capability

    New capabilities: audio transcription, speech to text

    36 tracked38 tracked · +audio transcription, speech to text

    source
  • version

    Whisper moved to gpt-transcribe

    Whispergpt-transcribe

    source
  • capability

    New capabilities: Speech to text, Subtitles, Translations

    33 tracked36 tracked · +Speech to text, Subtitles, Translations

    source
  • capability

    New capabilities: speech translation

    28 tracked29 tracked · +speech translation

    source
  • capability

    New capabilities: Speech-to-text, Translation

    26 tracked28 tracked · +Speech-to-text, Translation

    source
  • model

    Whisper added Whisper to its model stack

    Whisper large-v3, gpt-transcribe, whisper-1Whisper, Whisper large-v3, gpt-transcribe, whisper-1

    source
  • capability

    New capabilities: transcription

    25 tracked26 tracked · +transcription

    source
  • capability

    New capabilities: Audio transcription, Audio translation, Support for mp3, mp4, mpeg, mpga, m4a, wav, webm

    22 tracked25 tracked · +Audio transcription, Audio translation, Support for mp3, mp4, mpeg, mpga, m4a, wav, webm

    source
  • capability

    New capabilities: Speaker labels

    215

    source
  • model

    Whisper added whisper-1 to its model stack

    Whisper large-v3, gpt-transcribeWhisper large-v3, gpt-transcribe, whisper-1

    source
  • capability

    New capabilities: Subtitle formats, Translation into English

    195

    source
  • capability

    New capabilities: speech-to-text

    186

    source
  • capability

    New capabilities: Realtime transcription, Speech generation

    164

    source
  • version

    Whisper moved to gpt-transcribe

    large-v3gpt-transcribe

    source
  • capability

    New capabilities: file transcription, realtime transcription, translation

    106

    source
  • model

    Whisper added gpt-transcribe to its model stack

    Whisper large-v3Whisper large-v3, gpt-transcribe

    source
  • capability

    New capabilities: File transcription, Language detection, Support for mp3, mp4, mpeg, mpga, m4a, wav, and webm formats

    55

    source
  • model

    large-v3-turbo variant published for faster inference

    large-v3large-v3, large-v3-turbo

    source
  • version

    Whisper moved to Whisper

    gpt-transcribeWhisper

    source
  • capability

    New capabilities: speech recognition, multilingual transcription, zero-shot performance

    29 tracked33 tracked · +speech recognition, multilingual transcription, zero-shot performance, timestamps

    source
Subscribe to Whisper changes

Whisper compared

Straight head-to-head pages against the busiest products in Audio.

How to cite this page

Free to cite and reuse under CC BY 4.0. Permalink: https://tomorrow.aliensquad.ai/tools/whisper

APA
Tomorrow. (2026). Whisper — version, pricing and model stack [Data set entry]. AlienSquad. Retrieved 2026-09-17, from https://tomorrow.aliensquad.ai/tools/whisper
BibTeX
@misc{tomorrow-tools-whisper,
  author       = {{Tomorrow}},
  title        = {Whisper — version, pricing and model stack},
  year         = {2026},
  publisher    = {AlienSquad},
  howpublished = {\url{https://tomorrow.aliensquad.ai/tools/whisper}},
  note         = {Accessed: 2026-09-17}
}