Skip to content
Is there an AI for this?

Model

Whisper large-v3-turbo

Recorded as suitable for transcription, multilingual.

Source
Model registry API
Verified
25 Aug 2026
Confidence
High

01What this is

Author
openai
Family
Whisper
Task
speech-to-text
Hugging Face
openai/whisper-large-v3-turbo
OpenRouter
not listed
Also known as
openai/whisper-large-v3-turbo, openai-whisper-large-v3-turbo, whisper-large-v3-turbo, whisper turbo, whisper v3 turbo

02Registry facts

25 Aug 2026
LicenceFACT

mit

FreshRetrieved 25 Aug 2026
ParametersFACT

809M

FreshRetrieved 25 Aug 2026
Context lengthFACT

not returned by the registry

Evidence not verified
Downloads (30 days)FACT

7,658,635

FreshRetrieved 25 Aug 2026
LikesFACT

3,262

FreshRetrieved 25 Aug 2026
GatedFACT

no

FreshRetrieved 25 Aug 2026

Model

downloadsFACT

7658635

FreshRetrieved 25 Aug 2026
gatedFACT

false

FreshRetrieved 25 Aug 2026
likesFACT

3262

FreshRetrieved 25 Aug 2026
model licenseFACT

mit

FreshRetrieved 25 Aug 2026
parameter countFACT

808878080

FreshRetrieved 25 Aug 2026

Other

last modifiedFACT

2024-10-04T14:51:11.000Z

FreshRetrieved 25 Aug 2026
libraryFACT

transformers

FreshRetrieved 25 Aug 2026

03VRAM by quantisation

ASSESSMENT

Peak VRAM for serving this model, computed by lib/deployment/hardware.ts from weights + KV cache + 15% runtime overhead. Inputs: 809M parameters, 8,192 tokens of context (assumed — none fetched yet), 4 concurrent requests (assumed).

Decimal GB (10⁹ bytes), matching how GPU memory is advertised.
QuantisationWeightsTotal
fp161.62 GB7.42 GB
int80.81 GB6.49 GB

What this estimate assumes

  • Weights: 0.8B parameters × 2 bytes per parameter (fp16) = 1.62 GB.
  • No model configuration was supplied, so the KV cache uses the size heuristic for this band: 0.147 MB per token, from Qwen3-8B (36 layers × 8 KV heads × 128 head dim).
  • KV cache: 0.147 MB per token × 8,192 tokens of context × 4 concurrent requests = 4.83 GB, sized for every user holding a full context at once.
  • Overhead: 15% of weights plus KV cache for the runtime, activations and memory fragmentation = 0.97 GB.
  • GB means 10⁹ bytes, matching how GPU memory is advertised.
  • KV cache held at fp16 (2 bytes per element); quantising weights does not by itself quantise the cache.
  • Peak figure, not average: prefix caching and paged attention usually keep real usage lower, while speculative decoding, CUDA graphs and a second resident model push it higher.
  • "Concurrent" counts requests generated at the same instant, not people: 4 simultaneous requests is the worst case this estimate is built on — measure yours before buying.

04Hardware profiles

No hardware profile was sized for this answer.


05Hosted providers

No hosted provider has been fetched for this model. OpenRouter’s catalogue is read by pnpm sync; until it runs, this is empty rather than assumed.


06Deployment stacks

0 using it

No published deployment stack names this model yet.


07Evidence

7 verified
  1. 01–07
    Tier 4official repository / model cardHugging Face Hub

    Hugging Face model card data — openai/whisper-large-v3-turbo

    https://huggingface.co/api/models/openai/whisper-large-v3-turbo

    FreshRetrieved 25 Aug 2026sha256:e9e711191ca9

    7 records · · · · · ·

Improve this page

Sign in to contribute

From the field

0 deployments · 0 questions

Nobody has reported deploying this here yet, and no question has been opened against this page. Both appear once a reviewer accepts them.