- Source
- Model registry API
- Verified
- 25 Aug 2026
- Confidence
- High
01What this is
- Author
- unknown
- Family
- unknown
- Task
- unknown
- Hugging Face
- Qwen/Qwen2.5-7B-Instruct-AWQ
- OpenRouter
- not listed
- Also known as
- none recorded
02Registry facts
- LicenceFACT
apache-2.0
- FreshRetrieved 25 Aug 2026
- ParametersFACT
7.6B
- FreshRetrieved 25 Aug 2026
- Context lengthFACT
not returned by the registry
- Evidence not verified
- Downloads (30 days)FACT
4,109,975
- FreshRetrieved 25 Aug 2026
- LikesFACT
51
- FreshRetrieved 25 Aug 2026
- GatedFACT
no
- FreshRetrieved 25 Aug 2026
Model
- downloadsFACT
4109975
- FreshRetrieved 25 Aug 2026
- gatedFACT
false
- FreshRetrieved 25 Aug 2026
- likesFACT
51
- FreshRetrieved 25 Aug 2026
- model licenseFACT
apache-2.0
- FreshRetrieved 25 Aug 2026
- parameter countFACT
7615616512
- FreshRetrieved 25 Aug 2026
Other
- last modifiedFACT
2024-10-09T12:26:45.000Z
- FreshRetrieved 25 Aug 2026
- libraryFACT
transformers
- FreshRetrieved 25 Aug 2026
03VRAM by quantisation
Peak VRAM for serving this model, computed by lib/deployment/hardware.ts from weights + KV cache + 15% runtime overhead. Inputs: 7.6B parameters, 8,192 tokens of context (assumed — none fetched yet), 4 concurrent requests (assumed).
| Quantisation | Weights | Total |
|---|---|---|
| fp16 | 15.23 GB | 23.07 GB |
| int8 | 7.62 GB | 14.31 GB |
| int4 | 3.81 GB | 9.94 GB |
What this estimate assumes
- Weights: 7.6B parameters × 2 bytes per parameter (fp16) = 15.23 GB.
- No model configuration was supplied, so the KV cache uses the size heuristic for this band: 0.147 MB per token, from Qwen3-8B (36 layers × 8 KV heads × 128 head dim).
- KV cache: 0.147 MB per token × 8,192 tokens of context × 4 concurrent requests = 4.83 GB, sized for every user holding a full context at once.
- Overhead: 15% of weights plus KV cache for the runtime, activations and memory fragmentation = 3.01 GB.
- GB means 10⁹ bytes, matching how GPU memory is advertised.
- KV cache held at fp16 (2 bytes per element); quantising weights does not by itself quantise the cache.
- Peak figure, not average: prefix caching and paged attention usually keep real usage lower, while speculative decoding, CUDA graphs and a second resident model push it higher.
- "Concurrent" counts requests generated at the same instant, not people: 4 simultaneous requests is the worst case this estimate is built on — measure yours before buying.
04Hardware profiles
No hardware profile was sized for this answer.
05Hosted providers
No hosted provider has been fetched for this model. OpenRouter’s catalogue is read by pnpm sync; until it runs, this is empty rather than assumed.
06Deployment stacks
No published deployment stack names this model yet.
07Evidence
- 01–07Tier 4official repository / model cardHugging Face Hub
Hugging Face model card data — Qwen/Qwen2.5-7B-Instruct-AWQ
https://huggingface.co/api/models/Qwen/Qwen2.5-7B-Instruct-AWQ
FreshRetrieved 25 Aug 2026sha256:8035c16e21fc7 records · · · · · ·
Improve this page
Sign in to contribute
From the field
0 deployments · 0 questions
Nobody has reported deploying this here yet, and no question has been opened against this page. Both appear once a reviewer accepts them.