- Source
- Model registry API
- Verified
- 25 Aug 2026
- Confidence
- High
01What this is
- Author
- Qwen
- Family
- Qwen2.5
- Task
- text-generation
- Hugging Face
- Qwen/Qwen2.5-7B-Instruct
- OpenRouter
- qwen/qwen-2.5-7b-instruct
- Also known as
- Qwen2.5-7B-Instruct, qwen2.5 7b, qwen 7b instruct
02Registry facts
- LicenceFACT
apache-2.0
- FreshRetrieved 25 Aug 2026
- ParametersFACT
7.6B
- FreshRetrieved 25 Aug 2026
- Context lengthFACT
not returned by the registry
- Evidence not verified
- Downloads (30 days)FACT
11,430,234
- FreshRetrieved 25 Aug 2026
- LikesFACT
1,555
- FreshRetrieved 25 Aug 2026
- GatedFACT
no
- FreshRetrieved 25 Aug 2026
Model
- downloadsFACT
11430234
- FreshRetrieved 25 Aug 2026
- gatedFACT
false
- FreshRetrieved 25 Aug 2026
- likesFACT
1555
- FreshRetrieved 25 Aug 2026
- model licenseFACT
apache-2.0
- FreshRetrieved 25 Aug 2026
- parameter countFACT
7615616512
- FreshRetrieved 25 Aug 2026
Other
- last modifiedFACT
2025-01-12T02:10:10.000Z
- FreshRetrieved 25 Aug 2026
- libraryFACT
transformers
- FreshRetrieved 25 Aug 2026
03VRAM by quantisation
Peak VRAM for serving this model, computed by lib/deployment/hardware.ts from weights + KV cache + 15% runtime overhead. Inputs: 7.6B parameters, 8,192 tokens of context (assumed — none fetched yet), 4 concurrent requests (assumed).
| Quantisation | Weights | Total |
|---|---|---|
| fp16 | 15.23 GB | 23.07 GB |
| int8 | 7.62 GB | 14.31 GB |
| q4_k_mSized as 4-bit weights. Q4_K_M averages nearer 4.5 bits in practice, so read this as a floor. | 3.81 GB | 9.94 GB |
What this estimate assumes
- Weights: 7.6B parameters × 2 bytes per parameter (fp16) = 15.23 GB.
- No model configuration was supplied, so the KV cache uses the size heuristic for this band: 0.147 MB per token, from Qwen3-8B (36 layers × 8 KV heads × 128 head dim).
- KV cache: 0.147 MB per token × 8,192 tokens of context × 4 concurrent requests = 4.83 GB, sized for every user holding a full context at once.
- Overhead: 15% of weights plus KV cache for the runtime, activations and memory fragmentation = 3.01 GB.
- GB means 10⁹ bytes, matching how GPU memory is advertised.
- KV cache held at fp16 (2 bytes per element); quantising weights does not by itself quantise the cache.
- Peak figure, not average: prefix caching and paged attention usually keep real usage lower, while speculative decoding, CUDA graphs and a second resident model push it higher.
- "Concurrent" counts requests generated at the same instant, not people: 4 simultaneous requests is the worst case this estimate is built on — measure yours before buying.
04Hardware profiles
Cloud GPU instance — 1 × NVIDIA L4 24 GB
- GPU
- NVIDIA L4 24 GB (AWS G6, Azure NVadsA10/NCads equivalents, GCP G2)
- VRAM
- 24 GB
- System RAM
- 64 GB
- Storage
- 1000 GB
- CPU
- 8–16 vCPU
- Form factor
- Cloud instance
Indicative costUS$1 – US$2
indicative on-demand rental, USD per hour, Aug 2026 — verify against the provider price list for your region
The private-cloud counterpart of `onprem-small-24gb`: same model sizes, no capital outlay, and a region you choose explicitly. Running it continuously for a year usually costs more than buying the equivalent box, so it suits pilots, bursts and firms without a server room. The cloud provider becomes a data processor — a DPA and a documented region are required.
Sizing basis
- Weights: 7.6B parameters × 2 bytes per parameter (fp16) = 15.23 GB.
- No model configuration was supplied, so the KV cache uses the size heuristic for this band: 0.147 MB per token, from Qwen3-8B (36 layers × 8 KV heads × 128 head dim).
- KV cache: 0.147 MB per token × 8,192 tokens of context × 4 concurrent requests = 4.83 GB, sized for every user holding a full context at once.
- Overhead: 15% of weights plus KV cache for the runtime, activations and memory fragmentation = 3.01 GB.
- GB means 10⁹ bytes, matching how GPU memory is advertised.
- KV cache held at fp16 (2 bytes per element); quantising weights does not by itself quantise the cache.
- Peak figure, not average: prefix caching and paged attention usually keep real usage lower, while speculative decoding, CUDA graphs and a second resident model push it higher.
- "Concurrent" counts requests generated at the same instant, not people: 4 simultaneous requests is the worst case this estimate is built on — measure yours before buying.
On-premise single 24 GB GPU server
- GPU
- NVIDIA RTX 4090 24 GB (or NVIDIA L4 24 GB for a rack-mounted, 72 W alternative)
- VRAM
- 24 GB
- System RAM
- 64 GB
- Storage
- 2000 GB
- CPU
- 16-core x86 server CPU (AMD EPYC 7003/9004 or Intel Xeon Scalable)
- Form factor
- Tower server
Indicative costUS$4,000 – US$9,000
indicative build cost for the complete machine, USD, Aug 2026 — verify with a local supplier
The default box for a 20–60 person firm. Fits a 14B model at 4-bit with roughly 8 GB of KV cache left for concurrent chat, or an 8B model at fp16. NVMe storage sized for the model cache plus a document corpus and its embeddings. Add a UPS and an offsite backup target — this machine holds the whole knowledge base.
Sizing basis
- Weights: 7.6B parameters × 2 bytes per parameter (fp16) = 15.23 GB.
- No model configuration was supplied, so the KV cache uses the size heuristic for this band: 0.147 MB per token, from Qwen3-8B (36 layers × 8 KV heads × 128 head dim).
- KV cache: 0.147 MB per token × 8,192 tokens of context × 4 concurrent requests = 4.83 GB, sized for every user holding a full context at once.
- Overhead: 15% of weights plus KV cache for the runtime, activations and memory fragmentation = 3.01 GB.
- GB means 10⁹ bytes, matching how GPU memory is advertised.
- KV cache held at fp16 (2 bytes per element); quantising weights does not by itself quantise the cache.
- Peak figure, not average: prefix caching and paged attention usually keep real usage lower, while speculative decoding, CUDA graphs and a second resident model push it higher.
- "Concurrent" counts requests generated at the same instant, not people: 4 simultaneous requests is the worst case this estimate is built on — measure yours before buying.
CPU-only server (small models and embeddings)
- GPU
- unknown
- VRAM
- unknown
- System RAM
- 64 GB
- Storage
- 1000 GB
- CPU
- 16–32 core x86 server CPU with AVX-512
- Form factor
- Tower server
Indicative costUS$1,500 – US$4,000
indicative build cost for the complete machine, USD, Aug 2026 — verify with a local supplier
No GPU. Embedding models and document parsing run acceptably here, which is enough for a search-only pilot or a nightly batch pipeline. Chat generation with a 7–8B model at 4-bit works but reads at a few tokens per second — usable for one person testing, not for a team. The honest use of this profile is to prove the retrieval quality before buying a GPU.
Sizing basis
- Weights: 7.6B parameters × 2 bytes per parameter (fp16) = 15.23 GB.
- No model configuration was supplied, so the KV cache uses the size heuristic for this band: 0.147 MB per token, from Qwen3-8B (36 layers × 8 KV heads × 128 head dim).
- KV cache: 0.147 MB per token × 8,192 tokens of context × 4 concurrent requests = 4.83 GB, sized for every user holding a full context at once.
- Overhead: 15% of weights plus KV cache for the runtime, activations and memory fragmentation = 3.01 GB.
- GB means 10⁹ bytes, matching how GPU memory is advertised.
- KV cache held at fp16 (2 bytes per element); quantising weights does not by itself quantise the cache.
- Peak figure, not average: prefix caching and paged attention usually keep real usage lower, while speculative decoding, CUDA graphs and a second resident model push it higher.
- "Concurrent" counts requests generated at the same instant, not people: 4 simultaneous requests is the worst case this estimate is built on — measure yours before buying.
05Hosted providers
No hosted provider has been fetched for this model. OpenRouter’s catalogue is read by pnpm sync; until it runs, this is empty rather than assumed.
06Deployment stacks
No published deployment stack names this model yet.
07Evidence
- 01–07Tier 4official repository / model cardHugging Face Hub
Hugging Face model card data — Qwen/Qwen2.5-7B-Instruct
https://huggingface.co/api/models/Qwen/Qwen2.5-7B-Instruct
FreshRetrieved 25 Aug 2026sha256:c22d7792f0fc7 records · · · · · ·
Improve this page
Sign in to contribute
From the field
0 deployments · 0 questions
Nobody has reported deploying this here yet, and no question has been opened against this page. Both appear once a reviewer accepts them.