The right Ollama model is the smallest current model that handles your real workload without filling the memory available for inference. Ollama is a free program that downloads and runs a large language model (LLM) on local hardware. An LLM is a program trained to generate and interpret language, much like a very large predictive text system that can follow instructions.
In short: Check the exact download size in the table, then allow additional RAM (the computer's working memory) or video memory for the context and Ollama's runtime. Qwen3:4b, Llama 3.2 3B and Phi-4 Mini are sensible compact trials. Qwen3:8b, Qwen3.5:9b and Llama 3.1 8B occupy the next general-purpose tier, while Qwen2.5-Coder remains a practical coding-specific family. Smaller models usually run faster, but hardware, quantisation and context length can change the result.
Ollama Model Sizes and Download Table
The download size is not the same as the memory required while the model is running. It tells you how much storage the listed Ollama tag consumes when pulled, while runtime memory also includes the active context and software overhead.
A parameter is one of the learned values inside a model. Parameter counts such as 4B or 8B are comparable to the number of adjustable connections in a very large control panel: more controls can support greater capability, but they also require more storage and computation.
Ollama Model Download Sizes (Ollama library, September 2026)
| Ollama tag | Parameter size | Listed download size | Default for family | |
|---|---|---|---|---|
| Llama 3.1 | llama3.1:8b | 8B | 4.9GB | Yes |
| Llama 3.1 | llama3.1:70b | 70B | 43GB | No |
| Llama 3.1 | llama3.1:405b | 405B | 243GB | No |
| Llama 3.2 | llama3.2:1b | 1B | 1.3GB | No |
| Llama 3.2 | llama3.2:3b | 3B | 2.0GB | Yes |
| Llama 3.3 | llama3.3:70b | 70B | 43GB | Yes |
| Qwen2.5 | qwen2.5:0.5b | 0.5B | 398MB | No |
| Qwen2.5 | qwen2.5:1.5b | 1.5B | 986MB | No |
| Qwen2.5 | qwen2.5:3b | 3B | 1.9GB | No |
| Qwen2.5 | qwen2.5:7b | 7B | 4.7GB | Yes |
| Qwen2.5 | qwen2.5:14b | 14B | 9.0GB | No |
| Qwen2.5 | qwen2.5:32b | 32B | 20GB | No |
| Qwen2.5 | qwen2.5:72b | 72B | 47GB | No |
| Qwen2.5-Coder | qwen2.5-coder:0.5b | 0.5B | 398MB | No |
| Qwen2.5-Coder | qwen2.5-coder:1.5b | 1.5B | 986MB | No |
| Qwen2.5-Coder | qwen2.5-coder:3b | 3B | 1.9GB | No |
| Qwen2.5-Coder | qwen2.5-coder:7b | 7B | 4.7GB | Yes |
| Qwen2.5-Coder | qwen2.5-coder:14b | 14B | 9.0GB | No |
| Qwen2.5-Coder | qwen2.5-coder:32b | 32B | 20GB | No |
| Qwen3 | qwen3:0.6b | 0.6B | 523MB | No |
| Qwen3 | qwen3:1.7b | 1.7B | 1.4GB | No |
| Qwen3 | qwen3:4b | 4B | 2.5GB | No |
| Qwen3 | qwen3:8b | 8B | 5.2GB | Yes |
| Qwen3 | qwen3:14b | 14B | 9.3GB | No |
| Qwen3 | qwen3:30b | 30B | 19GB | No |
| Qwen3 | qwen3:32b | 32B | 20GB | No |
| Qwen3 | qwen3:235b | 235B | 142GB | No |
| Qwen3.5 | qwen3.5:0.8b | 0.8B | 1.0GB | No |
| Qwen3.5 | qwen3.5:2b | 2B | 2.7GB | No |
| Qwen3.5 | qwen3.5:4b | 4B | 3.4GB | No |
| Qwen3.5 | qwen3.5:9b | 9B | 6.6GB | Yes |
| Qwen3.5 | qwen3.5:27b | 27B | 17GB | No |
| Qwen3.5 | qwen3.5:35b | 35B | 24GB | No |
| Qwen3.5 | qwen3.5:122b | 122B | 81GB | No |
| Gemma 3 | gemma3:270m | 270M | 292MB | No |
| Gemma 3 | gemma3:1b | 1B | 815MB | No |
| Gemma 3 | gemma3:4b | 4B | 3.3GB | Yes |
| Gemma 3 | gemma3:12b | 12B | 8.1GB | No |
| Gemma 3 | gemma3:27b | 27B | 17GB | No |
| Mistral | mistral:7b | 7B | 4.4GB | Yes |
| Phi-4 Mini | phi4-mini:3.8b | 3.8B | 2.5GB | Yes |
| DeepSeek-R1 | deepseek-r1:1.5b | 1.5B | 1.1GB | No |
| DeepSeek-R1 | deepseek-r1:7b | 7B | 4.7GB | No |
| DeepSeek-R1 | deepseek-r1:8b | 8B | 5.2GB | Yes |
| DeepSeek-R1 | deepseek-r1:14b | 14B | 9.0GB | No |
| DeepSeek-R1 | deepseek-r1:32b | 32B | 20GB | No |
| DeepSeek-R1 | deepseek-r1:70b | 70B | 43GB | No |
| DeepSeek-R1 | deepseek-r1:671b | 671B | 404GB | No |
| gpt-oss | gpt-oss:20b | 20B | 14GB | Yes |
| gpt-oss | gpt-oss:120b | 120B | 65GB | No |
Sizes as listed on the Ollama library, checked September 2026.
The source rows are the official Ollama library pages for Llama 3.1, Llama 3.2, Llama 3.3, Qwen2.5, Qwen2.5-Coder, Qwen3, Qwen3.5, Gemma 3, Mistral, Phi-4 Mini, DeepSeek-R1 and gpt-oss.
How Ollama Model Names Work
An Ollama name follows the family:tag pattern. The family identifies the model line, while the tag selects a parameter size or a more specific variant. For example, qwen3:4b selects Qwen3 at the 4B tier.
Leaving out the tag selects that family's current default, also called latest. In the table, ollama pull qwen3 resolves to the 8B download, while ollama pull llama3.2 resolves to the 3B download. A default alias can change when its publisher updates the library, so scripts and repeatable deployments should specify the intended tag.
Quantisation is a method of storing model values at reduced numerical precision to lower memory and computation requirements. It is like compressing a detailed map for a phone: the smaller version is easier to carry, but some fine detail can be lost. The quantisation guide explains why two variants with the same parameter count can have different sizes and quality.
An instruct model has been tuned to respond to directions and conversation. A base model mainly predicts text continuations, like an autocomplete engine without the customer-service training. For chat, summarising and question answering, use the instruction-tuned or default conversational variant unless a development workflow specifically requires a base model.
How Much RAM or VRAM Does an Ollama Model Need?
The practical rule is that the model file plus working memory for context must fit in available RAM or video memory. RAM is the computer's general working area. Video memory, or VRAM, is memory attached to a graphics processor, like a separate workbench reserved for graphics and model calculations.
A context window is the amount of prompt and conversation the model can consider at once. Think of it as the number of pages that can remain open on a desk: a larger desk helps with long documents, but occupies more room. Ollama's documentation warns that increasing context length increases memory use.
Do not treat a 5.2GB download as proof that the model will run comfortably with exactly 5.2GB free. The operating system, Ollama, context cache, concurrent requests and other services all need memory. A system can also split work between RAM and VRAM, although keeping the model on a supported graphics processor usually gives better interactive performance.
Use this sequence before pulling a large tag:
- Check genuinely available memory after the operating system and normal services are running.
- Choose a model file that leaves room rather than consuming nearly all available memory.
- Start with a modest context length and one active request.
- Run
ollama pswhile the model is loaded to inspect its processor placement and allocated size. - Move to a larger tag only if the smaller model fails the actual workload.
The full Ollama model guide by 16GB, 32GB and 64GB RAM covers the hardware tiers in detail. The AI hardware selector can also narrow the hardware class before a purchase.
Which Ollama Model Is Fast?
Smaller models are usually faster than larger models on the same hardware and software configuration. There is no honest universal fastest-model figure because speed also depends on quantisation, memory bandwidth, processor support, context length, prompt length and whether the model fits fully in VRAM.
A compact model that stays entirely in fast memory can feel more responsive than a more capable model that spills into slower system memory. This is like keeping a job on the workbench instead of walking to a storeroom for every part. A longer context also takes more memory and makes prompt processing heavier.
For a quick first trial, compare llama3.2:3b, qwen3:4b and phi4-mini:3.8b on the same hardware. If those models are not accurate enough, move to qwen3:8b, qwen3.5:9b or another larger tag. This is a workload-fit sequence, not a benchmark ranking.
Reasoning models can take longer to produce an answer because their task may involve more generated reasoning. DeepSeek-R1 and gpt-oss can suit reasoning work, but they should not be selected merely because the family is more capable on difficult prompts. A short summarisation or classification task may be better served by a smaller conventional model.
Which Model to Pull for Each Task
General chat
For general conversation, begin with a current model in the 4B to 9B class that fits with headroom. qwen3:4b is a compact starting point, while qwen3:8b, qwen3.5:9b and llama3.1:8b offer larger alternatives. Test them with the same five real prompts because preferred tone and instruction following vary by workload.
Llama 3.1 remains widely available, but it is no longer the only sensible default. Qwen3.5 adds text and image input in its listed local tags, while Gemma 3 offers image-capable options from 4B upward. Do not download a multimodal model solely for text chat if a smaller text model already meets the requirement.
Coding
qwen2.5-coder:7b is a practical coding-specific download at 4.7GB. Its family also provides 3B, 14B and 32B tiers, making it easy to test the same model line against the available memory.
The 7B tag suits code explanation, small edits and debugging trials where a local assistant is useful but memory is limited. Larger coding models can improve difficult work, but repository-scale tasks also need more context and therefore more working memory. Always validate generated code with tests and review rather than treating model output as authoritative.
Summarising and rewriting
llama3.2:3b, qwen3:4b and gemma3:4b are compact candidates for summaries, rewriting and extracting information from modest inputs. Select the smallest one that preserves the facts and structure required by the document.
Long documents change the calculation. Even when a model advertises a large supported context window, increasing the active context consumes additional memory. Split documents into sensible sections when the task allows it instead of assuming the maximum context is free.
Reasoning and complex questions
deepseek-r1:8b is the current default DeepSeek-R1 download at 5.2GB. The larger distill tags step through 14B, 32B and 70B, while the full 671B tag is a 404GB download and is outside ordinary home hardware.
gpt-oss:20b is a newer 14GB reasoning and agentic option. It is relevant to systems with substantially more memory than a compact 4B or 8B model requires. Do not choose it for a low-memory server or when rapid short answers matter more than difficult reasoning.
Small or low-RAM hardware
gemma3:1b, llama3.2:1b, qwen3:1.7b, qwen3:4b and phi4-mini:3.8b provide useful steps across the compact tiers. Very small models suit classification, rewriting and simple extraction, but are more likely to lose instructions or produce weak answers on complex work.
The real trade-off is speed and fit against answer quality. A 1B model that responds immediately but fails the task is not efficient, while a 70B model that makes the system unusably slow is also a poor fit.
Three Common Model Selection Mistakes
The first mistake is equating download size with required runtime memory. The listed file is only the starting point, and long context or concurrent users can materially increase memory use.
The second mistake is pulling an unqualified latest tag into a permanent workflow. The alias is convenient for testing, but an explicit tag makes the chosen parameter tier visible and avoids an unnoticed move to a different default.
The third mistake is collecting many models before defining a test. Keeping three similar 8B chat models consumes storage without proving which one suits the workload. Prepare a small prompt set, compare answers, keep the winner and remove the rest.
A Practical Pull Checklist
Use this checklist to turn the size table into a decision:
- Define the task: chat, coding, summarising, vision or reasoning.
- Check free RAM and VRAM while normal services are running.
- Choose the smallest plausible tag with memory headroom.
- Pull an explicit tag, such as
ollama pull qwen3:4b. - Test five representative prompts and record accuracy, omissions and response time.
- Check placement and allocated memory with
ollama ps. - Increase the model size only when the smaller option fails a defined test.
- Remove unused downloads with
ollama rm model:tag.
Ollama's current command-line reference lists ollama pull for downloading, ollama ls for installed models, ollama ps for running models and ollama rm for removal. A graphical interface can be added later through the Ollama frontends and Open WebUI guide.
Australian Buyers: What You Need to Know
Australian users running Ollama continuously should compare idle power, active power and useful performance, not just the purchase price. Annual energy use can be estimated as average power in kilowatts multiplied by 24 hours, 365 days and the applicable electricity tariff. Measure at the wall under a representative mix of idle and active use because a maximum power rating is not an average.
A NAS, or network-attached storage device, connects to a router and provides shared storage and services to other devices. Think of it as a shared digital filing cabinet that can also run selected applications. Running Ollama on an existing NAS can avoid another always-on box, but a CPU-only NAS may deliver poor interactive speed and its memory is shared with storage services.
A mini-PC separates local AI from storage duties and may offer stronger processors, more memory bandwidth or supported graphics acceleration. It also adds another device to power and maintain. The mini-PC versus NAS comparison helps decide which arrangement suits the workload, while the local LLM on a NAS guide explains the limitations to check.
Downloading a model uses the listed amount of internet data once, apart from later updates. After that, local prompts do not need to upload the model over the NBN. Remote access to a home-hosted interface still depends on the connection's upload performance and network configuration.
CGNAT, or Carrier-Grade NAT, is a network arrangement some internet providers use that can prevent direct inbound access to a home connection. It works like an apartment building sharing one street address, where an unexpected visitor cannot reach a specific unit without another routing method. Use a properly secured remote-access design rather than exposing Ollama or a web interface directly to the internet.
Related reading: our NAS buyer's guide, our NAS vs cloud storage comparison, and our NAS explainer.
Use our free AI Hardware Requirements Calculator to size the hardware you need to run AI locally.
What is the fastest Ollama model?
There is no single fastest model across every computer. On the same hardware, a smaller model such as Gemma 3 1B, Llama 3.2 1B or Qwen3 1.7B will usually generate responses faster than a much larger model, but quantisation, context length, memory bandwidth and GPU placement also matter.
How much RAM does an Ollama model need?
Allow more memory than the listed download size because Ollama also needs working memory for context, runtime overhead and concurrent requests. Check genuinely free RAM or VRAM, begin with a shorter context and use ollama ps to inspect the loaded model rather than relying on the file size alone.
What is the Qwen2.5-Coder 7B download size in Ollama?
The official Ollama library lists qwen2.5-coder:7b at 4.7GB. Runtime memory will be higher than that file size, so a machine with exactly 4.7GB free should not be assumed to fit it comfortably.
How large are the Qwen3 4B and 8B models?
The Ollama library lists qwen3:4b at 2.5GB and qwen3:8b at 5.2GB. The 8B tag is the current family default, so ollama pull qwen3 downloads the 5.2GB model unless the default changes later.
What are the official Llama 3.2 model sizes in Ollama?
Ollama lists llama3.2:1b at 1.3GB and llama3.2:3b at 2.0GB. The 3B tag is the default and is downloaded when the command omits a tag.
Is the latest Ollama tag always the best choice?
No. latest is the family's default alias, not a guarantee that the model is best for a particular task or computer. Use it for exploration, but specify a parameter tag in repeatable setups so the intended size remains clear.
Can a 64GB computer run a 70B Ollama model?
A listed 70B download may be 43GB, but that does not guarantee a comfortable fit or useful speed on a 64GB system. The active context, operating system, runtime overhead, memory architecture and processor performance all matter, so check the exact tag and the full by-RAM guide before committing to the download.
Should unused Ollama models be deleted?
Yes, if local storage is constrained or two downloads serve the same workload. Run ollama ls to review installed models and ollama rm model:tag to remove an unused tag; it can be pulled again later if required.
Need a chat interface to use these models? Open WebUI runs in Docker on a NAS or mini-PC and serves every device on your network from a single installation.