How Much RAM and Storage Does a Practical Local AI System Need?
Local-AI memory planning starts with the actual model file and quantization, then adds context and runtime overhead while keeping enough RAM and SSD headroom for the operating system and other applications.
There is no universal local-AI rule such as "buy 32 GB of RAM."
A practical memory plan starts with the actual model file you intend to run.
Model files create the first storage requirement
Quantized model files can still be large.
Current llama.cpp documentation gives these Q4_K_M examples for Llama 3.1:
- 8B: about 4.9 GB in its high-level sizing table;
- 70B: about 43.1 GB.
The detailed quantization table currently lists the 8B Q4_K_M variant at roughly 4.58 GiB.
Those are examples for one model family and quantization method.
Use the exact file size of the model you plan to download.
RAM must hold more than the file
Inference needs memory for:
- model weights;
- context/KV cache;
- runtime overhead;
- the operating system;
- other applications.
That means a 5 GB model file does not imply that a machine with exactly 5 GB free RAM is comfortably sized.
Leave headroom.
Quantization changes the equation
Lower-bit quantization reduces model size and memory demand.
That can let an existing machine run a model that would otherwise be too large.
The tradeoff can include lower model quality or different performance characteristics.
Do not choose quantization purely by smallest file size.
Context length costs memory
Long context windows require additional memory.
If your real workload is short chats, the requirements can be much lower than a workflow that feeds very large documents or long conversation histories.
A model that runs comfortably at one context setting may become memory-constrained at another.
Storage adds up faster than RAM
Only one model may be loaded at a time, but all downloaded models occupy disk space.
If you keep:
- several model families;
- multiple quantizations;
- embedding models;
- temporary downloads;
- local indexes;
SSD use can grow quickly.
Plan enough free space for the operating system and normal applications too.
Swap is not free RAM
Linux swap or Windows virtual memory can prevent an immediate out-of-memory crash.
It does not make slow storage behave like fast system RAM.
Heavy swapping can make local inference painfully slow.
Treat swap as a fallback, not a substitute for adequate memory.
Use actual measurements
A practical sizing process is:
- choose a model;
- note its file size and quantization;
- choose a realistic context length;
- load it on the current machine;
- watch RAM/VRAM use;
- note how much memory remains for normal work.
Then decide whether an upgrade is justified.
For the broader CPU/GPU picture, see What Kind of Computer Do You Need to Run AI Locally?.
For deciding whether to spend the money at all, see When Running AI Locally Is Not Worth the Hardware Cost.
Local-AI sizing works best when it starts with a workload and model—not a shopping target.
- Categories: Computer Hardware & Upgrades
- Tags: #Local AI