What Kind of Computer Do You Need to Run AI Locally?
There is no single local-AI minimum specification: model size, quantization, context length and runtime determine whether a machine can use CPU, system RAM, GPU VRAM or a combination effectively.
There is no single answer to "What computer do I need for local AI?"
The hardware requirement depends on:
- model size;
- quantization;
- context length;
- runtime;
- number of concurrent models;
- whether you use CPU or GPU acceleration.
The best starting point is to choose a model that fits the hardware you already own.
CPU-only local AI is possible
A dedicated GPU is not mandatory.
llama.cpp supports CPU inference, and other runtimes can run smaller models primarily from system memory.
CPU-only operation is usually slower than a good supported GPU, but it can still be useful for:
- testing local AI;
- small models;
- low-volume personal use;
- machines with plenty of RAM but no strong GPU.
That makes an existing PC worth testing before buying new hardware.
System RAM matters
LM Studio currently recommends 16 GB or more RAM for typical Windows use and notes that smaller models may work with less on some platforms.
Do not turn that recommendation into a universal model limit.
The model, context and runtime all affect memory use.
More RAM gives you room for larger models and leaves more memory available for the operating system and other applications.
VRAM matters when using a dedicated GPU
A GPU can dramatically accelerate inference when the runtime supports the hardware and enough model data fits in VRAM.
LM Studio currently recommends at least 4 GB dedicated VRAM on Windows as a starting point.
That is not a guarantee that every useful model fits in 4 GB.
Larger models and longer contexts can require much more.
Models can be split between GPU and RAM
llama.cpp can offload part of a model to the GPU while keeping the rest in system memory.
That allows models larger than available VRAM to run.
The tradeoff is usually lower performance than keeping more of the model on the GPU.
A machine with modest VRAM and plenty of RAM may therefore run a model that does not fit entirely on the graphics card.
Quantization changes memory requirements
Quantization stores model weights at lower numerical precision.
That can reduce:
- file size;
- RAM use;
- VRAM use.
The tradeoff can include some loss of model quality depending on the quantization level and model.
This is why two copies of the "same" model can have very different memory requirements.
Context length also costs memory
Longer context means the model can keep more text available during a conversation or document task.
That extra context has a memory cost.
A model that fits comfortably with a modest context may become much heavier when configured for very long prompts.
Hardware sizing therefore has to include how you plan to use the model, not just its parameter count.
Storage adds up quickly
Local models are large files.
If you keep several variants, embeddings, caches and multiple models, storage use grows quickly.
An SSD is strongly preferable for loading large model files and general workstation responsiveness.
GPU support is runtime-specific
Ollama and llama.cpp support multiple acceleration paths, including NVIDIA CUDA and AMD-related backends.
Exact GPU support changes over time and is hardware-specific.
Do not buy a GPU based only on raw VRAM.
Confirm that your chosen runtime supports the card and driver stack.
Test before buying
Before building a dedicated AI workstation:
- install a local runtime;
- download a modest model;
- measure whether response speed is acceptable;
- watch RAM/VRAM use;
- identify the actual bottleneck.
Only then decide whether you need more memory, a supported GPU or an entirely new machine.
For the setup architecture, see Running an AI Assistant on Your Own Computer Instead of the Cloud.
For a Linux-centered machine, see Using Linux as the Foundation for a Local AI Workstation.