Llama 3.1 8B

Llama8.03B paramsLlama 3.1 Community License128K contextReleased 2024-07-23

The default first local model: small enough for an 8 GB card at Q4, good enough to be genuinely useful, and the standard against which starter AI builds are sized.

Last verified 2026-07 · Page created; quant sizes verified against the official GGUF repos.

Why this is the starter model

Llama 3.1 8B is the model most people run first: at Q4_K_M the weights are under 5 GB, so it fits comfortably on an 8 GB GPU with room for context, and it streams faster than reading speed on anything mid-range or better.

It is also the honest baseline for sizing a first AI build — if a machine cannot run an 8B at Q4 briskly, it is not an AI machine. Step up in VRAM, not in hype: the jump that matters is to 16 GB (13–14B class and long contexts), then 24 GB (32B class).

VRAM by quantization

QuantFile sizeVRAM @ 8K ctxVRAM @ 32K ctx8 GB16 GB24 GB32 GB48 GB
FP1616.1 GB19.6 GB23.1 GBDoesn't fitDoesn't fitFits comfortablyFits comfortablyFits comfortably
Q8_08.5 GB11.0 GB14.4 GBDoesn't fitFits comfortablyFits comfortablyFits comfortablyFits comfortably
Q5_K_M5.7 GB7.7 GB11.2 GBFits comfortablyFits comfortablyFits comfortablyFits comfortablyFits comfortably
Q4_K_M4.9 GB6.8 GB10.3 GBFits comfortablyFits comfortablyFits comfortablyFits comfortablyFits comfortably

Fit verdicts use the 8K context preset. Estimated from published weight sizes plus standard KV-cache math — see methodology.

Runs on

RTX 5060 (8 GB)
best fit Q5_K_MFits comfortably
RTX 3070 (8 GB)
best fit Q5_K_MFits comfortably
RTX 3060 Ti (8 GB)
best fit Q5_K_MFits comfortably
RTX 4060 Ti (8 GB)
best fit Q5_K_MFits comfortably
RTX 5050 (8 GB)
best fit Q5_K_MFits comfortably
RTX 4060 (8 GB)
best fit Q5_K_MFits comfortably
RTX 4070 Laptop (8 GB)
best fit Q5_K_MFits comfortably
RX 7600 (8 GB)
best fit Q5_K_MFits comfortably
RTX 4060 Laptop (8 GB)
best fit Q5_K_MFits comfortably
Arc A750 (8 GB)
best fit Q5_K_MFits comfortably
RTX 3080 (10 GB)
best fit Q5_K_MFits comfortably
Arc B570 (10 GB)
best fit Q5_K_MFits comfortably
RTX 5070 (12 GB)
best fit Q8_0Fits comfortably
RTX 4070 Ti (12 GB)
best fit Q8_0Fits comfortably
RTX 4070 Super (12 GB)
best fit Q8_0Fits comfortably
RTX 4070 (12 GB)
best fit Q8_0Fits comfortably
RX 7700 XT (12 GB)
best fit Q8_0Fits comfortably
RTX 3060 (12 GB)
best fit Q8_0Fits comfortably
Arc B580 (12 GB)
best fit Q8_0Fits comfortably
RTX 5080 (16 GB)
best fit Q8_0Fits comfortably
RTX 5070 Ti (16 GB)
best fit Q8_0Fits comfortably
RTX 4080 Super (16 GB)
best fit Q8_0Fits comfortably
RTX 4080 (16 GB)
best fit Q8_0Fits comfortably
RX 9070 XT (16 GB)
best fit Q8_0Fits comfortably
RTX 4070 Ti Super (16 GB)
best fit Q8_0Fits comfortably
RX 9070 (16 GB)
best fit Q8_0Fits comfortably
RTX 5060 Ti (16 GB)
best fit Q8_0Fits comfortably
RTX 4090 Laptop (16 GB)
best fit Q8_0Fits comfortably
RX 7900 GRE (16 GB)
best fit Q8_0Fits comfortably
RX 7800 XT (16 GB)
best fit Q8_0Fits comfortably
RTX 4060 Ti 16GB (16 GB)
best fit Q8_0Fits comfortably
RTX A4000 (16 GB)
best fit Q8_0Fits comfortably
Arc A770 (16 GB)
best fit Q8_0Fits comfortably
RX 7900 XT (20 GB)
best fit FP16Fits comfortably
RTX 4090 (24 GB)
best fit FP16Fits comfortably
RX 7900 XTX (24 GB)
best fit FP16Fits comfortably
RTX 5090 (32 GB)
best fit FP16Fits comfortably
RTX 6000 Ada (48 GB)
best fit FP16Fits comfortably

Unified-memory machines — Apple Silicon Mac Studio/MacBook Pro (Max/Ultra chips) and AMD Strix Halo mini-PCs — share system RAM and GPU memory in one pool, so a 64–128 GB configuration can hold models that would otherwise need multiple discrete GPUs. Throughput is bandwidth-bound like any GPU; it isn’t modeled in the table above yet.

Recommended builds

Three ways to size a machine for Llama 3.1 8B, from the smallest GPU that fits to the no-compromise option with headroom for the next model up.

BudgetSmallest GPU that fits at Q4 — runs it, no headroom for bigger contexts.
BalancedOne tier up — comfortable context length, room to run alongside other apps.
No-compromiseFits at Q8 or FP16, or leaves room to step up to the next model class.

Tokens/sec expectations

GPUBandwidthEst. tok/s (Q4_K_M)
RTX 50901792 GB/s~255
RTX 40901008 GB/s~143.4
RTX 6000 Ada960 GB/s~136.6
RX 7900 XTX960 GB/s~136.6
RTX 5080960 GB/s~136.6
RTX 5070 Ti896 GB/s~127.5

Estimated from memory bandwidth ÷ active weight size — a ceiling, not a measured figure. Real throughput depends on the inference engine, batch size, and context length.

Build it, or rent the compute?

Owning the hardware is a one-time cost that keeps paying off the more you actually use it — every run after the first is free, the data never leaves your machine, and the GPU still does everything else a PC does (gaming, editing, the rest of your workload) between AI sessions. Renting cloud GPU time trades that upfront cost for pay-as-you-go pricing with no idle hardware to maintain, which wins for occasional, bursty, or much-larger-than-consumer-VRAM jobs. For a model that fits comfortably in the VRAM tiers above and gets used regularly, owning the GPU tends to be the cheaper path over time — for anything larger than a single card can hold, renting is often the more practical starting point.

Frequently asked questions

Can I run Llama 3.1 8B on an 8 GB GPU?

Yes — Q4_K_M weights are about 4.9 GB, leaving room for a few thousand tokens of context on an 8 GB card. For the full 128K context or Q8 quality you want 16 GB or system-RAM offload at reduced speed.

Related models:Stable Diffusion XL