Back to oracle

Mimir Forge / Memory-budget estimator

Can my machine run this?

Enter your GPU, Mac, RAM, context length, and use case. ToolHalla estimates which local models fit — and when cloud GPU is the smarter call.

Verdict / Llama-3.1-8B-Instruct / Chat / RAG

Apple M4 / 24 GB Likely fits comfortably.

LLM directory data / 11.2 GB model memory at Q8_0

Memory estimate: likely fits at Q8_0 with roomy memory pressure at 8k context. Speed estimate: benchmark needed.

Unified memory systems are estimates, not direct VRAM matches.

Q8_0
Recommended quantization
Verdict
Likely
Recommended quantization
Q8_0
Expected speed estimate
Benchmark needed
Memory pressure
Roomy

Memory budget breakdown

01
Weights
8B @ Q8_0
11.2 GB
02
KV cache reserve
8k context / solo concurrency
0.2 GB
03
Headroom for OS + activations
If this drops below about 10%, expect swapping or OOM
6.6 GB

Lighter local alternative

If you want faster/lower-power local inference, consider smaller models.

Fit
Phi-3.5-mini
3.8B class / Q4 fits about 6 GB
3B class

When cloud is smarter

Cloud is usually smarter when you need long context, heavy concurrency, fast experiments, or high-memory models without buying hardware.

01
A100 80GB rental
Good fit for 70B class and longer context
$/hr sample

If you want to upgrade

01
RTX 3090 used
24 GB VRAM class, strong local AI value
used market
02
RTX 4090
24 GB VRAM class, fast consumer card
new/used
Confidence: medium — estimate based on memory requirements, not a live benchmark.

Estimates vary by runtime, quantization, context length, OS overhead, and backend.