MemVanta

CPU LLM memory benchmark explorer
same-model CPU A/B benchmark

Less resident memory.
Visible throughput trade-off.

Explore MemVanta's canonical OpenLLaMA 7B v2 Q4_0 benchmark against a pinned llama.cpp revision. The page is generated from committed benchmark evidence before publication.

Measured peak-RSS reduction
47.54%
for this benchmark only
MemVanta peak RSS
3.80 GiB
measured process peak resident set
llama.cpp peak RSS
7.24 GiB
same GGUF artifact, pinned comparison runtime
MemVanta prompt
2.92 tok/s
tokens / second
MemVanta generation
1.92 tok/s
tokens / second

Memory footprint

MemVanta
3.80 GiB
llama.cpp
7.24 GiB

Throughput comparison

Prompt processing
MemVanta
2.92
llama.cpp
11.95
Token generation
MemVanta
1.92
llama.cpp
7.97

Measured-memory budget explorer

Set a hypothetical RAM budget and compare it with the measured peak RSS values from this run.
Budget: 4.0 GiB

What the benchmark says

MemoryMemVanta measured 47.54% lower peak RSS.
Prompt speedllama.cpp was 4.10× faster in prompt processing.
Generation speedllama.cpp was 4.14× faster in token generation.
ScopeSpecific model, host, workload and pinned runtime revisions.

Reproducibility evidence

BenchmarkOpenLLaMA 7B v2 Q4_0 same-model CPU A/B
Model fileopenlm-research-open_llama_7b_v2-Q4_0.gguf
Model SHA-256892f6e2e840ed98bb7bc1d74da67ae0cca52b2d90b0316d0a7e66a37a5a60760
llama.cpp commit6503355df0eb4f65875012523263c302fe0088c1
MemVanta source commit2dffdde24bfa410eb7d4b2256c3c20731c825b19
Source artifactresults/openllama-7b-v2-ab/summary.json
This is an explanatory static Space. It does not run a 7B model in the browser. Lower measured RSS does not imply faster inference, and these measurements should not be generalized beyond the tested configuration without reproduction.