Explore MemVanta's canonical OpenLLaMA 7B v2 Q4_0 benchmark against a pinned llama.cpp revision. The page is generated from committed benchmark evidence before publication.
| Memory | MemVanta measured 47.54% lower peak RSS. |
|---|---|
| Prompt speed | llama.cpp was 4.10× faster in prompt processing. |
| Generation speed | llama.cpp was 4.14× faster in token generation. |
| Scope | Specific model, host, workload and pinned runtime revisions. |
| Benchmark | OpenLLaMA 7B v2 Q4_0 same-model CPU A/B |
|---|---|
| Model file | openlm-research-open_llama_7b_v2-Q4_0.gguf |
| Model SHA-256 | 892f6e2e840ed98bb7bc1d74da67ae0cca52b2d90b0316d0a7e66a37a5a60760 |
| llama.cpp commit | 6503355df0eb4f65875012523263c302fe0088c1 |
| MemVanta source commit | 2dffdde24bfa410eb7d4b2256c3c20731c825b19 |
| Source artifact | results/openllama-7b-v2-ab/summary.json |