Open dataset
Open-weight models, measured in the open.
One harness, one quantization class, hardware I own — every number reviewed in the llmfit dataset before it lands here.
- 410benchmark runs
- 237distinct models
- 6platforms, one harness
- 4-bitquantization
Hardware wanted
What does a unified-memory box actually deliver against discrete VRAM, on real models rather than a spec sheet? Lend me a unit for six weeks: same protocol, raw JSON published and submitted upstream. No editorial control, no embargo on a bad result. A number that can only come out flattering is worth nothing to your engineers either.
Currently after: unified-memory boxes · Ryzen AI Max+ 395 · Apple M silicon · Intel Arc
Method
Boring on purpose. Boring is what makes it comparable.
- One harness,
llmfit, two providers,llama.cppandmlx. - 4-bit everywhere: Q4_K_M GGUF, mlx-community 4bit on MLX.
- A warm-up pass before every measured run.
- Three runs per model, the average is published.
- Offload tuned until it fits, then held constant.
- Device confirmed by measuring VRAM, not by parsing a log.
© 2026 Akciali