Open dataset
Open-weight models, measured in the open.
No vendor slides, no peak numbers on hardware nobody has: one harness, one quantization, GPUs I actually own. Every raw number goes upstream to the llmfit community dataset and is reviewed before it lands here, so you can diff this page against the source.
- 189benchmark runs
- 80distinct models
- 4platforms, one harness
- Q4_K_Mquantization
Hardware wanted
What does a unified-memory box actually deliver against discrete VRAM, on real models rather than a spec sheet? Lend me a unit for six weeks: same protocol, raw JSON published and submitted upstream. No editorial control, no embargo on a bad result. A number that can only come out flattering is worth nothing to your engineers either.
Currently after: unified-memory boxes · GB10 class · Ryzen AI Max+ 395 · Apple M silicon · Intel Arc
Method
Boring on purpose. Boring is what makes it comparable.
- One harness,
llmfit 1.1.3, one provider,llama.cpp. - Q4_K_M GGUF everywhere, same builds.
- A warm-up pass before every measured run.
- Three runs per model, the average is published.
- Offload tuned until it fits, then held constant.
- Device confirmed by measuring VRAM, not by parsing a log.