Open dataset

Open-weight models, measured in the open.

No vendor slides, no peak numbers on hardware nobody has: one harness, one quantization, GPUs I actually own. Every raw number goes upstream to the llmfit community dataset and is reviewed before it lands here, so you can diff this page against the source.

Reviewed upstream llmfit 9 PRs merged #820 open

    Hardware wanted

    What does a unified-memory box actually deliver against discrete VRAM, on real models rather than a spec sheet? Lend me a unit for six weeks: same protocol, raw JSON published and submitted upstream. No editorial control, no embargo on a bad result. A number that can only come out flattering is worth nothing to your engineers either.

    Currently after: unified-memory boxes · GB10 class · Ryzen AI Max+ 395 · Apple M silicon · Intel Arc

    Mail me · LinkedIn

    Method

    Boring on purpose. Boring is what makes it comparable.

    1. One harness, llmfit 1.1.3, one provider, llama.cpp.
    2. Q4_K_M GGUF everywhere, same builds.
    3. A warm-up pass before every measured run.
    4. Three runs per model, the average is published.
    5. Offload tuned until it fits, then held constant.
    6. Device confirmed by measuring VRAM, not by parsing a log.