Open dataset

Open-weight models, measured in the open.

One harness, one quantization class, hardware I own — every number reviewed in the llmfit dataset before it lands here.

Reviewed upstream llmfit 23 PRs merged #1047 open

    Hardware wanted

    What does a unified-memory box actually deliver against discrete VRAM, on real models rather than a spec sheet? Lend me a unit for six weeks: same protocol, raw JSON published and submitted upstream. No editorial control, no embargo on a bad result. A number that can only come out flattering is worth nothing to your engineers either.

    Currently after: unified-memory boxes · Ryzen AI Max+ 395 · Apple M silicon · Intel Arc

    Method

    Boring on purpose. Boring is what makes it comparable.

    1. One harness, llmfit, two providers, llama.cpp and mlx.
    2. 4-bit everywhere: Q4_K_M GGUF, mlx-community 4bit on MLX.
    3. A warm-up pass before every measured run.
    4. Three runs per model, the average is published.
    5. Offload tuned until it fits, then held constant.
    6. Device confirmed by measuring VRAM, not by parsing a log.
    © 2026 Akciali