Read this first: what the benchmark actually measures
-
The benchmark runs your model on your hardware and reports what happened. A few principles it sticks to:
- Measured means measured. Every number is labelled —
Measured,Estimated,Mixed,Partial. If a run didn't finish, you getIncomplete, not a filled-in guess. - Nothing is invented. There are no synthetic values anywhere. A missing number shows as
—. - Derived numbers inherit the weakest input. Measured speed divided by an estimated memory figure is labelled
Mixed, neverMeasured.
Free readiness check vs. full benchmark: the readiness check in setup is free and always available — it tells you tokens/sec for your model on this machine. The full suite (Lizard vs Ollama vs llama.cpp across Q1–Q9) is the credit-based one.
Full explanation: https://lizard-llm.qendryx.com/benchmarks.html
When posting results, your CPU/GPU and model + quantization make them useful to everyone else.
- Measured means measured. Every number is labelled —
Hello! It looks like you're interested in this conversation, but you don't have an account yet.
Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.
With your input, this post could be even better 💗
Register Login