Watch a Native, Caterpillar and llama.cpp benchmark run
-
This sequence follows a provider comparison from a failed preflight to a successful run. It includes the exact command, warm-up state, partial progress and final rows, because an industrial benchmark should be reproducible and should not hide failure evidence.
Interactive gallery: https://lizard-llm.qendryx.com/benchmarks.html
A missing Ollama model stops before measurement

The attempted triple-provider run exits because no matching Ollama model is installed. Lizard records the reason and exit code instead of inventing an Ollama result.
The successful provider run begins

The rerun shows the exact providers, quantization, prompt budget and warm-server reuse while Lizard Native starts its first row.
Running and completed rows stay distinct

Mid-run, Lizard Native Q4 measures 12.8 tok/s and Caterpillar Q4 8.23 while llama.cpp is still warming. A running lane is not presented as finished.
The finished head-to-head

For this Llama 3.2 Q4 run, stock llama.cpp reaches 16.31 tok/s, Lizard Native 12.8, and Caterpillar 8.23. Setup and warm-up fields explain the wall-time difference.
Repeatable recipes instead of hidden presets

Single-pass, full-sweep and triple-stack recipes are visible and editable. The command center makes the intended comparison explicit before it consumes a run.
Question for you: Should the next public run include an installed Ollama baseline, and if so which exact Ollama model tag should we use?
Hello! It looks like you're interested in this conversation, but you don't have an account yet.
Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.
With your input, this post could be even better 💗
Register Login