<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Watch a Native, Caterpillar and llama.cpp benchmark run]]></title><description><![CDATA[<p dir="auto">This sequence follows a provider comparison from a failed preflight to a successful run. It includes the exact command, warm-up state, partial progress and final rows, because an industrial benchmark should be reproducible and should not hide failure evidence.</p>
<p dir="auto">Interactive gallery: <a href="https://lizard-llm.qendryx.com/benchmarks.html" rel="nofollow ugc">https://lizard-llm.qendryx.com/benchmarks.html</a></p>
<h2>A missing Ollama model stops before measurement</h2>
<p dir="auto"><img src="https://community.lizard-llm.qendryx.com/assets/uploads/benchmark-evidence-2026-07-22/Screenshot%202026-07-22%20213816.png" alt="Failed Ollama benchmark preflight" class=" img-fluid img-markdown" /></p>
<p dir="auto">The attempted triple-provider run exits because no matching Ollama model is installed. Lizard records the reason and exit code instead of inventing an Ollama result.</p>
<h2>The successful provider run begins</h2>
<p dir="auto"><img src="https://community.lizard-llm.qendryx.com/assets/uploads/benchmark-evidence-2026-07-22/Screenshot%202026-07-22%20213915.png" alt="Live Native Caterpillar and llama.cpp command" class=" img-fluid img-markdown" /></p>
<p dir="auto">The rerun shows the exact providers, quantization, prompt budget and warm-server reuse while Lizard Native starts its first row.</p>
<h2>Running and completed rows stay distinct</h2>
<p dir="auto"><img src="https://community.lizard-llm.qendryx.com/assets/uploads/benchmark-evidence-2026-07-22/Screenshot%202026-07-22%20213952.png" alt="Live provider cards while llama.cpp warms" class=" img-fluid img-markdown" /></p>
<p dir="auto">Mid-run, Lizard Native Q4 measures 12.8 tok/s and Caterpillar Q4 8.23 while llama.cpp is still warming. A running lane is not presented as finished.</p>
<h2>The finished head-to-head</h2>
<p dir="auto"><img src="https://community.lizard-llm.qendryx.com/assets/uploads/benchmark-evidence-2026-07-22/Screenshot%202026-07-22%20214012.png" alt="Finished llama.cpp Native and Caterpillar comparison" class=" img-fluid img-markdown" /></p>
<p dir="auto">For this Llama 3.2 Q4 run, stock llama.cpp reaches 16.31 tok/s, Lizard Native 12.8, and Caterpillar 8.23. Setup and warm-up fields explain the wall-time difference.</p>
<h2>Repeatable recipes instead of hidden presets</h2>
<p dir="auto"><img src="https://community.lizard-llm.qendryx.com/assets/uploads/benchmark-evidence-2026-07-22/Screenshot%202026-07-22%20214040.png" alt="Benchmark command recipe library" class=" img-fluid img-markdown" /></p>
<p dir="auto">Single-pass, full-sweep and triple-stack recipes are visible and editable. The command center makes the intended comparison explicit before it consumes a run.</p>
<hr />
<p dir="auto"><strong>Question for you:</strong> Should the next public run include an installed Ollama baseline, and if so which exact Ollama model tag should we use?</p>
]]></description><link>https://community.lizard-llm.qendryx.com/topic/28/watch-a-native-caterpillar-and-llama.cpp-benchmark-run</link><generator>RSS for Node</generator><lastBuildDate>Mon, 24 Aug 2026 01:37:04 GMT</lastBuildDate><atom:link href="https://community.lizard-llm.qendryx.com/topic/28.rss" rel="self" type="application/rss+xml"/><pubDate>Wed, 22 Jul 2026 15:58:46 GMT</pubDate><ttl>60</ttl></channel></rss>