<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[What the latest Native and Caterpillar results actually show]]></title><description><![CDATA[<p dir="auto">These analyzed captures put <strong>Lizard Native</strong> and <strong>Caterpillar</strong> in the foreground beside stock <strong>llama.cpp</strong> and optional <strong>Ollama</strong>. They also show the honest result: no provider wins every metric, and Caterpillar trails the other two lanes in these selected general-case runs.</p>
<p dir="auto">Interactive gallery: <a href="https://lizard-llm.qendryx.com/benchmarks.html" rel="nofollow ugc">https://lizard-llm.qendryx.com/benchmarks.html</a></p>
<h2>Wall time, throughput and decode are different metrics</h2>
<p dir="auto"><img src="https://community.lizard-llm.qendryx.com/assets/uploads/benchmark-evidence-2026-07-22/Screenshot%202026-07-22%20221126.png" alt="Runtime throughput and decode charts" class=" img-fluid img-markdown" /></p>
<p dir="auto">llama.cpp wins this run's wall time and end-to-end throughput. Lizard Native leads native decode at 15.198 tok/s versus Caterpillar's 9.201. The dashboard does not blend those numbers.</p>
<h2>No provider wins every dimension</h2>
<p dir="auto"><img src="https://community.lizard-llm.qendryx.com/assets/uploads/benchmark-evidence-2026-07-22/Screenshot%202026-07-22%20221157.png" alt="Provider heatmap radar and history charts" class=" img-fluid img-markdown" /></p>
<p dir="auto">The heatmap and radar show the trade: llama.cpp leads speed and memory efficiency here, while the native lanes carry their own decode and execution-lane evidence.</p>
<h2>The decision guide names the real winners</h2>
<p dir="auto"><img src="https://community.lizard-llm.qendryx.com/assets/uploads/benchmark-evidence-2026-07-22/Screenshot%202026-07-22%20221220.png" alt="Measured practical decision guide" class=" img-fluid img-markdown" /></p>
<p dir="auto">For this hardware, llama.cpp Q4 is the balanced choice at 16.31 tok/s and 3.925 seconds. Lizard Native Q4 is called out separately for best native decode at 15.198 tok/s. Ollama was not included and remains unmeasured.</p>
<h2>A larger model changes the gap</h2>
<p dir="auto"><img src="https://community.lizard-llm.qendryx.com/assets/uploads/benchmark-evidence-2026-07-22/Screenshot%202026-07-22%20222539.png" alt="Gemma 4 provider comparison" class=" img-fluid img-markdown" /></p>
<p dir="auto">For Gemma 4 E4B, llama.cpp records 4.37 tok/s, Lizard Native 4.08, and Caterpillar 1.41. The Ollama card says not included and unmeasured—not zero.</p>
<h2>Provider identity stays attached to every metric</h2>
<p dir="auto"><img src="https://community.lizard-llm.qendryx.com/assets/uploads/benchmark-evidence-2026-07-22/Screenshot%202026-07-22%20222551.png" alt="Full-width Gemma 4 provider cards" class=" img-fluid img-markdown" /></p>
<p dir="auto">Runtime, model load, warm-up, peak memory and efficiency stay under the provider that produced them. This is the evidence needed to evaluate Lizard as a local inference provider.</p>
<hr />
<p dir="auto"><strong>Question for you:</strong> What hardware and model should we run next to test where Caterpillar closes the gap—or where Native's decode path matters most?</p>
]]></description><link>https://community.lizard-llm.qendryx.com/topic/29/what-the-latest-native-and-caterpillar-results-actually-show</link><generator>RSS for Node</generator><lastBuildDate>Mon, 24 Aug 2026 01:37:24 GMT</lastBuildDate><atom:link href="https://community.lizard-llm.qendryx.com/topic/29.rss" rel="self" type="application/rss+xml"/><pubDate>Wed, 22 Jul 2026 15:59:47 GMT</pubDate><ttl>60</ttl></channel></rss>