<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[New benchmark evidence: completed provider comparisons]]></title><description><![CDATA[<p dir="auto">These four captures show why the dashboard keeps <strong>provider rows</strong> and <strong>analysis</strong> separate. The same machine can produce a different winner in a different run, and a composite recommendation is never allowed to erase the underlying values.</p>
<p dir="auto">Interactive gallery: <a href="https://lizard-llm.qendryx.com/benchmarks.html" rel="nofollow ugc">https://lizard-llm.qendryx.com/benchmarks.html</a></p>
<h2>Five rows, five completed measurements</h2>
<p dir="auto"><img src="https://community.lizard-llm.qendryx.com/assets/uploads/benchmark-evidence-2026-07-22/Screenshot%202026-07-22%20213232.png" alt="Completed five-row provider comparison" class=" img-fluid img-markdown" /></p>
<p dir="auto">Ollama records 8.56 tok/s, stock llama.cpp Q4 7.85, and Lizard Native Q4 5.61 in this run. Load, warm-up and memory evidence stay visible beside throughput.</p>
<h2>A second run changes the order</h2>
<p dir="auto"><img src="https://community.lizard-llm.qendryx.com/assets/uploads/benchmark-evidence-2026-07-22/Screenshot%202026-07-22%20213249.png" alt="Expanded ten-row provider comparison" class=" img-fluid img-markdown" /></p>
<p dir="auto">In this separate 10/10 run, llama.cpp Q4 reaches 6.23 tok/s, Lizard Native Q4 5.28, and Ollama 4.53. The result belongs to this model, machine and run—not a universal ranking.</p>
<h2>Leaderboard plus metric-specific winners</h2>
<p dir="auto"><img src="https://community.lizard-llm.qendryx.com/assets/uploads/benchmark-evidence-2026-07-22/Screenshot%202026-07-22%20213314.png" alt="Provider leaderboard and insight panel" class=" img-fluid img-markdown" /></p>
<p dir="auto">Ollama leads the composite score, Lizard Native Q4 takes best native decode and parity among the compared rows, and llama.cpp Q4 leads wall time and end-to-end tokens per second.</p>
<h2>Recommendations stay attached to source evidence</h2>
<p dir="auto"><img src="https://community.lizard-llm.qendryx.com/assets/uploads/benchmark-evidence-2026-07-22/Screenshot%202026-07-22%20213327.png" alt="Wide benchmark leaderboard and evidence panel" class=" img-fluid img-markdown" /></p>
<p dir="auto">Scores, badges, execution lanes and raw provider values remain on the same screen. The recommendation annotates the run; it does not replace it.</p>
<hr />
<p dir="auto"><strong>Question for you:</strong> On your hardware, which decision should come first: shortest wall time, steady decode speed, peak memory, or answer parity?</p>
]]></description><link>https://community.lizard-llm.qendryx.com/topic/26/new-benchmark-evidence-completed-provider-comparisons</link><generator>RSS for Node</generator><lastBuildDate>Mon, 24 Aug 2026 03:20:07 GMT</lastBuildDate><atom:link href="https://community.lizard-llm.qendryx.com/topic/26.rss" rel="self" type="application/rss+xml"/><pubDate>Wed, 22 Jul 2026 15:56:46 GMT</pubDate><ttl>60</ttl></channel></rss>