<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[From GGUF selection to local inference: the benchmark setup]]></title><description><![CDATA[<p dir="auto">A benchmark is only meaningful when the exact model file and runtime state are known. These five screens show the path from repository metadata to a locally activated model. Planning estimates are shown before the run; measured evidence is recorded later.</p>
<p dir="auto">Interactive gallery: <a href="https://lizard-llm.qendryx.com/benchmarks.html" rel="nofollow ugc">https://lizard-llm.qendryx.com/benchmarks.html</a></p>
<h2>Inspect repository and compatibility metadata</h2>
<p dir="auto"><img src="https://community.lizard-llm.qendryx.com/assets/uploads/benchmark-evidence-2026-07-22/Screenshot%202026-07-22%20213514.png" alt="Llama 3.2 GGUF model details" class=" img-fluid img-markdown" /></p>
<p dir="auto">Lizard shows the model family, author, supported quantizations and whether the file can run on the native engine before anything is downloaded.</p>
<h2>Choose the exact quantized file</h2>
<p dir="auto"><img src="https://community.lizard-llm.qendryx.com/assets/uploads/benchmark-evidence-2026-07-22/Screenshot%202026-07-22%20213538.png" alt="GGUF quantization variants" class=" img-fluid img-markdown" /></p>
<p dir="auto">The selected Q4_K_M file has explicit download size, estimated RAM and planning speed. Those estimates are not relabeled as measurements later.</p>
<h2>Download identity remains visible</h2>
<p dir="auto"><img src="https://community.lizard-llm.qendryx.com/assets/uploads/benchmark-evidence-2026-07-22/Screenshot%202026-07-22%20213551.png" alt="Selected GGUF download progress" class=" img-fluid img-markdown" /></p>
<p dir="auto">The repository and exact GGUF filename remain on screen while the file downloads, so the later provider rows can be traced back to their intended weights.</p>
<h2>Cold is different from warm</h2>
<p dir="auto"><img src="https://community.lizard-llm.qendryx.com/assets/uploads/benchmark-evidence-2026-07-22/Screenshot%202026-07-22%20213620.png" alt="Cold local chat model state" class=" img-fluid img-markdown" /></p>
<p dir="auto">The chat identifies that the model is cold and offers an explicit wake action. Benchmark load and inference timings preserve the same distinction.</p>
<h2>Activate the model for local inference</h2>
<p dir="auto"><img src="https://community.lizard-llm.qendryx.com/assets/uploads/benchmark-evidence-2026-07-22/Screenshot%202026-07-22%20213650.png" alt="Locally activated Llama 3.2 model" class=" img-fluid img-markdown" /></p>
<p dir="auto">The activated model becomes available to the CLI, chat and local API. The screen names Lizard Native and Caterpillar as the two on-device engines.</p>
<hr />
<p dir="auto"><strong>Question for you:</strong> Which model and quantization did Lizard recommend for your machine, and did the measured result match the planning estimate?</p>
]]></description><link>https://community.lizard-llm.qendryx.com/topic/27/from-gguf-selection-to-local-inference-the-benchmark-setup</link><generator>RSS for Node</generator><lastBuildDate>Mon, 24 Aug 2026 03:19:27 GMT</lastBuildDate><atom:link href="https://community.lizard-llm.qendryx.com/topic/27.rss" rel="self" type="application/rss+xml"/><pubDate>Wed, 22 Jul 2026 15:57:46 GMT</pubDate><ttl>60</ttl></channel></rss>