<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Lizard inference engineering: Budget against free VRAM, not the sticker]]></title><description><![CDATA[<p dir="auto"><strong>Inference Engineering · Day 2 · Morning</strong></p>
<p dir="auto">The number on the GPU box is not your inference budget.</p>
<p dir="auto">Windows, the display pipeline, browsers and other applications already occupy graphics memory. Planning against total VRAM can turn a model that looks safe into a load failure or an eviction loop.</p>
<p dir="auto">Lizard asks the driver what is available now, then budgets model weights and runtime needs against that live figure. If the fit is unsafe, it chooses a supported path and reports why.</p>
<p dir="auto">Capacity planning should begin with free memory, not advertised memory.</p>
<p dir="auto"><strong>Engineering fact:</strong> Lizard checks the graphics driver's currently available memory rather than assuming the GPU's advertised total is free.</p>
<p dir="auto"><a href="https://lizard-llm.qendryx.com/product.html" rel="nofollow ugc">Read the relevant Lizard page</a></p>
<p dir="auto">#VRAM #GPUEngineering #LocalAI #LizardLLM #InferenceEngineering</p>
<p dir="auto">&lt;!-- lizard-marketing-slot:day-02-am --&gt;</p>
]]></description><link>https://community.lizard-llm.qendryx.com/topic/37/lizard-inference-engineering-budget-against-free-vram-not-the-sticker</link><generator>RSS for Node</generator><lastBuildDate>Mon, 24 Aug 2026 01:37:22 GMT</lastBuildDate><atom:link href="https://community.lizard-llm.qendryx.com/topic/37.rss" rel="self" type="application/rss+xml"/><pubDate>Sat, 25 Jul 2026 01:00:11 GMT</pubDate><ttl>60</ttl></channel></rss>