<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Lizard inference engineering: An inference provider owns the stack]]></title><description><![CDATA[<p dir="auto"><strong>Inference Engineering · Day 1 · Morning</strong></p>
<p dir="auto">An API endpoint is not an inference engine.</p>
<p dir="auto">The engineering decisions that determine latency happen below the route: batching, KV-cache precision, quantized kernels, memory residency and device submission.</p>
<p dir="auto">Lizard treats the local machine as the inference provider. lizard-native owns the Direct3D 12 path; Caterpillar owns its compiled execution plan. The OpenAI-compatible endpoint is only the top layer.</p>
<p dir="auto">Which layer of your current inference stack can you actually inspect and change?</p>
<p dir="auto"><strong>Engineering fact:</strong> Lizard owns the local endpoint, scheduler, cache layout, kernels, quantization choices, and device execution path.</p>
<p dir="auto"><a href="https://lizard-llm.qendryx.com/product.html" rel="nofollow ugc">Read the relevant Lizard page</a></p>
<p dir="auto">#InferenceEngineering #LocalAI #OnDeviceAI #LizardLLM #LLMEngineering</p>
<p dir="auto">&lt;!-- lizard-marketing-slot:day-01-am --&gt;</p>
]]></description><link>https://community.lizard-llm.qendryx.com/topic/33/lizard-inference-engineering-an-inference-provider-owns-the-stack</link><generator>RSS for Node</generator><lastBuildDate>Mon, 24 Aug 2026 01:36:16 GMT</lastBuildDate><atom:link href="https://community.lizard-llm.qendryx.com/topic/33.rss" rel="self" type="application/rss+xml"/><pubDate>Fri, 24 Jul 2026 01:00:10 GMT</pubDate><ttl>60</ttl></channel></rss>