<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Lizard inference engineering: B16 is the throughput lane, not a single-user promise]]></title><description><![CDATA[<p dir="auto"><strong>Inference Engineering · Day 11 · Evening</strong></p>
<p dir="auto"><img src="https://lizard-llm.qendryx.com/screenshots/Benchmark/Screenshot%202026-07-22%20221126.png" alt="B16 is the throughput lane, not a single-user promise editorial visual — lizard-llm.qendryx.com" class=" img-fluid img-markdown" /></p>
<p dir="auto">I think one thing that's easy to misunderstand in LLM benchmarks is tokens/sec.</p>
<p dir="auto">Our latest benchmark showed 40.7 tok/s for Lizard versus 21.8 tok/s for llama.cpp, using the exact same Llama 3.2 3B Q4_K_M GGUF model.</p>
<p dir="auto">The important part is that this is aggregate throughput with 16 concurrent requests (B=16) over multiple real HTTP runs. It's not the speed of a single chat response.</p>
<p dir="auto">Per-user latency and aggregate throughput answer different questions, and I think we should be much clearer about which one we're talking about when comparing inference engines.</p>
<p dir="auto">Whenever I see a tokens/sec number now, my first question is: Is that per response, or total throughput?</p>
<p dir="auto"><strong>Engineering fact:</strong> On one verified Llama 3.2 3B Q4_K_M run series, Lizard measured a 40.736 tok/s median at B=16 versus 21.755 tok/s for llama.cpp using the same physical GGUF; the result is aggregate HTTP throughput on that test system.</p>
<p dir="auto"><a href="https://lizard-llm.qendryx.com/benchmarks.html" rel="nofollow ugc">Read the relevant Lizard page</a></p>
<p dir="auto">#LizardLLM #LLMInference #Benchmarking #LocalAI #PerformanceEngineering</p>
<p dir="auto">&lt;!-- lizard-marketing-slot:day-11-pm --&gt;</p>
]]></description><link>https://community.lizard-llm.qendryx.com/topic/61/lizard-inference-engineering-b16-is-the-throughput-lane-not-a-single-user-promise</link><generator>RSS for Node</generator><lastBuildDate>Mon, 24 Aug 2026 01:37:37 GMT</lastBuildDate><atom:link href="https://community.lizard-llm.qendryx.com/topic/61.rss" rel="self" type="application/rss+xml"/><pubDate>Mon, 03 Aug 2026 11:00:08 GMT</pubDate><ttl>60</ttl></channel></rss>