Skip to content
  • Categories
  • Recent
  • Tags
  • Popular
  • World
  • Users
  • Groups
Skins
  • Light
  • Brite
  • Cerulean
  • Cosmo
  • Flatly
  • Journal
  • Litera
  • Lumen
  • Lux
  • Materia
  • Minty
  • Morph
  • Pulse
  • Sandstone
  • Simplex
  • Sketchy
  • Spacelab
  • United
  • Yeti
  • Zephyr
  • Dark
  • Cyborg
  • Darkly
  • Quartz
  • Slate
  • Solar
  • Superhero
  • Vapor

  • Default (No Skin)
  • No Skin
Collapse

Lizard-LLM Community

  1. Home
  2. Lizard Native
  3. Lizard inference engineering: Stop rebuilding descriptors per token

Lizard inference engineering: Stop rebuilding descriptors per token

Scheduled Pinned Locked Moved Lizard Native
1 Posts 1 Posters 0 Views 1 Watching
  • Oldest to Newest
  • Newest to Oldest
  • Most Votes
Reply
  • Reply as topic
Log in to reply
This topic has been deleted. Only users with topic management privileges can see it.
  • L Offline
    L Offline
    lizardadmin
    wrote last edited by lizardadmin
    #1

    Inference Engineering · Day 6 · Morning

    Gemma 4 provider comparison — lizard-llm.qendryx.com

    A decode loop can look GPU-bound and still waste time on the host.

    In the lizard-native Gemma path, roughly 400 weight descriptors were being rebuilt for every generated token. That meant the runtime kept reconstructing the same per-token metadata before the next step could move forward. The GPU had work to do, but so did the scaffolding around it, and that scaffolding sat on the critical path.

    The fix was to stop rebuilding those descriptors and keep them around for reuse. On the Gemma test runs, that improved decode by 13–26% and brought first-token response back by about 20%.

    The useful lesson is that decode tuning is rarely only about the kernel. If the loop repeats cleanly, inspect the state that gets recreated every pass: descriptor setup, graph metadata, plan assembly, session bookkeeping, cache wiring. Those are easy to ignore because they do not look like compute, but they can still decide how fast the next token appears.

    I have learned to ask a boring question before touching the math: what is being rebuilt on every token that could have been carried forward from the previous one?

    In practice, persistent per-token state often buys a cleaner win than shaving a few cycles off the math itself. Less allocation, less synchronization pressure, fewer host-side stalls, and a shorter path back into the decode loop. When the hot path is stable, that kind of cleanup tends to show up immediately in the first-token feel as well as steady-state throughput.

    Engineering fact: The lizard-native Gemma path stopped rebuilding roughly 400 weight descriptors per token; test runs improved decode 13–26% and first-token response about 20%.

    Lizard The AI Runtime You'll Own—Not Rent.

    Receive two professional Windows AI runtimes with lifetime updates. Run AI at native speed, keep every conversation private, and stay independent with intelligent hardware optimization and no cloud dependency.

    Read the relevant Lizard page

    #localai #inferencetuning #decodeperformance #windowsai #d3d12

    <!-- lizard-marketing-slot:day-06-am -->

    1 Reply Last reply
    0

    Hello! It looks like you're interested in this conversation, but you don't have an account yet.

    Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.

    With your input, this post could be even better 💗

    Register Login
    Reply
    • Reply as topic
    Log in to reply
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes


    • Login

    • Don't have an account? Register

    • Login or register to search.
    Powered by NodeBB Contributors
    • First post
      Last post
    0
    • Categories
    • Recent
    • Tags
    • Popular
    • World
    • Users
    • Groups