Skip to content
  • Categories
  • Recent
  • Tags
  • Popular
  • World
  • Users
  • Groups
Skins
  • Light
  • Brite
  • Cerulean
  • Cosmo
  • Flatly
  • Journal
  • Litera
  • Lumen
  • Lux
  • Materia
  • Minty
  • Morph
  • Pulse
  • Sandstone
  • Simplex
  • Sketchy
  • Spacelab
  • United
  • Yeti
  • Zephyr
  • Dark
  • Cyborg
  • Darkly
  • Quartz
  • Slate
  • Solar
  • Superhero
  • Vapor

  • Default (No Skin)
  • No Skin
Collapse

Lizard-LLM Community

  1. Home
  2. Lizard Native
  3. Technical overview: how Lizard Native keeps Direct3D 12 inference resident

Technical overview: how Lizard Native keeps Direct3D 12 inference resident

Scheduled Pinned Locked Moved Lizard Native
1 Posts 1 Posters 3 Views 1 Watching
  • Oldest to Newest
  • Newest to Oldest
  • Most Votes
Reply
  • Reply as topic
Log in to reply
This topic has been deleted. Only users with topic management privileges can see it.
  • L Offline
    L Offline
    lizardadmin
    wrote on last edited by
    #1

    Lizard Native is Lizard's stable, Windows-first native provider. For a supported model, it uploads quantized tensors to GPU-resident buffers and executes decode through Direct3D 12 compute shaders. The model-fit decision happens before execution and uses the exact GGUF identity plus the machine's available memory.

    Wireframe Lizard Native resident-weight architecture

    The layers

    1. Hardware and model-fit planning — CPU, memory, Direct3D 12 capabilities, model family, quantization, and live budget are evaluated before a native lane is selected.
    2. Resident model tensors — supported quantized weights stay in device buffers while the model is active instead of being reloaded for every request.
    3. Compute shaders — quantized operations run through Direct3D 12 compute; compiled shader artifacts are cached on disk.
    4. Memory discipline — residency is budgeted, and compatible UMA hardware can use a shared-memory path.
    5. Provider identity — a native row stays labeled Lizard Native. A llama.cpp or Ollama fallback is not reported as native.

    The boundary matters

    Lizard Native implements a tested native subset; it does not claim that every GGUF architecture and quantization runs on this provider. Unsupported combinations stay eligible for an explicit bundled llama.cpp or optional Ollama fallback.

    Read the complete layer map and comparison: https://lizard-llm.qendryx.com/technical-overview.html

    Question: On your Windows hardware, is the limiting factor available memory, shader execution, model coverage, or per-token scheduling?

    1 Reply Last reply
    0

    Hello! It looks like you're interested in this conversation, but you don't have an account yet.

    Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.

    With your input, this post could be even better 💗

    Register Login
    Reply
    • Reply as topic
    Log in to reply
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes


    • Login

    • Don't have an account? Register

    • Login or register to search.
    Powered by NodeBB Contributors
    • First post
      Last post
    0
    • Categories
    • Recent
    • Tags
    • Popular
    • World
    • Users
    • Groups