Skip to content
  • Categories
  • Recent
  • Tags
  • Popular
  • World
  • Users
  • Groups
Skins
  • Light
  • Brite
  • Cerulean
  • Cosmo
  • Flatly
  • Journal
  • Litera
  • Lumen
  • Lux
  • Materia
  • Minty
  • Morph
  • Pulse
  • Sandstone
  • Simplex
  • Sketchy
  • Spacelab
  • United
  • Yeti
  • Zephyr
  • Dark
  • Cyborg
  • Darkly
  • Quartz
  • Slate
  • Solar
  • Superhero
  • Vapor

  • Default (No Skin)
  • No Skin
Collapse

Lizard-LLM Community

  1. Home
  2. Benchmarks
  3. Lizard inference engineering: B16 is the throughput lane, not a single-user promise

Lizard inference engineering: B16 is the throughput lane, not a single-user promise

Scheduled Pinned Locked Moved Benchmarks
1 Posts 1 Posters 0 Views 1 Watching
  • Oldest to Newest
  • Newest to Oldest
  • Most Votes
Reply
  • Reply as topic
Log in to reply
This topic has been deleted. Only users with topic management privileges can see it.
  • L Offline
    L Offline
    lizardadmin
    wrote last edited by
    #1

    Inference Engineering · Day 11 · Evening

    B16 is the throughput lane, not a single-user promise editorial visual — lizard-llm.qendryx.com

    I think one thing that's easy to misunderstand in LLM benchmarks is tokens/sec.

    Our latest benchmark showed 40.7 tok/s for Lizard versus 21.8 tok/s for llama.cpp, using the exact same Llama 3.2 3B Q4_K_M GGUF model.

    The important part is that this is aggregate throughput with 16 concurrent requests (B=16) over multiple real HTTP runs. It's not the speed of a single chat response.

    Per-user latency and aggregate throughput answer different questions, and I think we should be much clearer about which one we're talking about when comparing inference engines.

    Whenever I see a tokens/sec number now, my first question is: Is that per response, or total throughput?

    Engineering fact: On one verified Llama 3.2 3B Q4_K_M run series, Lizard measured a 40.736 tok/s median at B=16 versus 21.755 tok/s for llama.cpp using the same physical GGUF; the result is aggregate HTTP throughput on that test system.

    Read the relevant Lizard page

    #LizardLLM #LLMInference #Benchmarking #LocalAI #PerformanceEngineering

    <!-- lizard-marketing-slot:day-11-pm -->

    1 Reply Last reply
    0

    Hello! It looks like you're interested in this conversation, but you don't have an account yet.

    Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.

    With your input, this post could be even better 💗

    Register Login
    Reply
    • Reply as topic
    Log in to reply
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes


    • Login

    • Don't have an account? Register

    • Login or register to search.
    Powered by NodeBB Contributors
    • First post
      Last post
    0
    • Categories
    • Recent
    • Tags
    • Popular
    • World
    • Users
    • Groups