Skip to content
  • Categories
  • Recent
  • Tags
  • Popular
  • World
  • Users
  • Groups
Skins
  • Light
  • Brite
  • Cerulean
  • Cosmo
  • Flatly
  • Journal
  • Litera
  • Lumen
  • Lux
  • Materia
  • Minty
  • Morph
  • Pulse
  • Sandstone
  • Simplex
  • Sketchy
  • Spacelab
  • United
  • Yeti
  • Zephyr
  • Dark
  • Cyborg
  • Darkly
  • Quartz
  • Slate
  • Solar
  • Superhero
  • Vapor

  • Default (No Skin)
  • No Skin
Collapse

Lizard-LLM Community

  1. Home
  2. Benchmarks
  3. What the latest Native and Caterpillar results actually show

What the latest Native and Caterpillar results actually show

Scheduled Pinned Locked Moved Benchmarks
1 Posts 1 Posters 4 Views 1 Watching
  • Oldest to Newest
  • Newest to Oldest
  • Most Votes
Reply
  • Reply as topic
Log in to reply
This topic has been deleted. Only users with topic management privileges can see it.
  • L Offline
    L Offline
    lizardadmin
    wrote on last edited by
    #1

    These analyzed captures put Lizard Native and Caterpillar in the foreground beside stock llama.cpp and optional Ollama. They also show the honest result: no provider wins every metric, and Caterpillar trails the other two lanes in these selected general-case runs.

    Interactive gallery: https://lizard-llm.qendryx.com/benchmarks.html

    Wall time, throughput and decode are different metrics

    Runtime throughput and decode charts

    llama.cpp wins this run's wall time and end-to-end throughput. Lizard Native leads native decode at 15.198 tok/s versus Caterpillar's 9.201. The dashboard does not blend those numbers.

    No provider wins every dimension

    Provider heatmap radar and history charts

    The heatmap and radar show the trade: llama.cpp leads speed and memory efficiency here, while the native lanes carry their own decode and execution-lane evidence.

    The decision guide names the real winners

    Measured practical decision guide

    For this hardware, llama.cpp Q4 is the balanced choice at 16.31 tok/s and 3.925 seconds. Lizard Native Q4 is called out separately for best native decode at 15.198 tok/s. Ollama was not included and remains unmeasured.

    A larger model changes the gap

    Gemma 4 provider comparison

    For Gemma 4 E4B, llama.cpp records 4.37 tok/s, Lizard Native 4.08, and Caterpillar 1.41. The Ollama card says not included and unmeasured—not zero.

    Provider identity stays attached to every metric

    Full-width Gemma 4 provider cards

    Runtime, model load, warm-up, peak memory and efficiency stay under the provider that produced them. This is the evidence needed to evaluate Lizard as a local inference provider.


    Question for you: What hardware and model should we run next to test where Caterpillar closes the gap—or where Native's decode path matters most?

    1 Reply Last reply
    0

    Hello! It looks like you're interested in this conversation, but you don't have an account yet.

    Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.

    With your input, this post could be even better 💗

    Register Login
    Reply
    • Reply as topic
    Log in to reply
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes


    • Login

    • Don't have an account? Register

    • Login or register to search.
    Powered by NodeBB Contributors
    • First post
      Last post
    0
    • Categories
    • Recent
    • Tags
    • Popular
    • World
    • Users
    • Groups