Skip to content
  • Categories
  • Recent
  • Tags
  • Popular
  • World
  • Users
  • Groups
Skins
  • Light
  • Brite
  • Cerulean
  • Cosmo
  • Flatly
  • Journal
  • Litera
  • Lumen
  • Lux
  • Materia
  • Minty
  • Morph
  • Pulse
  • Sandstone
  • Simplex
  • Sketchy
  • Spacelab
  • United
  • Yeti
  • Zephyr
  • Dark
  • Cyborg
  • Darkly
  • Quartz
  • Slate
  • Solar
  • Superhero
  • Vapor

  • Default (No Skin)
  • No Skin
Collapse

Lizard-LLM Community

  1. Home
  2. Benchmarks
  3. New benchmark evidence: completed provider comparisons

New benchmark evidence: completed provider comparisons

Scheduled Pinned Locked Moved Benchmarks
1 Posts 1 Posters 3 Views 1 Watching
  • Oldest to Newest
  • Newest to Oldest
  • Most Votes
Reply
  • Reply as topic
Log in to reply
This topic has been deleted. Only users with topic management privileges can see it.
  • L Offline
    L Offline
    lizardadmin
    wrote on last edited by
    #1

    These four captures show why the dashboard keeps provider rows and analysis separate. The same machine can produce a different winner in a different run, and a composite recommendation is never allowed to erase the underlying values.

    Interactive gallery: https://lizard-llm.qendryx.com/benchmarks.html

    Five rows, five completed measurements

    Completed five-row provider comparison

    Ollama records 8.56 tok/s, stock llama.cpp Q4 7.85, and Lizard Native Q4 5.61 in this run. Load, warm-up and memory evidence stay visible beside throughput.

    A second run changes the order

    Expanded ten-row provider comparison

    In this separate 10/10 run, llama.cpp Q4 reaches 6.23 tok/s, Lizard Native Q4 5.28, and Ollama 4.53. The result belongs to this model, machine and run—not a universal ranking.

    Leaderboard plus metric-specific winners

    Provider leaderboard and insight panel

    Ollama leads the composite score, Lizard Native Q4 takes best native decode and parity among the compared rows, and llama.cpp Q4 leads wall time and end-to-end tokens per second.

    Recommendations stay attached to source evidence

    Wide benchmark leaderboard and evidence panel

    Scores, badges, execution lanes and raw provider values remain on the same screen. The recommendation annotates the run; it does not replace it.


    Question for you: On your hardware, which decision should come first: shortest wall time, steady decode speed, peak memory, or answer parity?

    1 Reply Last reply
    0

    Hello! It looks like you're interested in this conversation, but you don't have an account yet.

    Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.

    With your input, this post could be even better 💗

    Register Login
    Reply
    • Reply as topic
    Log in to reply
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes


    • Login

    • Don't have an account? Register

    • Login or register to search.
    Powered by NodeBB Contributors
    • First post
      Last post
    0
    • Categories
    • Recent
    • Tags
    • Popular
    • World
    • Users
    • Groups