Skip to content
  • Categories
  • Recent
  • Tags
  • Popular
  • World
  • Users
  • Groups
Skins
  • Light
  • Brite
  • Cerulean
  • Cosmo
  • Flatly
  • Journal
  • Litera
  • Lumen
  • Lux
  • Materia
  • Minty
  • Morph
  • Pulse
  • Sandstone
  • Simplex
  • Sketchy
  • Spacelab
  • United
  • Yeti
  • Zephyr
  • Dark
  • Cyborg
  • Darkly
  • Quartz
  • Slate
  • Solar
  • Superhero
  • Vapor

  • Default (No Skin)
  • No Skin
Collapse

Lizard-LLM Community

  1. Home
  2. Caterpillar
  3. Lizard inference engineering: More CPU threads can be slower

Lizard inference engineering: More CPU threads can be slower

Scheduled Pinned Locked Moved Caterpillar
1 Posts 1 Posters 0 Views 1 Watching
  • Oldest to Newest
  • Newest to Oldest
  • Most Votes
Reply
  • Reply as topic
Log in to reply
This topic has been deleted. Only users with topic management privileges can see it.
  • L Offline
    L Offline
    lizardadmin
    wrote last edited by lizardadmin
    #1

    Inference Engineering · Day 5 · Morning

    Repeatable command recipes — lizard-llm.qendryx.com

    Caterpillar now defaults to physical cores, and that change came from a simple result: on a four-core test laptop, using every logical thread made decode slower. Switching to the physical-core count improved measured decode by about 9% and also made the timing less noisy.

    The important part is not the exact laptop. It is the shape of the workload. Decode can be limited by memory bandwidth long before it is limited by raw thread count. Once that happens, hyperthreads stop looking like extra capacity and start looking like contention. More schedulable threads can mean more pressure on the same caches, the same memory channels, and the same execution resources.

    That is why the default changed. A local runtime should choose the thread count that matches the bottleneck, not the number that looks best on a spec sheet. The override still matters for experiments and unusual machines, but the default should bias toward the setting that is most likely to behave well without tuning.

    I have seen this pattern enough times to treat thread count as part of measurement, not just configuration. If a workload is bandwidth-bound, the “more threads” instinct can be the wrong first move.

    What’s the first sign you use to tell whether extra logical threads are helping, or just adding noise?

    Engineering fact: Caterpillar defaults to physical cores; on a four-core test laptop this improved decode by about 9% versus using every logical thread.

    Lizard The AI Runtime You'll Own—Not Rent.

    Receive two professional Windows AI runtimes with lifetime updates. Run AI at native speed, keep every conversation private, and stay independent with intelligent hardware optimization and no cloud dependency.

    Read the relevant Lizard page

    #localai #cpu #performance #benchmarking #systems

    <!-- lizard-marketing-slot:day-05-am -->

    1 Reply Last reply
    0

    Hello! It looks like you're interested in this conversation, but you don't have an account yet.

    Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.

    With your input, this post could be even better 💗

    Register Login
    Reply
    • Reply as topic
    Log in to reply
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes


    • Login

    • Don't have an account? Register

    • Login or register to search.
    Powered by NodeBB Contributors
    • First post
      Last post
    0
    • Categories
    • Recent
    • Tags
    • Popular
    • World
    • Users
    • Groups