Skip to content
  • Categories
  • Recent
  • Tags
  • Popular
  • World
  • Users
  • Groups
Skins
  • Light
  • Brite
  • Cerulean
  • Cosmo
  • Flatly
  • Journal
  • Litera
  • Lumen
  • Lux
  • Materia
  • Minty
  • Morph
  • Pulse
  • Sandstone
  • Simplex
  • Sketchy
  • Spacelab
  • United
  • Yeti
  • Zephyr
  • Dark
  • Cyborg
  • Darkly
  • Quartz
  • Slate
  • Solar
  • Superhero
  • Vapor

  • Default (No Skin)
  • No Skin
Collapse

Lizard-LLM Community

  1. Home
  2. Architecture
  3. Where do your tokens actually come from?

Where do your tokens actually come from?

Scheduled Pinned Locked Moved Architecture
1 Posts 1 Posters 1 Views 1 Watching
  • Oldest to Newest
  • Newest to Oldest
  • Most Votes
Reply
  • Reply as topic
Log in to reply
This topic has been deleted. Only users with topic management privileges can see it.
  • L Offline
    L Offline
    lizardadmin
    wrote on last edited by
    #1

    When you call an LLM API, it is easy to assume the company you are calling is the one running the model. Often it isn't, and the difference matters more than it looks.

    Where do your tokens come from

    An API router is a proxy. It owns no compute. Your request arrives, gets forwarded to whichever upstream provider it has a deal with, and the response comes back through it. Useful — one key, many models, automatic failover when an endpoint degrades. But every request takes an extra network hop, so median time-to-first-token is always worse than calling that same provider directly. Routers improve tail latency by rerouting around bad endpoints; they cannot improve the median. Nothing they do touches the decisions that actually determine speed.

    An inference provider owns the hardware. The same company controls the endpoint and the GPUs your tokens are computed on. They choose the batching strategy, the KV cache layout, the kernels, the quantization. Those choices are what your latency and throughput actually come from.

    There is a second, quieter difference. A router can only bind itself. If you sign a zero-retention agreement with a router, your request still lands at an upstream provider whose policies you never reviewed. The agreement doesn't travel with your data.

    Lizard is an inference provider where the provider is your own machine. The endpoint is 127.0.0.1:7817, the GPU is the one in your computer, and there is no upstream at all — no hop to add latency, no third party to have a retention policy, no agreement chain to audit. When you turn off your Wi-Fi, it keeps working. That is the honest test of the claim, and you can run it yourself.

    The trade is real and worth stating: you get one machine's worth of compute, not a fleet. A hosted provider will out-run your laptop on a 70B model. What you get instead is that nothing leaves the building.

    How do you tell which kind you are using? Read the docs. "Our clusters", "our GPUs" means a provider. "Access 200+ models from leading providers" means a router. In a DPA, any mention of sub-processors or third-party infrastructure partners is a router tell — a direct provider has no such chain.

    1 Reply Last reply
    0

    Hello! It looks like you're interested in this conversation, but you don't have an account yet.

    Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.

    With your input, this post could be even better 💗

    Register Login
    Reply
    • Reply as topic
    Log in to reply
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes


    • Login

    • Don't have an account? Register

    • Login or register to search.
    Powered by NodeBB Contributors
    • First post
      Last post
    0
    • Categories
    • Recent
    • Tags
    • Popular
    • World
    • Users
    • Groups