Lizard inference engineering: Integrated GPUs need a shared-memory plan
-
Inference Engineering · Day 19 · Morning

Integrated GPUs change the memory question, not just the adapter name.
On a shared-memory system, the model, KV cache, application, compositor, and CPU all compete inside one physical budget. Lizard distinguishes this configuration when planning fit and execution rather than treating advertised graphics memory as an isolated pool.
The useful number is the memory currently available to the whole workload.
How does your runtime budget an iGPU while the desktop and other applications are active?
Engineering fact: Lizard's hardware planning distinguishes integrated-memory and discrete-memory systems so model fit and execution choices can account for shared RAM instead of treating every GPU like a separate VRAM device.
#LizardNative #LizardLLM #LocalAI #GGUF #InferenceEngineering
<!-- lizard-marketing-slot:day-19-am -->
Hello! It looks like you're interested in this conversation, but you don't have an account yet.
Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.
With your input, this post could be even better 💗
Register Login