The resident-GPU Direct3D 12 engine: questions, tips, issues.
This category can be followed from the open social web via the handle lizard-native@community.lizard-llm.qendryx.com
-
Lizard Native benchmark CLI and HTTP command reference
Watching Ignoring Scheduled Pinned Locked Moved0 Votes1 Posts3 Views -
Lizard inference engineering: Context length belongs in lived performance history
Watching Ignoring Scheduled Pinned Locked Moved0 Votes1 Posts0 Views -
Lizard inference engineering: Chat metrics should be scoped to the active provider
Watching Ignoring Scheduled Pinned Locked Moved0 Votes1 Posts0 Views -
Lizard inference engineering: Prompt caching should follow conversation structure
Watching Ignoring Scheduled Pinned Locked Moved0 Votes1 Posts0 Views -
Lizard inference engineering: Warm the selected model before the first conversation
Watching Ignoring Scheduled Pinned Locked Moved0 Votes1 Posts0 Views -
Lizard inference engineering: CPU-only should remain a first-class path
Watching Ignoring Scheduled Pinned Locked Moved0 Votes1 Posts0 Views -
Lizard inference engineering: Integrated GPUs need a shared-memory plan
Watching Ignoring Scheduled Pinned Locked Moved0 Votes1 Posts0 Views -
Lizard inference engineering: Capacity still matters when no benchmark exists
Watching Ignoring Scheduled Pinned Locked Moved0 Votes1 Posts1 Views -
Lizard inference engineering: Let measured B16 and B8 results choose the lane
Watching Ignoring Scheduled Pinned Locked Moved0 Votes1 Posts0 Views -
Lizard inference engineering: TTFT and completion time diagnose different pain
Watching Ignoring Scheduled Pinned Locked Moved0 Votes1 Posts0 Views -
Lizard inference engineering: Separate user-visible speed from native decode speed
Watching Ignoring Scheduled Pinned Locked Moved0 Votes1 Posts0 Views -
Lizard inference engineering: Make KV-cache precision visible
Watching Ignoring Scheduled Pinned Locked Moved0 Votes1 Posts0 Views -
Lizard inference engineering: Use UMA without a fake copy
Watching Ignoring Scheduled Pinned Locked Moved0 Votes1 Posts0 Views -
Lizard inference engineering: Stop rebuilding descriptors per token
Watching Ignoring Scheduled Pinned Locked Moved0 Votes1 Posts0 Views -
Lizard inference engineering: Stop oversized models from thrashing
Watching Ignoring Scheduled Pinned Locked Moved0 Votes1 Posts0 Views -
Lizard inference engineering: Budget against free VRAM, not the sticker
Watching Ignoring Scheduled Pinned Locked Moved0 Votes1 Posts0 Views -
Lizard inference engineering: Keep model weights resident
Watching Ignoring Scheduled Pinned Locked Moved0 Votes1 Posts0 Views -
Technical overview: how Lizard Native keeps Direct3D 12 inference resident
Watching Ignoring Scheduled Pinned Locked Moved0 Votes1 Posts3 Views -
Which graphics cards work — post yours
Watching Ignoring Scheduled Pinned Locked Moved0 Votes1 Posts1 Views -
Running a model that's bigger than your graphics card
Watching Ignoring Scheduled Pinned Locked Moved0 Votes1 Posts0 Views