Lizard inference engineering: Verify several drafted tokens in one pass
-
Inference Engineering · Day 5 · Evening

What if one pass over the model weights could accept several tokens?
Caterpillar's speculative path drafts a short sequence, then verifies multiple candidates in a single weight pass. It is lossless: rejected drafts fall back to the normal path, so accepted output remains the output the model would have produced.
On predictable or repetitive text, our tests reached up to 2× speed. On ordinary prose, the path steps aside rather than forcing a bad optimization.
Optimization should be conditional when the workload is conditional.
Engineering fact: Caterpillar's lossless speculative path can verify multiple drafted tokens per weight pass on predictable output and steps aside on ordinary prose.
#SpeculativeDecoding #CaterpillarEngine #LLMInference #LocalAI #LizardLLM
<!-- lizard-marketing-slot:day-05-pm -->
Hello! It looks like you're interested in this conversation, but you don't have an account yet.
Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.
With your input, this post could be even better 💗
Register Login