[research] KV-Cache Grafting: smaller model beats larger one, 93.3% vs 89.2% #304
Closed
Replies: 1 comment
|
This discussion was automatically closed because it expired on 2026-07-27T10:12:12.042Z.
|
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
🔬 The Finding
Researchers introduce KV-Cache Grafting — pre-computing verified knowledge as a byte-exact KV-state artifact and "grafting" it into fresh inference contexts with zero weight changes. On AIME 2025, a frozen Gemma-4-12B jumps from 80.0% → 93.3% accuracy after grafting a verified solution library, surpassing its 31B sibling's 89.2% — all with dramatically lower per-query compute. The restore is cryptographically verified (SHA-256 bit-exact, zero KL divergence).
⚙️ What It Means for Agentic Workflows
🔗 Source
Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting — 15 Jul 2026
All reactions