Release packed RecurQuant cache with held-out MBPP confirmation - #1
Merged
Conversation
# Conflicts: # README.md
Labeeb2339
marked this pull request as ready for review
July 22, 2026 13:59
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed
0.2.0a1package and generated evidence charts.Held-out result
On the pinned
Qwen/Qwen3.5-0.8B-Baserevision, the frozen layer-0 INT8 plus 17-layer INT4 layout reduced task-macro excess NLL from2.949743to0.803713versus uniform INT4, a 72.75% reduction, over 500 MBPP test tasks and 30,244 teacher-forced reference-code tokens.The paired mixed-vs-uniform improvement was
2.1460nats/token with 95% CI[2.0922, 2.1999]. The layout used exactly2,564,096resident recurrent-state bytes. Every preregistered quality gate passed.This is a scoped teacher-forced fidelity and resident-byte result. It is not a novelty, breakthrough, generated-code correctness, speed, peak-memory, whole-model-memory, or cross-model claim.
Verification
The raw 34.4 MB checkpoint is kept out of Git and will be attached to the
v0.2.0a1prerelease as a 9.1 MB archive so the full raw-array verification remains available.