Lossless model compression experiment: GLM-5.2 in 25% less memory

hambandit2 pts0 comments

Lossless Model Compression Experiment

Lossless BF16<br>Compression Research

A full GLM-5.2 scan found 30.168% K15 charged-format accounting. A separate byte-split representation was decoded bit-for-bit across all 59,509 BF16 tensors at 24.967% reduction. Those are distinct evidence classes.

GLM-5.2 753B · BF16<br>K15 ACCOUNTING

1,403 GiB

−423 GiB<br>30.17% SMALLER

[ FIG. 00 ] GLM-5.2 753B K15 CHARGED-FORMAT ACCOUNTING: 1,403.19 TO 979.87 GiB, 30.168%. SEPARATE BYTE-SPLIT EXACT INVERSE: 24.967%.

How it works

[ FIG. 01 ]

01 / Where the K15 estimate comes from<br>Most weights share a few sign-exponent symbols

Every BF16 weight is 16 bits: a sign, 8 exponent bits, 7 mantissa bits. In a trained model the exponents are wildly repetitive; a handful of values cover nearly every weight.

The K15 accounting replaces the 9-bit sign-and-exponent symbol with a 4-bit code into a 15-entry table. Rare 9-bit symbols go to an exact escape stream; the 7-bit mantissa stays raw. With all side costs charged, the full scan prices this layout at 11.173 bits per weight.

02 / What was decoded exactly<br>The byte-split path round-trips

The streamed validator reconstructed the high byte from its codebook, indices, and escape stream, kept the low byte verbatim, and compared every BF16 tensor with the source.

All 59,509 tensors matched bit-for-bit. That decoded representation costs 12.005 bits per weight, or 24.967% less than BF16. The 30.168% K15 layout is separately charged accounting and was not independently decoded at GLM scale.

03 / Runtime evidence and open work<br>A dense prototype, not a serving result

A separate dense 12-bit prototype reconstructed codes in registers and measured 0.733 times BF16 GEMV time on an A40. Sparse escape correction was validated separately but was not fused into or included in that timing.

An exact lossless speedup, a physical GLM-scale K15 container, and end-to-end serving integration therefore remain open.

04 / Prior work and distinction<br>The exponent insight and fused concept are prior art

ZipNN and DFloat11 established lossless BF16 exponent compression in this size range. DFloat11 uses variable-length Huffman codes and restores BF16 weights before matrix multiplication. ZipServ is the closest prior runtime design: it already demonstrated fixed-length coding with direct register reconstruction. This experiment uses a different representation, per-tensor 4-bit codes over 15 joint sign-and-exponent symbols with sparse exact escapes, and does not claim the shared exponent redundancy, roughly 11-bit range, or fused concept as new.

>>> Reproduction protocol<br>Don&rsquo;t trust it. Run it.

One script, no GPU. It streams the checkpoint shard by shard from Hugging Face (~1.4 TB, deleted as it goes), verifies the byte-split reconstruction bit-for-bit, and separately computes the fully charged K15 size accounting.

Open REPRODUCE.md ->

One command<br>Copy

uv run verify.py zai-org/GLM-5.2

GPUnot required

Inputthe BF16 checkpoint, streamed (~1.4 TB)

Exact pathbyte-split, all 59,509 tensors

K15 pathcomplete bit accounting, no GLM-scale decode

Verifiedblind gate, 2026-07-12

bf16 accounting exponent byte lossless charged

Related Articles