Konstantin on X: "In Kimi K3, a caption's embedding points almost perfectly backwards from its own image
That's not a bug. It's a coordinate system choice, and you can undo it with a single rotation馃У" / X<br>Post
Log inSign up
Post
Konstantin on X: "In Kimi K3, a caption's embedding points almost perfectly backwards from its own image
That's not a bug. It's a coordinate system choice, and you can undo it with a single rotation馃У"
Konstantin
@advprop
In Kimi K3, a caption's embedding points almost perfectly backwards from its own image
That's not a bug. It's a coordinate system choice, and you can undo it with a single rotation馃У
span:not(:empty)~span:not(:empty)]:before:content-['路'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0">11:33 PM 路 Aug 15, 202613Views
span:not(:empty)~span:not(:empty)]:before:content-['路'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Konstantin
@advprop
17m
a linear classifier separates image from text embeddings with 100% accuracy in all 3 models we tested (Kimi K3, Inkling, Qwen3-Omni).
It just reads off each model's coordinate conventions, not what the modalities actually share.<br>(full writeup with interactive figures is at Show more
span:not(:empty)~span:not(:empty)]:before:content-['路'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Konstantin
@advprop
17m
Where the conventions come from: each modality is packed in its own cone.
Kimi K3 packs all text into 1.6掳. Qwen3-Omni does the same to images (2.0掳). Same architecture family, opposite choice.
The encoder-free Inkling from @thinkymachines keeps both wide.
span:not(:empty)~span:not(:empty)]:before:content-['路'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Konstantin
@advprop
17m
So we tried the dumbest possible fix: whiten each modality, then fit ONE rank-32 rotation on training pairs.
No scaling. No translation. No MLP. A rotation is not allowed to move or reshape anything, so it can only turn the image basis into the text basis.
span:not(:empty)~span:not(:empty)]:before:content-['路'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Konstantin
@advprop
17m
Result: held-out images retrieve their own captions at
9.9x chance in Kimi K3<br>9.2x in Qwen3-Omni<br>3.2x in Inkling
A pure rotation, fitted on 96 pairs, recovers that much of the "modality gap".
span:not(:empty)~span:not(:empty)]:before:content-['路'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Konstantin
@advprop
17m
If you compare image and text activations directly using probes, similarity metrics, steering, the native coordinates will lie to you. Part of the "gap" is basis mismatch, not missing content.
This research is done at @whitecircle -- we are hiring!
Log in or sign up for X<br>See what鈥檚 happening and join the conversation<br>Continue with phoneContinue with AppleContinue with Google<br>or<br>Log in with username or email
Relevant people
Konstantin@advpropFollow<br>head of applied research @whitecircle
Trending now