Image and text do not share a coordinate system

apsys1 pts0 comments

Konstantin on X: "In Kimi K3, a caption's embedding points almost perfectly backwards from its own image

That's not a bug. It's a coordinate system choice, and you can undo it with a single rotation馃У" / X<br>Post

Log inSign up

Post

Konstantin on X: "In Kimi K3, a caption's embedding points almost perfectly backwards from its own image

That's not a bug. It's a coordinate system choice, and you can undo it with a single rotation馃У"

Konstantin

@advprop

In Kimi K3, a caption's embedding points almost perfectly backwards from its own image

That's not a bug. It's a coordinate system choice, and you can undo it with a single rotation馃У

span:not(:empty)~span:not(:empty)]:before:content-['路'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0">11:33 PM 路 Aug 15, 202613Views

span:not(:empty)~span:not(:empty)]:before:content-['路'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Konstantin

@advprop

17m

a linear classifier separates image from text embeddings with 100% accuracy in all 3 models we tested (Kimi K3, Inkling, Qwen3-Omni).

It just reads off each model's coordinate conventions, not what the modalities actually share.<br>(full writeup with interactive figures is at Show more

span:not(:empty)~span:not(:empty)]:before:content-['路'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Konstantin

@advprop

17m

Where the conventions come from: each modality is packed in its own cone.

Kimi K3 packs all text into 1.6掳. Qwen3-Omni does the same to images (2.0掳). Same architecture family, opposite choice.

The encoder-free Inkling from @thinkymachines keeps both wide.

span:not(:empty)~span:not(:empty)]:before:content-['路'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Konstantin

@advprop

17m

So we tried the dumbest possible fix: whiten each modality, then fit ONE rank-32 rotation on training pairs.

No scaling. No translation. No MLP. A rotation is not allowed to move or reshape anything, so it can only turn the image basis into the text basis.

span:not(:empty)~span:not(:empty)]:before:content-['路'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Konstantin

@advprop

17m

Result: held-out images retrieve their own captions at

9.9x chance in Kimi K3<br>9.2x in Qwen3-Omni<br>3.2x in Inkling

A pure rotation, fitted on 96 pairs, recovers that much of the "modality gap".

span:not(:empty)~span:not(:empty)]:before:content-['路'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Konstantin

@advprop

17m

If you compare image and text activations directly using probes, similarity metrics, steering, the native coordinates will lie to you. Part of the "gap" is basis mismatch, not missing content.

This research is done at @whitecircle -- we are hiring!

Log in or sign up for X<br>See what鈥檚 happening and join the conversation<br>Continue with phoneContinue with AppleContinue with Google<br>or<br>Log in with username or email

Relevant people

Konstantin@advpropFollow<br>head of applied research @whitecircle

Trending now

span empty before konstantin image content

Related Articles