Mage a Lightweight, Research-Friendly Multimodal Model Family

nmstoker1 pts0 comments

Mage — A Lightweight, Research-Friendly Multimodal Model Family

Mage

A Lightweight, Research-Friendly Multimodal Model Family

Microsoft Mage Team

Code

🤗<br>Models

Mage is a family of lightweight, research-friendly multimodal models built at a fixed 4B-parameter budget, sharing a codec-aligned efficiency philosophy — spend representation capacity where the signal is — across both visual understanding and generation . Both models are compact enough to train, fine-tune, and deploy on modest hardware, yet remain competitive with much larger open systems.

Vision–Language · Understanding<br>Mage-VL

An Efficient Codec-Native Streaming Multimodal Foundation Model

A codec-native, from-scratch VLM for image & video understanding — reads video the way a codec does (anchor/predicted frames, 16×16 patches), with bio-inspired proactive streaming.

Enter project page →

Generation · Editing<br>Mage-Flow

An Efficient Native-Resolution Foundation Model for Image Generation and Editing

A compact 4B generative stack (Mage-VAE + Native-Resolution MMDiT) for text-to-image generation and instruction-based editing at native resolution, with Base / RL / 4-step Turbo variants.

Enter project page →

mage multimodal model native lightweight research

Related Articles