Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
4B parameters
0.59s Turbo generation
1.02s Turbo editing
4 Turbo steps
0.88 Turbo GenEval
8.271 GEdit-EN
Latency measured at 1024² on a single NVIDIA A100. Benchmark results follow the unified protocol in the technical report.
Abstract
Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow , a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The stack is built from two co-designed components: Mage-VAE , a lightweight high-fidelity latent tokenizer, and a Native-Resolution Multimodal Diffusion Transformer trained with rectified flow matching. Mage-VAE uses one-step diffusion-style encoding and decoding with anchor-latent regularization, preserving the reconstruction quality of strong public VAEs while reducing tokenization cost by more than an order of magnitude. Together with native-resolution packing and stack-level CUDA kernel fusion, the stack supports flexible-resolution training and improves end-to-end training throughput by about 2.5×. Built on this foundation, we develop a complete model family with Base, RL-aligned, and Turbo variants for both generation and editing. Diffusion-NFT improves prompt following, text rendering, aesthetic quality, and editing fidelity, while few-step distillation with adversarial perceptual guidance produces 4-step Turbo models for low-latency inference. Despite its compact scale, Mage-Flow and Mage-Flow-Edit achieve competitive performance across standard generation and editing benchmarks. More importantly, the Turbo variants make high-resolution generation and editing practical for interactive use: at 1024² resolution on a single NVIDIA A100 GPU, Mage-Flow-Turbo generates an image in 0.59 s, and Mage-Flow-Edit-Turbo edits an image in 1.02 s, while maintaining a small memory footprint. These results show that careful tokenizer–backbone–system co-design can deliver strong high-resolution generation and editing within an efficient 4B model family.
Full Family & Speed
Mage-VAE Tokenizer
Native-Resolution MMDiT
Full family & speed
The same compact stack powers generation and editing, each available as Base, aligned, and 4-step Turbo variants. Stack-level optimization and quality-preserving distillation make high-resolution inference interactive.
0.59s<br>Turbo generation
1.02s<br>Turbo editing
18–20 GB<br>Peak inference memory
Quality vs. latency & memory at 1024² on a single A100 — GenEval (generation, left) and GEdit-EN (editing, right). Mage-Flow sits at the fast, low-memory frontier.
Mage-VAE — efficient latent tokenizer
One-step encoding and decoding remove the high-resolution tokenizer bottleneck while preserving a generation-ready latent space. Mage-VAE matches strong reconstruction quality with ~12× / ~22× fewer encode / decode MACs per pixel .
~12×<br>Fewer encode MACs
~22×<br>Fewer decode MACs
One-step diffusion encode/decode with anchor-latent regularization.
Native-Resolution Multimodal DiT
Variable-length image and text packing avoids rigid resolution buckets and unnecessary padding. One shared 4B backbone supports flexible canvases from 512 to 2048 pixels , including extreme aspect ratios, while packed CFG evaluates both branches in one forward pass.
Packed native-resolution image and text tokens through the shared NR-MMDiT.
Benchmark highlights
Text-to-image generation<br>4 steps
0.88<br>GenEval · Mage-Flow-Turbo
0.873<br>CVTG-2K · Mage-Flow-Turbo
ModelParamsGenEvalCVTG-2K
Mage-Flow-Turbo4B0.880.873<br>Z-Image-Turbo6B0.820.859<br>Qwen-Image20B0.870.829<br>FLUX.2-Klein-9B9B0.860.424
Instruction-based editing<br>4 steps
8.271<br>GEdit-EN · Edit-Turbo
8.264<br>GEdit-CN · Edit-Turbo
ModelParamsGEdit-ENGEdit-CN
Mage-Flow-Edit-Turbo4B8.2718.264<br>FireRed-Image-Edit-1.020B7.9437.887<br>JoyAI-Image-Edit16B8.2768.125<br>Qwen-Image-Edit-251120B7.8777.819
View more text-to-image results
20 models · 13 columns. GenEval, CVTG-2K, OneIG, and LongText use 0–1 scales; DPG and TIIF use 0–100 scales.
View more image editing results
20 models · 9 columns. ImgEdit uses a 0–5 scale, GEdit a 0–10 scale, and TextEdit a 0–25 scale.
Qualitative results
Text-to-image
Instruction-based editing
SOURCE
BACKGROUND CHANGE
Background change<br>Color alteration<br>Count change<br>Extract<br>Canny<br>Colorization<br>Depth<br>HED<br>Normal map<br>Segmentation<br>Sketch<br>Material alteration<br>Motion change<br>Style change<br>Subject addition<br>Subject removal<br>Subject replacement<br>Tone transfer<br>Viewpoint
Editing diversity — select an edit type to compare the shared source image with the corresponding Mage-Flow-Edit result.
Contributors
Xinjie Zhang*†,<br>Peng Zhang*,<br>Shicheng Zheng*,<br>Jinghao Guo*,<br>Zhaoyang Jia*,<br>Yifei Shen*,<br>Xun Guo,<br>Yuxuan Luo,<br>Jiahao Li,<br>Wenxuan Xie,<br>Fanyi Pu,<br>Xiaoyi Zhang,<br>Kaichen Zhang,<br>Zongyu Guo,<br>Tianci Bi,<br>Dongnan Gui,<br>Zhening Liu,<br>Zimo Wen,<br>Zihan Zheng,<br>Senqiao Yang,<br>Xiao Li,<br>Jinglu Wang,<br>Bin...