Mage-Flow: Efficient Native-Resolution Foundation Model for Image Generation

macote1 pts0 comments

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

4B parameters

0.59s Turbo generation

1.02s Turbo editing

4 Turbo steps

0.88 Turbo GenEval

8.271 GEdit-EN

Latency measured at 1024² on a single NVIDIA A100. Benchmark results follow the unified protocol in the technical report.

Abstract

Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow , a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The stack is built from two co-designed components: Mage-VAE , a lightweight high-fidelity latent tokenizer, and a Native-Resolution Multimodal Diffusion Transformer trained with rectified flow matching. Mage-VAE uses one-step diffusion-style encoding and decoding with anchor-latent regularization, preserving the reconstruction quality of strong public VAEs while reducing tokenization cost by more than an order of magnitude. Together with native-resolution packing and stack-level CUDA kernel fusion, the stack supports flexible-resolution training and improves end-to-end training throughput by about 2.5×. Built on this foundation, we develop a complete model family with Base, RL-aligned, and Turbo variants for both generation and editing. Diffusion-NFT improves prompt following, text rendering, aesthetic quality, and editing fidelity, while few-step distillation with adversarial perceptual guidance produces 4-step Turbo models for low-latency inference. Despite its compact scale, Mage-Flow and Mage-Flow-Edit achieve competitive performance across standard generation and editing benchmarks. More importantly, the Turbo variants make high-resolution generation and editing practical for interactive use: at 1024² resolution on a single NVIDIA A100 GPU, Mage-Flow-Turbo generates an image in 0.59 s, and Mage-Flow-Edit-Turbo edits an image in 1.02 s, while maintaining a small memory footprint. These results show that careful tokenizer–backbone–system co-design can deliver strong high-resolution generation and editing within an efficient 4B model family.

Full Family & Speed

Mage-VAE Tokenizer

Native-Resolution MMDiT

Full family & speed

The same compact stack powers generation and editing, each available as Base, aligned, and 4-step Turbo variants. Stack-level optimization and quality-preserving distillation make high-resolution inference interactive.

0.59s<br>Turbo generation

1.02s<br>Turbo editing

18–20 GB<br>Peak inference memory

Quality vs. latency & memory at 1024² on a single A100 — GenEval (generation, left) and GEdit-EN (editing, right). Mage-Flow sits at the fast, low-memory frontier.

Mage-VAE — efficient latent tokenizer

One-step encoding and decoding remove the high-resolution tokenizer bottleneck while preserving a generation-ready latent space. Mage-VAE matches strong reconstruction quality with ~12× / ~22× fewer encode / decode MACs per pixel .

~12×<br>Fewer encode MACs

~22×<br>Fewer decode MACs

One-step diffusion encode/decode with anchor-latent regularization.

Native-Resolution Multimodal DiT

Variable-length image and text packing avoids rigid resolution buckets and unnecessary padding. One shared 4B backbone supports flexible canvases from 512 to 2048 pixels , including extreme aspect ratios, while packed CFG evaluates both branches in one forward pass.

Packed native-resolution image and text tokens through the shared NR-MMDiT.

Benchmark highlights

Text-to-image generation<br>4 steps

0.88<br>GenEval · Mage-Flow-Turbo

0.873<br>CVTG-2K · Mage-Flow-Turbo

ModelParamsGenEvalCVTG-2K

Mage-Flow-Turbo4B0.880.873<br>Z-Image-Turbo6B0.820.859<br>Qwen-Image20B0.870.829<br>FLUX.2-Klein-9B9B0.860.424

Instruction-based editing<br>4 steps

8.271<br>GEdit-EN · Edit-Turbo

8.264<br>GEdit-CN · Edit-Turbo

ModelParamsGEdit-ENGEdit-CN

Mage-Flow-Edit-Turbo4B8.2718.264<br>FireRed-Image-Edit-1.020B7.9437.887<br>JoyAI-Image-Edit16B8.2768.125<br>Qwen-Image-Edit-251120B7.8777.819

View more text-to-image results

20 models · 13 columns. GenEval, CVTG-2K, OneIG, and LongText use 0–1 scales; DPG and TIIF use 0–100 scales.

View more image editing results

20 models · 9 columns. ImgEdit uses a 0–5 scale, GEdit a 0–10 scale, and TextEdit a 0–25 scale.

Qualitative results

Text-to-image

Instruction-based editing

SOURCE

BACKGROUND CHANGE

Background change<br>Color alteration<br>Count change<br>Extract<br>Canny<br>Colorization<br>Depth<br>HED<br>Normal map<br>Segmentation<br>Sketch<br>Material alteration<br>Motion change<br>Style change<br>Subject addition<br>Subject removal<br>Subject replacement<br>Tone transfer<br>Viewpoint

Editing diversity — select an edit type to compare the shared source image with the corresponding Mage-Flow-Edit result.

Contributors

Xinjie Zhang*†,<br>Peng Zhang*,<br>Shicheng Zheng*,<br>Jinghao Guo*,<br>Zhaoyang Jia*,<br>Yifei Shen*,<br>Xun Guo,<br>Yuxuan Luo,<br>Jiahao Li,<br>Wenxuan Xie,<br>Fanyi Pu,<br>Xiaoyi Zhang,<br>Kaichen Zhang,<br>Zongyu Guo,<br>Tianci Bi,<br>Dongnan Gui,<br>Zhening Liu,<br>Zimo Wen,<br>Zihan Zheng,<br>Senqiao Yang,<br>Xiao Li,<br>Jinglu Wang,<br>Bin...

mage image turbo editing flow resolution

Related Articles