FLUX-Reason-6M
FLUX-Reason-6M is a massive, 6-million-scale text-to-image dataset engineered to instill complex reasoning capabilities in generative models. This dataset was created to bridge the performance gap between open-source and leading closed-source text-to-image systems.
This dataset contains:
- 6 million high-quality, reasoning-focused images synthesized by the state-of-the-art FLUX.1-dev model.
- 20 million bilingual (English and Chinese) descriptions, providing a rich, multi-faceted annotation for each image.
- Pioneering Generation Chain-of-Thought (GCoT) prompts that provide detailed, step-by-step breakdowns of the image generation process, moving beyond simple descriptions to explain compositional and semantic logic.
- A systematic organization across six key reasoning characteristics: Imagination, Entity, Text rendering, Style, Affection, and Composition.
The creation of this dataset was a significant undertaking, requiring 15,000 A100 GPU days. We are releasing it to provide the community with a resource previously unattainable outside of large industrial labs.
See our paper for more details!
Dataset Architectural Design
The core of FLUX-Reason-6M is its multidimensional framework, designed to teach models the foundational principles of visual reasoning. Each image is annotated with multiple labels and caption types.
The Six Characteristics
- Imagination: Captions and images representing surreal, fantastical, or abstract concepts that push beyond literal interpretations (e.g., βa city made of glass where rivers of light flow").
- Entity: Focuses on knowledge-grounded depiction of specific real-world objects, beings, or named entities with high fidelity (e.g., βLionel Messi dribbling past defenders in the World Cup finalβ).
- Text rendering: Addresses the common weakness of text generation in images, providing clean data for typographic control with explicit instruct