Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout Paper • 2609.09123 • Published 13 days ago • 53
H3-World: Turning Language Understanding into World Control Paper • 2609.01560 • Published 20 days ago • 51
SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models Paper • 2609.02886 • Published 19 days ago • 115
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published Jul 30 • 157
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Paper • 2607.07508 • Published Jul 8 • 33
ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding Paper • 2507.14533 • Published Jul 19, 2025 • 9
Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think Paper • 2410.06940 • Published Oct 9, 2024 • 14
Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis Paper • 2603.06507 • Published Mar 6 • 7
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU Paper • 2607.19191 • Published Jul 21 • 114
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Paper • 2607.19064 • Published Jul 21 • 78
Bridging Supervised Learning and Reinforcement Learning in Math Reasoning Paper • 2505.18116 • Published May 23, 2025 • 5
Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence Paper • 2607.16401 • Published Jul 17 • 45
RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources Paper • 2606.29538 • Published Jul 16 • 144
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget Paper • 2607.13125 • Published Jul 18 • 139
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published Jul 10 • 83
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Paper • 2607.07675 • Published Jul 8 • 64
DataComp-VLM: Improved Open Datasets for Vision-Language Models Paper • 2606.28551 • Published Jun 26 • 52
ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog Paper • 2607.04438 • Published Jul 5 • 61