Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See Paper • 2608.17744 • Published 13 days ago • 15
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Paper • 2608.17426 • Published 14 days ago • 157
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published 18 days ago • 280
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published about 1 month ago • 263
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations Paper • 2607.28956 • Published Jul 31 • 97
CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Paper • 2608.02589 • Published 29 days ago • 25
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering Paper • 2607.28568 • Published Jul 30 • 185
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published Jul 30 • 310