StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling Paper • 2608.15089 • Published 8 days ago • 435
Scaling Properties of Text Conditioning in Visual Generation Paper • 2607.29679 • Published 23 days ago • 40
UniVR: Thinking in Visual Space for Unified Visual Reasoning Paper • 2607.12800 • Published Jul 14 • 32
SPIRAL: Self-Evolving Action-Conditioned Video Generation via Reflective Planning Agents Paper • 2603.08403 • Published May 21
UniVR: Thinking in Visual Space for Unified Visual Reasoning Paper • 2607.12800 • Published Jul 14 • 32
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published Jul 10 • 88
PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking Paper • 2607.00115 • Published Jun 30 • 13
The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Paper • 2606.24937 • Published Jun 22 • 19
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Paper • 2606.19534 • Published Jun 17 • 65
LSVOS 2025 Challenge Report: Recent Advances in Complex Video Object Segmentation Paper • 2510.11063 • Published Oct 13, 2025 • 1
RMP-SAM: Towards Real-Time Multi-Purpose Segment Anything Paper • 2401.10228 • Published Jan 18, 2024
RecTok: Reconstruction Distillation along Rectified Flow Paper • 2512.13421 • Published Dec 15, 2025 • 5
EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing Paper • 2512.11715 • Published Dec 12, 2025
WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World Paper • 2512.10958 • Published Dec 11, 2025 • 1
Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future Paper • 2512.16760 • Published Dec 18, 2025 • 15