-
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Paper • 2403.09611 • Published • 130 -
Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset
Paper • 2403.09029 • Published • 57 -
GiT: Towards Generalist Vision Transformer through Universal Language Interface
Paper • 2403.09394 • Published • 26
Xijia Tao
Cie1
AI & ML interests
Multimodal tool-calling agents, Diffusion large language models
Recent Activity
upvoted a paper about 11 hours ago
RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents upvoted a paper about 2 months ago
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models