Towards Pixel-level VLM Perception via Simple Points Prediction
Tianhui Song
sthui
AI & ML interests
None yet
Recent Activity
upvoted a paper 11 days ago
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs authored a paper 11 days ago
HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers upvoted a paper about 2 months ago
HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers