I join a steath-mode start-up and work as Research Scientist recently.
Researcher for Embodied AI, LLM/VLM Post-training, Social AI.
I join a steath-mode start-up and work as Research Scientist recently.
Researcher for Embodied AI, LLM/VLM Post-training, Social AI.
🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
[WACV 2024 Survey Paper] Multimodal Large Language Models for Autonomous Driving
[CVPR 2026] GazeAnywhere: Gaze Target Estimation Anywhere with Concepts
Unofficial implemention of lanenet model for real time lane detection Pytorch Version
[CVPR 2024] MAPLM: A Large-Scale Vision-Language Dataset for Map and Traffic Scene Understanding
Cognition-MLLM: Visual Cognition in Multimodal LLMs
JavaScript 9