Learning Video LLM with Streaming Speech Transcription at Scale (CVPR 2025)
Joya Chen PRO
chenjoya
AI & ML interests
Video LLM
Recent Activity
authored a paper 11 days ago
Rethinking Expressivity and Efficiency in Test-Time Training updated a dataset about 1 month ago
chenjoya/Live-WhisperX-526K upvoted a paper 2 months ago
Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning