Home
People
Events
Research
Publications
Contact
News
1
CARE: Multi-Task Pretraining for Latent Continuous Action Representation in Robot Control
Recent advances in Vision-Language-Action (VLA) models have shown promise for robot control, but their dependence on action supervision …
Jiaqi Shi
,
Xulong Zhang
,
Xiaoyang Qu
,
Jianzong Wang
Cite
arXiv
IEEE
From Knowing to Doing Precisely: A General Self-Correction and Termination Framework for VLA Models
While vision-language-action (VLA) models for embodied agents integrate perception, reasoning, and control, they remain constrained by …
Wentao Zhang
,
Aolan Sun
,
Wentao Mo
,
Xiaoyang Qu
,
Yuxin Zheng
,
Jianzong Wang
Cite
arXiv
IEEE
Head-Aware Visual Cropping: Enhancing Fine-Grained VQA with Attention-Guided Subimage
Multimodal Large Language Models (MLLMs) show strong performance in Visual Question Answering (VQA) but remain limited in fine-grained …
Junfei Xie
,
Peng Pan
,
Xulong Zhang
Cite
arXiv
IEEE
MirrorTalk: Forging Personalized Avatars via Disentangled Style and Hierarchical Motion Control
Synthesizing personalized talking faces that uphold and highlight a speaker’s unique style while maintaining lip-sync accuracy …
Renjie Lu
,
Xulong Zhang
,
Xiaoyang Qu
,
Jianzong Wang
,
Shangfei Wang
Cite
arXiv
IEEE
Mita: A Hierarchical Multi-Agent Collaboration Framework with Memory-Integrated and Task Allocation
Recent advances in large language models (LLMs) have substantially accelerated the development of embodied agents. LLM-based …
Xiaojie Zhang
,
Jianhan Wu
,
Xiaoyang Qu
,
Jianzong Wang
Cite
arXiv
IEEE
Triage: Hierarchical Visual Budgeting for Efficient Video Reasoning in Vision-Language Models
Vision-Language Models (VLMs) face significant computational challenges in video processing due to massive data redundancy, which …
Anmin Wang
,
Nan Zhang
,
Wei Tao
,
Xiaoyang Qu
,
Guokuan Li
,
Jiguang Wan
,
Jianzong Wang
Cite
arXiv
IEEE
Vista: Scene-Aware Optimization for Streaming Video Question Answering under Post-Hoc Queries
Streaming video question answering (Streaming Video QA) poses distinct challenges for multimodal large language models (MLLMs), as …
Haocheng Lu
,
Nan Zhang
,
Wei Tao
,
Xiaoyang Qu
,
Guokuan Li
,
Jiguang Wan
,
Jianzong Wang
Cite
AAAI
Turbo-TTS: Enhancing Diffusion Model TTS with an Improved ODE Solver
This paper introduces Turbo-TTS, a novel diffusion-based model for text-to-speech (TTS) synthesis. Diffusion models leverage stochastic …
Xulong Zhang
,
Jiashu Wang
,
Xiaoyang Qu
,
Hui Tian
,
Jianzong Wang
Cite
Springer
EMO-RL: Emotion-Rule-Based Reinforcement Learning Enhanced Audio-Language Model for Generalized Speech Emotion Recognition
Although Large Audio-Language Models (LALMs) have exhibited outstanding performance in auditory understanding, their performance in …
Pengcheng Li~
,
Botao Zhao
,
Zuheng Kang
,
Junqing Peng
,
Xiaoyang Qu
,
Yayun He
,
Jianzong Wang
Cite
arXiv
ACL
Federated Domain Generalization with Domain-specific Soft Prompts Generation
Prompt learning has become an efficient paradigm for adapting CLIP to downstream tasks. Compared with traditional fine-tuning, prompt …
Jianhan Wu
,
Xiaoyang Qu
,
Zhangcheng Huang
,
Jianzong Wang
Cite
arXiv
theCVF
«
»
Cite
×