Han Wang

I am an ELLIS Ph.D. student in Computer Science, advised by Prof. Nicu Sebe at the University of Trento and Prof. Hilde Kuehne at the University of Tuebingen.

I previously worked as a Senior AI Engineer at ByteDance AML and Seed, where I built multimodal large language models for video understanding and industrial-scale training. My research interests include LLMs, MLLMs, video understanding, efficient visual token compression, encoder-free vision-language models, document understanding, and diffusion/flow LLMs.

Portrait of Han Wang

Research

I work on multimodal intelligence that can see, read, track, summarize, and reason over long videos and documents. A recurring thread in my work is making MLLMs both more capable and more practical: reducing visual tokens, removing heavy vision encoders, improving object-level video perception, and scaling training systems to industrial data and model sizes.

Education and Experience

Education

Ph.D. in Computer Science
2025 - 2028 expected
ELLIS Ph.D. student, University of Trento and University of Tuebingen.

M.A. in Electrical and Computer Engineering
2020 - 2023
Shanghai Jiao Tong University.

B.A. in Electronics Engineering; second degree in Mathematics
2016 - 2020
Beihang University.

Industrial Experience

Senior AI Engineer, ByteDance AML, Seed
2023 - 2025

  • Built MLLM systems for video summarization, relevance, grounding, tracking, and RL for MLLMs.
  • Developed billion-scale training frameworks and trained models with up to 1,024 GPUs.
  • Led VoRA2, an internal encoder-free MLLM trained on 800M samples.

AI Research Intern, ByteDance
2022 - 2023

AI Research Intern, AI Research Center, HIKVISION Inc.
2022

Featured Publications

Figure from Deep Trajectory Supervision
Deep Trajectory Supervision: Deep Supervision Strikes Back
Han Wang, et al.
ICML 2026

Deep trajectory supervision aligns intermediate representations with the intrinsic dynamics of the inference flow.

Figure from Dynamic-VLM
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
Han Wang, Yuxiang Nie, Yongjie Ye, Deng GuanYu, Yanjie Wang, Shuai Li, Haiyang Yu, Jinghui Lu, Can Huang
ICCV 2025

Dynamically compresses video tokens so VideoLLMs can balance frame-level details and long-context understanding.

Figure from ChineseVideoBench
ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering
Yuxiang Nie*, Han Wang*, Yongjie Ye, Haiyang Yu, Weitao Jia, Tao Zeng, Hao Feng, Xiang Fei, Yang Li, Xiaohui Lv, Guozhi Tang, Jingqun Tang, Jinghui Lu, Zehui Dai, Jiacong Wang, Dingkang Yang, An-Lan Wang, Can Huang
Preprint

A Chinese video question answering benchmark for evaluating MLLMs across language, culture, and video understanding tasks. * Equal contribution.

Figure from GLOMA
GLOMA: Global Video Text Spotting with Morphological Association
Han Wang, Yanjie Wang, Yang Li, Can Huang
ICLR 2025

Models video text spotting as global associations and uses Wasserstein distance for morphology-aware tracking.

Figure from Vision as LoRA
Vision as LoRA
Han Wang, Yongjie Ye, Bingru Li, Yuxiang Nie, Jinghui Lu, Jingqun Tang, Yanjie Wang, Can Huang
Preprint

Explores an encoder-free MLLM where visual capability is encoded through lightweight LoRA layers inside the LLM.

Figure from Elysium
Elysium: Exploring Object-level Perception in Videos via MLLMs
Han Wang, Yongjie Ye, Yanjie Wang, Yuxiang Nie, Can Huang
ECCV 2024

Introduces ElysiumTrack-1M and studies single-object tracking, referring tracking, and video referring expression generation with MLLMs.

Figure from PTSEFormer
PTSEFormer: Progressive Temporal-Spatial Enhanced Transformer Towards Video Object Detection
Han Wang, Jun Tang, Xiaodong Liu, Shanyan Guan, Rong Xie, Li Song
ECCV 2022

Aggregates temporal and spatial information progressively for video object detection.

Algorithm figure from Sports Tracking
A Large-scale Sports Tracking Dataset and Progressive Re-detection Based Sports Tracking
Han Wang, Xiaojun Zhou, Qinyu Xu, Huaqiang Ren, Rong Xie, Li Song
VCIP 2022

Builds a large-scale sports tracking dataset and proposes progressive re-detection for sports tracking.