Han Wang
I am an ELLIS Ph.D. student in Computer Science, advised by Prof. Nicu Sebe at the University of Trento and Prof. Hilde Kuehne at the University of Tuebingen.
I previously worked as a Senior AI Engineer at ByteDance AML and Seed, where I built multimodal large language models for video understanding and industrial-scale training. My research interests include LLMs, MLLMs, video understanding, efficient visual token compression, encoder-free vision-language models, document understanding, and diffusion/flow LLMs.
Research
I work on multimodal intelligence that can see, read, track, summarize, and reason over long videos and documents. A recurring thread in my work is making MLLMs both more capable and more practical: reducing visual tokens, removing heavy vision encoders, improving object-level video perception, and scaling training systems to industrial data and model sizes.
Education and Experience
Education
Ph.D. in Computer Science
2025 - 2028 expected
ELLIS Ph.D. student, University of Trento and University of Tuebingen.
M.A. in Electrical and Computer Engineering
2020 - 2023
Shanghai Jiao Tong University.
B.A. in Electronics Engineering; second degree in Mathematics
2016 - 2020
Beihang University.
Industrial Experience
Senior AI Engineer, ByteDance AML, Seed
2023 - 2025
- Built MLLM systems for video summarization, relevance, grounding, tracking, and RL for MLLMs.
- Developed billion-scale training frameworks and trained models with up to 1,024 GPUs.
- Led VoRA2, an internal encoder-free MLLM trained on 800M samples.
AI Research Intern, ByteDance
2022 - 2023
AI Research Intern, AI Research Center, HIKVISION Inc.
2022
Featured Publications
![]() |
Deep Trajectory Supervision: Deep Supervision Strikes Back
ICML 2026
Deep trajectory supervision aligns intermediate representations with the intrinsic dynamics of the inference flow. |
![]() |
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
ICCV 2025
Dynamically compresses video tokens so VideoLLMs can balance frame-level details and long-context understanding. |
![]() |
ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering
Preprint
A Chinese video question answering benchmark for evaluating MLLMs across language, culture, and video understanding tasks. * Equal contribution. |
![]() |
GLOMA: Global Video Text Spotting with Morphological Association
ICLR 2025
Models video text spotting as global associations and uses Wasserstein distance for morphology-aware tracking. |
![]() |
Vision as LoRA
Preprint
Explores an encoder-free MLLM where visual capability is encoded through lightweight LoRA layers inside the LLM. |
![]() |
Elysium: Exploring Object-level Perception in Videos via MLLMs
ECCV 2024
Introduces ElysiumTrack-1M and studies single-object tracking, referring tracking, and video referring expression generation with MLLMs. |
![]() |
PTSEFormer: Progressive Temporal-Spatial Enhanced Transformer Towards Video Object Detection
ECCV 2022
Aggregates temporal and spatial information progressively for video object detection. |
|
|
A Large-scale Sports Tracking Dataset and Progressive Re-detection Based Sports Tracking
VCIP 2022
Builds a large-scale sports tracking dataset and proposes progressive re-detection for sports tracking. |






