Multimodal Vision AI Lab

Hohai University

The Multimodal Vision AI Lab focuses on research in multimodal artificial intelligence, vision-language models, document intelligence, OCR, handwriting recognition, compact multimodal models, and multimodal post-training.

Research Areas

Open Research

We develop open datasets, models, benchmarks, evaluation tools, and educational resources for multimodal vision and document intelligence research.

Projects

Projects, datasets, models, and benchmarks will be released progressively through this organization.

Links