Yuyi Zhang

Yuyi Zhang

My name is Yuyi Zhang (张宇一). I am a Ph.D. student at South China University of Technology, working on document intelligence, computer vision, and multimodal learning. I am advised by Prof. Lianwen Jin at the Deep Learning and Vision Computing Lab (SCUT-DLVCLab).

My goal is to build intelligent systems that can accurately read, understand, generate, and restore complex visual documents.

My current research interests include:

  • Multimodal OCR Large Language Models
  • Document Intelligence and Visual Document Understanding
  • AIGC, Visual Text Generation and Editing
  • Historical Document Recognition and Restoration

Google Scholar · Updated September 2026

News

Education

South China University of Technology

Ph.D. Student, School of Electronic and Information Engineering

Advisor: Prof. Lianwen Jin

South China University of Technology

M.S. Student, Deep Learning and Vision Computing Lab

Advisor: Prof. Lianwen Jin

China University of Geosciences (Wuhan)

Bachelor’s Degree

Experience

Huawei

Research Intern

IntSig Information Co., Ltd. (合合信息)

Research Intern

Awards

Selected Publications [Full List]

UniHIR unified historical inscription restoration pipeline
Draft, Verify, Restore: Self-Refining Historical Inscription Restoration with a Unified MLLM
ACL 2026 MainMLLMRestoration

One unified MLLM iteratively drafts, verifies, and restores damaged historical inscriptions.

Inscriptions are invaluable cultural heritage, yet centuries of fractures, erosion, and oxidation have rendered many partially illegible. UniHIR is the first unified multimodal large language model for end-to-end historical inscription restoration. Its Draft-Guided Localization and Hierarchical Self-Refinement designs enable iterative damage localization, content prediction, and self-correction, while UHIRFactory and HIR-Bench support memory-efficient instruction tuning for high-resolution inputs. Experiments show superior text-restoration accuracy and appearance quality with consistent page-level typography and style.

Yuyi Zhang*, Junle Liu*, Peirong Zhang*, Jianliang Liu, Zhenhua Yang, Lianwen Jin✉

* Equal contribution.

Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2026

AutoHDR three-stage restoration pipeline
Reviving Cultural Heritage: A Novel Approach for Comprehensive Historical Document Restoration
ACL 2025 MainHistorical DocumentsRestoration

AutoHDR restores full historical pages through a historian-inspired three-stage workflow.

Historical documents suffer from tears, water erosion, and oxidation, while existing restoration methods are often limited to a single modality or small regions. This work introduces FPHDR, a full-page dataset containing 1,633 real and 6,543 synthetic images, and AutoHDR, a three-stage system combining OCR-assisted damage localization, vision-language context prediction, and patch-autoregressive appearance restoration. Its modular design supports human-machine collaboration. On severely damaged documents, AutoHDR raises OCR accuracy from 46.83% to 84.05%, and to 94.25% with human intervention.

Yuyi Zhang, Peirong Zhang, Zhenhua Yang, Pengyu Yan, Yongxin Shi, Pengwei Liu, Fengjun Guo, Lianwen Jin✉

Annual Meeting of the Association for Computational Linguistics (ACL Main), 2025

MegaHan97K dataset scale comparison
MegaHan97K: A Large-Scale Dataset for Mega-Category Chinese Character Recognition with over 97K Categories
PR 2025OCRDataset

A 97,455-category benchmark for mega-scale Chinese character recognition.

Mega-category Chinese character recognition remains underexplored because existing datasets cover only a fraction of the latest GB18030-2022 standard. MegaHan97K introduces a large-scale dataset spanning 97,455 character categories across handwritten, historical, and synthetic subsets. It is the first dataset to fully support GB18030-2022 and provides balanced samples to reduce long-tail bias. Extensive benchmarks reveal new challenges in storage, morphologically similar character recognition, and zero-shot learning, establishing a foundation for future mega-category OCR research.

Yuyi Zhang*, Yongxin Shi*, Peirong Zhang, Yixin Zhao, Zhenhua Yang, Lianwen Jin✉

* Equal contribution.

Pattern Recognition, 2025

Schematic overview of the HierCode hierarchical codebook
HierCode: A Lightweight Hierarchical Codebook for Zero-shot Chinese Text Recognition
PR 2025OCRZero-shot Recognition

A lightweight hierarchical codebook enables efficient zero-shot Chinese text recognition.

Chinese text recognition is challenged by intricate character structures, a vast vocabulary, and unseen characters that conventional one-hot encoding cannot represent. HierCode introduces a lightweight hierarchical codebook using multi-hot representations, binary-tree encoding, and prototype learning to capture shared radicals and structures. The representation supports zero-shot recognition of out-of-vocabulary characters and efficient line-level matching with visual features. Experiments on handwritten, scene, document, web, and ancient text benchmarks show state-of-the-art conventional and zero-shot recognition with fewer parameters and fast inference.

Yuyi Zhang*, Yuanzhi Zhu*, Dezhi Peng, Peirong Zhang, Zhenhua Yang, Zhibo Yang, Cong Yao, Lianwen Jin✉

* Equal contribution.

Pattern Recognition, Volume 158, 110963, 2025

PosterVerse workflow
PosterVerse: A Full-Workflow Framework for Commercial-Grade Poster Generation with HTML-Based Scalable Typography
AAAI 2026 OralAIGCPoster Generation

A full-workflow poster generator with accurate, editable, and scalable HTML typography.

Commercial poster generation requires both appealing graphics and precise, editable text, yet existing systems often provide incomplete workflows and unreliable typography. PosterVerse automates the process through blueprint creation with a fine-tuned language model, graphical background generation with customized diffusion models, and unified layout-text rendering using an MLLM-powered HTML engine. It also introduces PosterDNA, an HTML-based commercial poster dataset supporting scalable typography. Experiments demonstrate visually compelling posters with accurate text alignment and flexible, customizable layouts.

Junle Liu, Peirong Zhang, Yuyi Zhang, et al., Lianwen Jin✉

AAAI Conference on Artificial Intelligence (AAAI Oral), 2026

DiffHDR restoration results on damaged historical documents
Predicting the Original Appearance of Damaged Historical Documents
AAAI 2025 OralDiffusionRestoration

DiffHDR predicts and reconstructs the original appearance of damaged historical documents.

Historical documents contain invaluable cultural information but frequently suffer from missing characters, damaged paper, and ink erosion. This work formulates Historical Document Repair as the task of predicting a document's original appearance and introduces HDR28K, a dataset of 28,552 damaged-repaired image pairs with character-level annotations and diverse degradations. DiffHDR augments a diffusion model with semantic and spatial guidance plus a character perceptual loss for coherent content and appearance. It substantially outperforms previous methods on real damaged documents and also generalizes to document editing and text-block generation.

Zhenhua Yang, Dezhi Peng, Yongxin Shi, Yuyi Zhang, Chongyu Liu, Lianwen Jin✉

AAAI Conference on Artificial Intelligence (AAAI Oral), 2025

FontDiffuser generated Chinese character styles
FontDiffuser: One-Shot Font Generation via Denoising Diffusion with Multi-Scale Content Aggregation and Style Contrastive Learning
AAAI 2024DiffusionFont Generation

A one-shot diffusion framework generates complete font libraries from a style reference.

One-shot font generation aims to create a complete font library from a small number of style references while preserving each source character's content. FontDiffuser reframes font imitation as an image-to-image denoising diffusion process. Its Multi-scale Content Aggregation block combines global and local cues to preserve intricate strokes, while Style Contrastive Refinement disentangles and supervises font style through contrastive learning. Experiments show state-of-the-art generation across diverse characters and styles, particularly for complex glyphs and large style variations.

Zhenhua Yang, Dezhi Peng, Yuxin Kong, Yuyi Zhang, Cong Yao, Lianwen Jin✉

AAAI Conference on Artificial Intelligence, 2024

Open-Source Projects

UniHIR

A unified MLLM for self-refining historical inscription restoration through drafting, verification, and restoration.

MLLMRestorationACL 2026

AutoHDR

A full-page historical document restoration pipeline inspired by the workflow of expert historians.

Historical DocumentsRestoration

MegaHan97K

A large-scale dataset for Chinese character recognition with more than 97,000 character categories.

OCRDataset97K+ Classes

PosterVerse

A full-workflow framework for commercial-grade poster generation with scalable typography.

AIGCTypography

GPT-4V OCR Evaluation

A quantitative and in-depth evaluation of OCR capabilities in multimodal large models.

Multimodal OCREvaluation

Honors

2024

National Scholarship

Ministry of Education of China

2022

National Scholarship

Ministry of Education of China

Miscellaneous

Professional Service

Reviewer: Pattern Recognition · AAAI · ACM Multimedia (ACM MM)

Interests

Piano · Guitar · Playing in bands · Improvisational accompaniment · Football · Badminton

Contact & Collaboration

Let’s build intelligent visual document systems together.

I welcome research discussions and open-source collaboration in document intelligence, multimodal OCR, AIGC, and cultural heritage digitization.