Skip to main navigation Skip to search Skip to main content

High-fidelity text refinement for ControlNet-guided latent diffusion in document inpainting

Research output: Journal PublicationArticlepeer-review

Abstract

Blind document image inpainting (DII) aims to restore degraded document scans without prior knowledge of noise locations, yet existing methods either produce over-smoothed text or introduce pseudo-characters. We propose a novel high-fidelity, text-guided restoration framework based on a ControlNet-guided Latent Diffusion Model (CGLDM). First, we extract and refine noisy OCR outputs using a two-stage Vision Language Model (VLM) and Large Language Model (LLM) pipeline, leveraging both global document context and local text cues to deliver near ground-truth textual fidelity. Next, these refined text features, together with an initial visually restored image, condition a latent diffusion process that progressively denoises and reconstructs clean image latents. To suppress patch-wise background inconsistencies inherent in high-resolution processing, we introduce an explicit feature-alignment loss of diffusion model training that enforces agreement between the predicted features and the VAE-encoded features of the ground-truth image. Extensive experiments on the FUNSD-ZH dataset demonstrate that our approach outperforms state-of-the-art methods in OCR legibility and maintains competitive image-level fidelity.
Original languageEnglish
Article number133491
JournalExpert Systems with Applications
VolumeVolume 332, Part A
Publication statusPublished Online - 1 Jul 2026

Free Keywords

  • Blind document image inpainting
  • Text-guided document restoration
  • OCR correction
  • Latent diff usion models
  • ControlNet
  • Vision language models

Fingerprint

Dive into the research topics of 'High-fidelity text refinement for ControlNet-guided latent diffusion in document inpainting'. Together they form a unique fingerprint.

Cite this