PrismAlign: Prior-Steered Multi-View VLM Alignment for Hallucination-Robust Table OCR
In the authors' words
Table extraction suffers from frequent structural errors and semantic hallucinations. We propose PrismAlign, a multi-VLM framework aligning diverse visual perspectives to resolve ambiguity. It integrates priors of table logic to assess output plausibility, decoupling structural alignment from cell content alignment. A Bayesian decision strategy maximizes alignment accuracy by exploiting the correlation between extraction errors and computable rule violations. Evaluated on open-source and custom VLMs, PrismAlign reduces hallucinations and achieves state-of-the-art performance on OmniDocBench 1.5, as well as on the table category of CC-OCR and PureDocBench.
Appeared: Monday, September 21. arXiv. Preprint, not yet peer-reviewed.
Authors' comment: Accepted by EMNLP industry track