Document Type
Thesis - Open Access
Award Date
2026
Degree Name
Master of Science (MS)
Department / School
Electrical Engineering and Computer Science
First Advisor
Chulwoo Pack
Abstract
Open-domain multimodal document question answering requires retrieving a small set of answer-bearing pages from a large document collection before a vision-language model can generate an answer. Although modern retrieval systems may return hundreds or thousands of candidate pages, the answer model typically receives only the highest-ranked few. Consequently, relevant evidence may already be present in the candidate pool yet remain unavailable to the answer model because it is ranked below the reader cutoff. This thesis identifies this retrieval-to-reader gap as a page evidence promotion problem and proposes a framework for diagnosing and addressing it. First, page-level pseudo-supervision is derived from MultiModalQA evidence annotations inherited by M3DocVQA, exported page text, table and image evidence metadata, and the M3DocVQA document-ID/URL mapping. The final label construction uses adaptive exact evidence matching inside the annotated supporting documents, which makes it possible to distinguish failures of evidence discovery from failures to promote already-discovered evidence. Second, graph page-preserving propagation combines dense and sparse retrieval seeds while keeping page-level outputs for the reader. Third, a content-aware promotion model reorders candidate pages using rank, source, document/page structure, and question-page content evidence. On the M3DocVQA open-domain benchmark development set, the fraction of labeled questions with pseudo-page evidence among the first four pages (page@4) increases from 0.5558 for Dense (ColPali top-1000) and 0.6376 for graph page-preserving retrieval to 0.7715 with CAPP on GPP. A conservative transfer setting applies the learned promotion signal to four additional document retrieval datasets and gives small, cutoff-dependent changes over dense retrieval, with positive page@4 gains on three datasets and a near-neutral loss on one. When the top-4 pages are passed to the Qwen2-VL answer model, CAPP on GPP achieves the strongest downstream result, improving EM/F1 from 34.74/40.15 for Dense (ColPali top-1000) and 37.69/43.47 for GPP to 39.41/45.69. These findings show that broad evidence discovery alone is insufficient for open-domain multimodal document question answering. Explicitly promoting useful pages into the limited reader context improves both evidence coverage and answer quality, while retrieval-stage and answer-stage objectives remain related but not identical.
Publisher
South Dakota State University
Recommended Citation
Erfanshekooh, Abdolhossein, "Page Evidence Promotion for Open-Domain Multimodal Document Question Answering" (2026). Electronic Theses and Dissertations. 2155.
https://openprairie.sdstate.edu/etd2/2155