Multimodal QUD: Inquisitive Questions from Scientific Figures
Di cosa parla
Si cerca di far sì che i modelli formulino le domande che nascono guardando una figura in un articolo scientifico: domande non ancora risolte, rilevanti per le conclusioni e basate su ciò che la figura mostra. Il dataset MQUD raccoglie 1.250 domande da 56 articoli, di cui 708 annotate dagli autori. Senza adattamento, i modelli multimodali aperti tendono a porre domande risolvibili solo dalla figura; addestrati su MQUD, pongono domande più inquisitive e mirate al ruolo della figura.
Cosa permette di osservare
Consente di esplorare se e come i modelli che combinano testo e immagini individuano le domande più utili per interpretare i risultati nelle figure e come cambiano le loro domande se vengono addestrati su esempi mirati.
Dalla fonte
Discourse comprehension in complex documents often involves continuously posing and resolving Questions Under Discussion (QUDs). While QUD frameworks have so far focused on text, scientific literature is inherently multimodal: figures convey discourse goals distinct from their textual counterparts, thus invoking implicit questions that the surrounding text answers. In scientific discovery, knowing the right questions to ask is as important as knowing how to answer them, yet this capability remains largely absent in current models. In this work, we extend QUD to multimodal discourse in scientific literature, targeting questions evoked by figures that are (1) inquisitive, i.e., not resolved in the prior context; (2) salient, i.e., relevant to the paper's research claims and addressed later in the paper; (3) grounded in visual insights. To benchmark model capability to generate such questi…