文档视觉问题回答挑战2020

论文标题

文档视觉问题回答挑战2020

Document Visual Question Answering Challenge 2020

论文作者

Mathew, Minesh, Tito, Ruben, Karatzas, Dimosthenis, Manmatha, R., Jawahar, C. V.

论文摘要

本文介绍了文档的视觉问题回答挑战的结果，该挑战是“深度学习时代的文本和文档”的一部分，在CVPR 2020中。挑战引入了一个新问题 - 对文档图像的视觉问题回答。挑战包括两个任务。第一个任务涉及在单个文档图像上询问问题。另一方面，第二个任务设置为检索任务，其中问题是在图像集合上提出的。对于任务1，引入了一个新的数据集，其中包括50,000个问题 - S）对定义了12,767个文档图像。对于任务2，已经创建了另一个数据集，其中包括共享相同文档模板的14,362个文档图像的20个问题。

This paper presents results of Document Visual Question Answering Challenge organized as part of "Text and Documents in the Deep Learning Era" workshop, in CVPR 2020. The challenge introduces a new problem - Visual Question Answering on document images. The challenge comprised two tasks. The first task concerns with asking questions on a single document image. On the other hand, the second task is set as a retrieval task where the question is posed over a collection of images. For the task 1 a new dataset is introduced comprising 50,000 questions-answer(s) pairs defined over 12,767 document images. For task 2 another dataset has been created comprising 20 questions over 14,362 document images which share the same document template.

下载PDF全文

下载文献需遵守相关版权规定

论文标题