RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Construction of a Synthetic Dataset for Visual Question Answering on Infographic Images

    한글로보기

    https://www.riss.kr/link?id=T17570643

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
      • URL 복사
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Visual question answering (VQA) on infographics requires models to read dense text, interpret visual elements, and aggregate evidence across non-linear layouts. Existing benchmarks mainly focus on chart-centric or statistics-rich infographics, leaving poster-style, illustration-driven fact-sheet infographics underrepresented. This thesis introduces SQuADv2-VQA, a synthesized infographic VQA dataset built from SQuADv2. The dataset contains 20,233 high-resolution infographics and 199,279 question-answer (QA) pairs, including answerable and unanswerable questions inherited from the source corpus and supplementary synthetic QA items for training. We construct the dataset with a data synthesis framework in which a large language model generates structured infographic specifications grounded in source passages and QA annotations, a layout-guided generator renders the corresponding images, and a rule-based spatial reasoning engine produces supplementary QA items from the resulting layout metadata. We evaluate modern open-source large vision-language models on SQuADv2-VQA-test and related text-centric VQA datasets. The results show that the proposed dataset is challenging, particularly for complex layouts, longer answer spans, and unanswerable questions. We further show that fine-tuning on the proposed dataset improves test-set performance while maintaining comparable results on related benchmarks. Additional analysis suggests robustness under cross-generator and template-bias shifts. Overall, this work extends infographic VQA evaluation to an underrepresented visual style and provides a training resource for improving LVLM understanding of dense illustrated fact-sheet infographics.
    번역하기

    Visual question answering (VQA) on infographics requires models to read dense text, interpret visual elements, and aggregate evidence across non-linear layouts. Existing benchmarks mainly focus on chart-centric or statistics-rich infographics, lea...

    Visual question answering (VQA) on infographics requires models to read dense text, interpret visual elements, and aggregate evidence across non-linear layouts. Existing benchmarks mainly focus on chart-centric or statistics-rich infographics, leaving poster-style, illustration-driven fact-sheet infographics underrepresented. This thesis introduces SQuADv2-VQA, a synthesized infographic VQA dataset built from SQuADv2. The dataset contains 20,233 high-resolution infographics and 199,279 question-answer (QA) pairs, including answerable and unanswerable questions inherited from the source corpus and supplementary synthetic QA items for training. We construct the dataset with a data synthesis framework in which a large language model generates structured infographic specifications grounded in source passages and QA annotations, a layout-guided generator renders the corresponding images, and a rule-based spatial reasoning engine produces supplementary QA items from the resulting layout metadata. We evaluate modern open-source large vision-language models on SQuADv2-VQA-test and related text-centric VQA datasets. The results show that the proposed dataset is challenging, particularly for complex layouts, longer answer spans, and unanswerable questions. We further show that fine-tuning on the proposed dataset improves test-set performance while maintaining comparable results on related benchmarks. Additional analysis suggests robustness under cross-generator and template-bias shifts. Overall, this work extends infographic VQA evaluation to an underrepresented visual style and provides a training resource for improving LVLM understanding of dense illustrated fact-sheet infographics.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    인포그래픽에 대한 시각 질의응답은 모델이 밀집된 텍스트를 읽고, 시각적 요소를 해석하며, 비선형적인 레이아웃 전반에 걸쳐 증거를 종합할 것을 요구한다. 기존 벤치마크는 주로 차트 중심 또는 통계 정보가 풍부한 인포그래픽에 초점을 맞추고 있어, 포스터형 구조를 가지며 일러스트레이션이 중심이 되는 팩트시트 형태의 인포그래픽은 상대적으로 충분히 다루어지지 않았다. 본 논문에서는 SQuADv2를 기반으로 구축한 합성 인포그래픽 질의응답 데이터셋인 SQuADv2-VQA를 제안한다. 이 데이터셋은 20,233개의 고해상도 인포그래픽과 199,279개의 질의응답 쌍으로 구성되며, 원천 말뭉치로부터 가져온 답변 가능 질문과 답변 불가능 질문뿐만 아니라, 학습을 위해 추가적으로 생성한 합성 질의응답 문항을 포함한다. 본 연구에서는 데이터 합성 프레임워크를 통해 데이터셋을 구축한다. 먼저 대규모 언어 모델이 원문 지문과 질의응답 주석에 기반하여 구조화된 인포그래픽 명세를 생성하고, 레이아웃 기반 생성기가 이에 대응하는 이미지를 렌더링한다. 이후 규칙 기반 공간 추론 엔진이 생성된 레이아웃 메타데이터를 활용하여 추가 질의응답 문항을 생성한다. 제안 데이터셋의 난이도와 유용성을 검증하기 위해, 최신 오픈소스 대규모 시각-언어 모델을 SQuADv2-VQA-test와 관련 텍스트 중심 시각 질의응답 데이터셋에서 평가하였다. 실험 결과, 제안한 데이터셋은 특히 복잡한 레이아웃, 긴 정답 구간, 그리고 답변 불가능 질문에서 높은 난이도를 보이는 것으로 나타났다. 또한 제안 데이터셋을 활용한 파인튜닝은 관련 벤치마크에서 유사한 수준의 성능을 유지하면서도 테스트 세트 성능을 향상시키는 것으로 확인되었다. 추가 분석 결과, 제안 데이터셋은 생성기 변화와 템플릿 편향 변화가 존재하는 상황에서도 강건성을 보이는 것으로 나타났다. 종합하면, 본 연구는 기존에 충분히 다루어지지 않았던 시각적 스타일의 인포그래픽으로 시각 질의응답 평가 범위를 확장하고, 밀집된 정보를 포함한 일러스트레이션 기반 팩트시트 인포그래픽에 대한 대규모 시각-언어 모델의 이해 능력을 향상시키기 위한 학습 자원을 제공한다.
    번역하기

    인포그래픽에 대한 시각 질의응답은 모델이 밀집된 텍스트를 읽고, 시각적 요소를 해석하며, 비선형적인 레이아웃 전반에 걸쳐 증거를 종합할 것을 요구한다. 기존 벤치마크는 주로 차트 중...

    인포그래픽에 대한 시각 질의응답은 모델이 밀집된 텍스트를 읽고, 시각적 요소를 해석하며, 비선형적인 레이아웃 전반에 걸쳐 증거를 종합할 것을 요구한다. 기존 벤치마크는 주로 차트 중심 또는 통계 정보가 풍부한 인포그래픽에 초점을 맞추고 있어, 포스터형 구조를 가지며 일러스트레이션이 중심이 되는 팩트시트 형태의 인포그래픽은 상대적으로 충분히 다루어지지 않았다. 본 논문에서는 SQuADv2를 기반으로 구축한 합성 인포그래픽 질의응답 데이터셋인 SQuADv2-VQA를 제안한다. 이 데이터셋은 20,233개의 고해상도 인포그래픽과 199,279개의 질의응답 쌍으로 구성되며, 원천 말뭉치로부터 가져온 답변 가능 질문과 답변 불가능 질문뿐만 아니라, 학습을 위해 추가적으로 생성한 합성 질의응답 문항을 포함한다. 본 연구에서는 데이터 합성 프레임워크를 통해 데이터셋을 구축한다. 먼저 대규모 언어 모델이 원문 지문과 질의응답 주석에 기반하여 구조화된 인포그래픽 명세를 생성하고, 레이아웃 기반 생성기가 이에 대응하는 이미지를 렌더링한다. 이후 규칙 기반 공간 추론 엔진이 생성된 레이아웃 메타데이터를 활용하여 추가 질의응답 문항을 생성한다. 제안 데이터셋의 난이도와 유용성을 검증하기 위해, 최신 오픈소스 대규모 시각-언어 모델을 SQuADv2-VQA-test와 관련 텍스트 중심 시각 질의응답 데이터셋에서 평가하였다. 실험 결과, 제안한 데이터셋은 특히 복잡한 레이아웃, 긴 정답 구간, 그리고 답변 불가능 질문에서 높은 난이도를 보이는 것으로 나타났다. 또한 제안 데이터셋을 활용한 파인튜닝은 관련 벤치마크에서 유사한 수준의 성능을 유지하면서도 테스트 세트 성능을 향상시키는 것으로 확인되었다. 추가 분석 결과, 제안 데이터셋은 생성기 변화와 템플릿 편향 변화가 존재하는 상황에서도 강건성을 보이는 것으로 나타났다. 종합하면, 본 연구는 기존에 충분히 다루어지지 않았던 시각적 스타일의 인포그래픽으로 시각 질의응답 평가 범위를 확장하고, 밀집된 정보를 포함한 일러스트레이션 기반 팩트시트 인포그래픽에 대한 대규모 시각-언어 모델의 이해 능력을 향상시키기 위한 학습 자원을 제공한다.

    더보기

    목차 (Table of Contents)

    • 1 Introduction 1
    • 1.1 Background 1
    • 1.2 Motivation 3
    • 1.3 Contributions 5
    • 1.4 Thesis Outline 6
    • 1 Introduction 1
    • 1.1 Background 1
    • 1.2 Motivation 3
    • 1.3 Contributions 5
    • 1.4 Thesis Outline 6
    • 2 Related Work 7
    • 2.1 Visual Question Answering 7
    • 2.2 Visual Question Answering on Infographics 10
    • 2.3 Large Vision-Language Models 12
    • 2.4 Data Synthesis for Multimodal Understanding 15
    • 3 SQuADv2-VQA Dataset Construction 17
    • 3.1 Dataset Overview 17
    • 3.2 Data Synthesis Framework 20
    • 3.3 Dataset Statistics and Analysis 33
    • 4 Experiments 37
    • 4.1 Datasets 37
    • 4.2 Evaluation Metrics 39
    • 4.3 Experimental Setup and Training Configuration 41
    • 4.4 Experimental Results 42
    • 5 Conclusion and Future Work 57
    • 5.1 Conclusion 57
    • 5.2 Future Work 58
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼