RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Supporting Iterative Prompt Refinement for Text-to-Image Generative AI Systems = 텍스트 기반 이미지 생성형 인공지능 시스템의 반복적 프롬프트 지원

    한글로보기

    https://www.riss.kr/link?id=T17451047

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Text-to-image generative AI models have become increasingly accessible and capable of producing high-quality visuals from natural language prompts. However, users still face challenges when refining prompts through laborious trial-and-error. This work first reports a mixed-methods formative study with novice users (N = 22) that captured trial-level prompt behaviors, ratings, and think-aloud data across two text-to-image tasks. The analysis revealed three recurring challenges: users frequently shifted their goals in response to AI outputs yet hesitated to explore, had difficulty recognizing when they had already reached their favorite image, and often blamed themselves for poor wording while being unsure what to write next. Based on these findings, this study introduces iteraTree, a tree-based interactive interface that structures iterative text-to-image generation. iteraTree presents prompt-image history as a branching tree, with tools to revisit earlier results, compare pinned candidates in a Best Images panel, and receive GPT-based suggestions for refinement. A within-subjects comparative study (N = 14) with a conversational baseline showed that iteraTree made it easier for users to track iterations, compare outputs, decide when to stop, and formulate prompts, while perceived output quality remained comparable across conditions. These results suggest that interface-level support for history, comparison, and lightweight suggestions can better align generative AI systems with real-world creative workflows by making iterative exploration more transparent and controllable.
    번역하기

    Text-to-image generative AI models have become increasingly accessible and capable of producing high-quality visuals from natural language prompts. However, users still face challenges when refining prompts through laborious trial-and-error. This work...

    Text-to-image generative AI models have become increasingly accessible and capable of producing high-quality visuals from natural language prompts. However, users still face challenges when refining prompts through laborious trial-and-error. This work first reports a mixed-methods formative study with novice users (N = 22) that captured trial-level prompt behaviors, ratings, and think-aloud data across two text-to-image tasks. The analysis revealed three recurring challenges: users frequently shifted their goals in response to AI outputs yet hesitated to explore, had difficulty recognizing when they had already reached their favorite image, and often blamed themselves for poor wording while being unsure what to write next. Based on these findings, this study introduces iteraTree, a tree-based interactive interface that structures iterative text-to-image generation. iteraTree presents prompt-image history as a branching tree, with tools to revisit earlier results, compare pinned candidates in a Best Images panel, and receive GPT-based suggestions for refinement. A within-subjects comparative study (N = 14) with a conversational baseline showed that iteraTree made it easier for users to track iterations, compare outputs, decide when to stop, and formulate prompts, while perceived output quality remained comparable across conditions. These results suggest that interface-level support for history, comparison, and lightweight suggestions can better align generative AI systems with real-world creative workflows by making iterative exploration more transparent and controllable.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    텍스트 기반 이미지 생성형 인공지능 모델은 자연어 프롬프트만으로도 고품질 시각 콘텐츠를 생성할 수 있을 만큼 보편화되었지만, 사용자는 여전히 반복적인 시행착오를 거치며 프롬프트를 수정하는 과정에서 어려움을 겪는다. 본 연구는 먼저 초보 사용자 22명을 대상으로 한 혼합 방법 형성 연구를 수행하여 두 가지 텍스트 기반 이미지 생성 과제에서의 프롬프트 수정 행동, 프롬프트 시도 단위 측정, 질적 데이터를 수집하였다. 분석 결과, 사용자는 (1) 생성 결과에 따라 목표를 자주 변경하면서도 탐색하기를 주저하고, (2) 이미 최선의 결과에 도달했음에도 이를 인지하지 못한 채 반복을 이어 가며, (3) 프롬프팅에 확신이 없을 때 그 책임을 스스로의 표현 능력 탓으로 돌리는 경향이 있음을 확인하였다. 이러한 발견을 바탕으로, 본 연구는 반복적인 텍스트 기반 이미지 생성을 구조화하기 위한 트리 기반 인터랙티브 인터페이스 iteraTree를 제안한다. iteraTree는 프롬프트와 이미지 생성 과정을 분기 구조로 시각화하고, 이전 결과로 되돌아가거나 그 지점에서 분기할 수 있는 기능, 후보 이미지를 모아 비교하는 패널, GPT 기반 프롬프트 수정 제안을 제공한다. 이후 14명을 대상으로 한 비교 실험에서, iteraTree는 기준 시스템에 비해 반복 과정 추적, 출력 간 비교, 중단 시점 결정, 프롬프트 작성의 용이성 측면에서 더 높은 평가를 받았으며, 산출물의 품질 평정은 두 조건 간에 유사하게 나타났다. 이러한 결과는 생성 모델 자체를 변경하지 않더라도, 이력 관리와 비교, 가벼운 제안을 제공하는 인터페이스 설계를 통해 실제 창의적 작업 흐름에 더 잘 부합하도록 지원하고, 반복적 탐색 과정을 보다 투명하고 통제 가능하게 만들 수 있음을 시사한다.
    번역하기

    텍스트 기반 이미지 생성형 인공지능 모델은 자연어 프롬프트만으로도 고품질 시각 콘텐츠를 생성할 수 있을 만큼 보편화되었지만, 사용자는 여전히 반복적인 시행착오를 거치며 프롬프트...

    텍스트 기반 이미지 생성형 인공지능 모델은 자연어 프롬프트만으로도 고품질 시각 콘텐츠를 생성할 수 있을 만큼 보편화되었지만, 사용자는 여전히 반복적인 시행착오를 거치며 프롬프트를 수정하는 과정에서 어려움을 겪는다. 본 연구는 먼저 초보 사용자 22명을 대상으로 한 혼합 방법 형성 연구를 수행하여 두 가지 텍스트 기반 이미지 생성 과제에서의 프롬프트 수정 행동, 프롬프트 시도 단위 측정, 질적 데이터를 수집하였다. 분석 결과, 사용자는 (1) 생성 결과에 따라 목표를 자주 변경하면서도 탐색하기를 주저하고, (2) 이미 최선의 결과에 도달했음에도 이를 인지하지 못한 채 반복을 이어 가며, (3) 프롬프팅에 확신이 없을 때 그 책임을 스스로의 표현 능력 탓으로 돌리는 경향이 있음을 확인하였다. 이러한 발견을 바탕으로, 본 연구는 반복적인 텍스트 기반 이미지 생성을 구조화하기 위한 트리 기반 인터랙티브 인터페이스 iteraTree를 제안한다. iteraTree는 프롬프트와 이미지 생성 과정을 분기 구조로 시각화하고, 이전 결과로 되돌아가거나 그 지점에서 분기할 수 있는 기능, 후보 이미지를 모아 비교하는 패널, GPT 기반 프롬프트 수정 제안을 제공한다. 이후 14명을 대상으로 한 비교 실험에서, iteraTree는 기준 시스템에 비해 반복 과정 추적, 출력 간 비교, 중단 시점 결정, 프롬프트 작성의 용이성 측면에서 더 높은 평가를 받았으며, 산출물의 품질 평정은 두 조건 간에 유사하게 나타났다. 이러한 결과는 생성 모델 자체를 변경하지 않더라도, 이력 관리와 비교, 가벼운 제안을 제공하는 인터페이스 설계를 통해 실제 창의적 작업 흐름에 더 잘 부합하도록 지원하고, 반복적 탐색 과정을 보다 투명하고 통제 가능하게 만들 수 있음을 시사한다.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • List of Figures v
    • List of Tables vi
    • 1 Introduction 1
    • 1.1 Background 1
    • Abstract i
    • List of Figures v
    • List of Tables vi
    • 1 Introduction 1
    • 1.1 Background 1
    • 1.2 Research Motivation and Overview 2
    • 1.3 Organization of the Thesis 4
    • 2 Related Work 6
    • 2.1 Human-AI Interaction in Creative Tasks 6
    • 2.2 Opportunities and Barriers in Text-to-image Interaction 8
    • 2.3 Understanding Decision-Making in Iterative Exploration 10
    • 2.4 Prompt-Based Interaction and Refinement 12
    • 2.5 Summary and Research Gap 16
    • 3 Experiment 1 Formative Study 18
    • 3.1 Overview 18
    • 3.2 Method 19
    • 3.2.1 Participants 19
    • 3.2.2 Text-to-Image System 21
    • 3.2.3 Experimental Design 21
    • 3.2.4 Experimental Procedure 24
    • 3.2.5 Analysis Method 27
    • 3.3 Findings 28
    • 3.3.1 User Goal Evolution 30
    • 3.3.2 Decision-Making Difficulty 32
    • 3.3.3 Prompt Refinement Patterns 35
    • 4 System Design 38
    • 4.1 Design Goals 38
    • 4.2 User Scenario 40
    • 4.3 User Interface 42
    • 4.4 Implementation 45
    • 5 Experiment 2 Evaluation 49
    • 5.1 Overview 49
    • 5.2 Method 50
    • 5.2.1 Participants 50
    • 5.2.2 Experimental Design 51
    • 5.2.3 Experimental Procedure 52
    • 5.2.4 Measures 55
    • 5.2.5 Analysis Method 58
    • 5.3 Results 60
    • 5.3.1 Support for Exploration 63
    • 5.3.2 Support for Decision-Making 66
    • 5.3.3 Support for Prompt Refinement 70
    • 6 Discussion 75
    • 6.1 Interpretation of Key Results 75
    • 6.2 Design Implications 78
    • 6.3 Generalization of IteraTree 81
    • 6.4 Limitations and Future Work 83
    • 7 Conclusion 86
    • Bibliography 88
    • 한국어 초록 105
    • Acknowledgement 106
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼