In a future society characterized by uncertainty and unpredictability, fostering problem-solving competence to respond flexibly to complex situations is essential. However, the issue of knowledge acquired in school remaining as inert knowledge, failin...
In a future society characterized by uncertainty and unpredictability, fostering problem-solving competence to respond flexibly to complex situations is essential. However, the issue of knowledge acquired in school remaining as inert knowledge, failing to transfer to real-world contexts, has been raised as a serious concern. Accordingly, many educational scholars have emphasized the importance of authentic problem solving, which enables students to effectively apply and utilize knowledge outside the school environment. Authentic problems resemble real-world issues due to their high ill-structured nature, authenticity, and openness; however, their high complexity often leads to difficulties in problem solving. To mitigate these difficulties, solving problems through collaboration is crucial. Recently, technological advancements have expanded the scope of collaboration to include human-AI collaboration, leading to an increasing number of cases where humans and AI form teams to solve problems. However, research on how to configure human-AI teams for effective learning remains insufficient. If learning proceeds without consideration of team composition, inefficient or ineffective interactions and learning processes may recur.
Therefore, this study aims to examine the effects on interaction and learning by manipulating the number of human participants (one, two, and three) in a context where humans solve authentic problems with AI, and to identify the effective group size for learning. The research questions are as follows: First, What are the differences in human-AI team interactions across different group sizes? Second, How do learning and problem-solving outcomes of human-AI teams differ by group size? Third, How do learners perceive the human-AI interaction and the group size in human–AI teaming?
To address these research questions, experimental research was conducted. The participants were 55 undergraduate and graduate students aged 18 or older, who were randomly assigned to three groups (1-person, 2-person, and 3-person groups). Pre-tests were conducted on factors that could affect interaction to ensure homogeneity among groups. Learners were assigned the task of 'writing a policy proposal to address economic slowdown caused by an aging population' and were allowed to use ChatGPT during the problem-solving process.
To analyze the quantity and patterns of human-AI interaction during the problem-solving process, this study utilized an interaction analysis framework consisting of prompting (information seeking, monitoring, result generation), AI response reviewing, AI response copying, outlining, as well as result drafting, reviewing, and revising. The problem solution writing and pre/post-knowledge tests were evaluated based on the developed rubric. To statistically identify differences between groups in the collected quantitative data, non-parametric tests such as the Kruskal-Wallis test, Mann–Whitney U, Wilcoxon signed-rank test, Friedman test, and Spearman correlation analysis, as well as Repeated Measures ANOVA, were used. In addition, interviews were conducted with all participants to investigate their perceptions of human-AI interaction and the impact of group size on learning and problem solving. The interview results were analyzed using thematic analysis.
The experimental results showed partial differences in human-AI interaction and learning depending on group size. According to descriptive statistics, the 1-person group spent more time on total prompting, result generation prompting, and copying AI responses in the first third of the problem-solving process compared to other groups. Also, the one-person group initiated result-sheet generation prompting earlier than the other two groups, and nearly half of the learners entered the assignment text verbatim into the AI to request a generated output. In contrast, the 2-person and 3-person groups spent considerable time in the first third refining outlines through human-to-human discussion, and the AI was also used for this purpose of outline refinement.
Epistemic Network Analysis (ENA) of interaction patterns also confirmed differences between groups. The 1-person group showed strong connections between AI answer copying and result-related behaviors in the early and middle phases, entering the main result drafting phase faster than other groups. The 2-person and 3-person groups showed strong connections between outline drafting and monitoring behaviors. The 2-person group showed strong connections between result generation prompting and monitoring behaviors entering the middle phase, indicating they used AI for result generation from the middle phase onwards. However, the 3-person group maintained stronger connections between outline drafting and monitoring until the middle phase compared to the 2-person group, resulting in the latest transition to the main result drafting phase. This was because the larger group size required more time to coordinate opinions among members. Additionally, an imbalance in the proportion of speaking time among learners appeared in the 3-person group. Considering the characteristic of multi-person groups where AI input prompts are decided through discussion and consensus, learners with low speaking proportions participated less in the prompt decision process, which may have limited their human-AI interaction.
Regarding the knowledge test, all three groups showed an increase in post-test scores compared to pre-test scores, with the 2-person group showing the largest change, but the difference between groups was not statistically significant. In the evaluation of problem solution
writing results, scores were highest in the order of 2-person, 3-person, and 1-person groups, with a statistically significant difference between the 2-person and 1-person groups. The 2-person group was effective in maximizing the benefits of multi-person groups—such as joint review of AI responses and solutions, mutual monitoring of opinions, and emotional exchange—while preventing issues of indiscriminate acceptance of AI and coordination costs. In learner interviews, the 2-person group was most frequently mentioned as the most effective group for learning and problem solving. Conversely, the 1-person group was perceived as unfavorable for learning and problem solving due to the occurrence of dependence on AI, which negatively affects learning, and the absence of a partner to jointly review AI responses. The 3-person group was appreciated for hearing diverse opinions, but concerns were identified regarding the time and effort required for coordination and the potential for free riders.
Meanwhile, most learners perceived AI as a passive tool that performs tasks only upon human request, rather than as a teammate. Positive aspects, such as AI aiding cognitive offloading and expansive thinking, were raised alongside negative aspects, such as inducing metacognitive laziness and standardized answers. As a result of correlation analysis between knowledge test changes/problem solution writing scores and time ratios by interaction behavior, in the 1-person group, higher time spent on total prompting and result generation prompting correlated with lower knowledge acquisition and problem-solving scores. This implies that uncritical reliance on AI can negatively affect learning and problem solving. However, such correlations did not appear in the 2-person and 3-person groups. In multi-person groups, human-to-human discussion appeared to prevent dependence on AI and uncritical acceptance of AI. This suggests that the same interaction behavior can have different effects on learning and problem solving depending on group size.
This study identified the effective group size for solving authentic problems using AI by analyzing interaction and learning according to group size. It holds theoretical significance in deriving the characteristics of group composition necessary to minimize the side effects and maximize the effects of AI. It also has practical significance in establishing the foundation for designing adaptive learning environments in educational settings by analyzing not just the learning results but also the process.