As the perception that traditional statistical education does not sufficiently support students' statistical thinking due to its focus on technology and calculation spreads, statistical literacy has become an increasingly important educational goal. A...
As the perception that traditional statistical education does not sufficiently support students' statistical thinking due to its focus on technology and calculation spreads, statistical literacy has become an increasingly important educational goal. According to previous studies, citizens with statistical literacy must have the Ability of Selective Reading (ASR) and the Ability of Imaginative Reading (AIR). ASR can be defined as the ability of statistical producers to understand the process of deriving compressed statistical information by using statistical, mathematical, and contextual knowledge and to draw statistical inferences. On the other hand, AIR can be said to be the ability to imagine what information is missing in the selective reading process (e.g., whether there is a hidden bias or whether an alternative model or interpretation is possible).
Data modeling has been studied by many researchers as a way to educate elementary school students. Data structuring and representing data in data modeling, as well as informal inferential reasoning, can support the expression of ASR. Imagining the hidden aspects of phenomena in models, in reasoning, and in data can support the expression of AIR. Yet, the greatest difficulty in introducing data modeling activities for statistical literacy education is supporting the development of models for data variability. In this study, we propose MDS-based data modeling activities to support variability modeling among elementary school students. The purpose of this study is to examine whether these modeling activities can serve as a practical method for statistical literacy education through literature and to explore the characteristics of ASR and AIR that emerge when elementary school students participate in MDS-based data modeling activities.
In this investigation, the difficulties that elementary school students may encounter in data modeling activities were identified, and pedagogical strategies to address them were reviewed. Regarding the modeling of variability within the sample, elementary school students tended to judge variability based on the variability of data on the basis of ‘extreme values’ or ‘range’ and ‘bumpiness’ on the vertical axis. These trends hinder the understanding of variability within a sample as a structure of signal and noise. Hence, it is necessary to guide learners to pay attention to changes along the horizontal axis of the graph and to support them in measuring the extent of their deviation from the center. In the pedagogical strategy for this, it is proposed to introduce a new model, such as the hat plot, use interactive visualization, and develop a measure of variability. A hat plot can be a useful model for helping learners recognize variability as a change in the horizontal axis and measure the degree to which individual data are off-center. Interactive visualization supports learners in exploring the structure of a hat plot on their own and developing it as their own model. Lastly, by developing indicators that capture the distribution’s characteristics, data can be structured into signals and noise, and a model can be developed based on these structures.
Regarding sampling variability modeling, elementary school students tend to be attracted only to signals or noise. If the learner is attracted only to the signal, it can be argued that the given data provide complete information about the population; if attracted only to noise, it can be considered that the given data provide no information about the population at all. Furthermore, learners may have difficulty forming a hierarchical image of the sample. A pedagogical strategy is needed to support learners in forming hierarchical images of samples and in understanding the sampling process at different levels. In this strategy, it is proposed to experience simple probabilistic situations and to represent empirical sampling distributions as overlapping hat plots. By presenting the initial model in a simple probabilistic situation, the learner can clearly recognize the limitations of the initial model based on a deterministic or relativistic perspective. In addition, by exploring and expressing the empirical sampling distribution as overlapping hat plots, sampling variability can be structured into signals and noise.
In this study, an MDS-based data modeling activity for elementary school students was proposed by reorganizing these pedagogical strategies according to the framework of MDS, which consists of model-eliciting activity, model-exploration activity, and model-application activity.
This activity consisted of MDS 1, dealing with sample variability modeling, and MDS 2, dealing with sampling variability modeling, and was conducted for 10 weeks of instruction for 5th and 6th-grader-students of an elementary school. By collecting data from the focus group and analyzing them using thematic analysis and constant comparative analysis, the researcher confirmed the expression of ASR and AIR among elementary school students during these modeling activities.
The findings of this study regarding ASR expression patterns of students through MDS-based data modeling activities are as follows. ASR is expressed in a way that the structure of the model introduced in the model exploration activity promotes the restructuring of signals and noise, and that the model performs informal inference reasoning based on these structures. The structure of the hat plot, in which the average is the center, and the average value of the distance from the average to each point corresponds to the length of the crown, promoted the restructuring of signals and noise into the structure of data that deviated from the average and the average in how students perceived signals as the most frequent values. In addition, the structure of the empirical sampling distribution, in which the reciprocating distance of the sample becomes shorter as the sample size increases, promoted the restructuring of signals and noise by stabilizing the method of structuring only signals or noise and variations in the surrounding it. However, it is difficult to understand the role of data structure in inference only by exploring the structure of a new model. It can be seen that the new model's structure and the three stages of MDS promoted ASR expression.
The students' AIR expression patterns that appear through MDS-based data modeling activities are as follows. First, AIR can be expressed as a model that imagines the raw model and presents an alternative model to assess the appropriateness of the reasoning. To demonstrate, students were able to point out that the inference that ‘the effects of the two treatments are the same’ is based solely on the ‘average model’ and to criticize its appropriateness. This shows that AIR can be expressed when the learner can distinguish between inference and the raw model on which it is based. Second, AIR can be expressed by imagining the structure of raw data based on the model when explaining it with the existing model fails at the stage of model application. The intuitive image of the sampling distribution, formed from the empirical sampling distribution, supports structuring the sampling distribution into noise such as stabilizing signals and mutations in the surrounding area. However, some students had difficulty explaining the small sample change just by forming these images. In other words, if you pay attention only to the superficial form of the empirical sampling distribution, confusion may occur and hinder understanding of the data’s deep structure. Through these progresses, this study found that it was necessary to be able to imagine the structure of the raw data by confirming the simulation results, explaining the similarity to ball drawing activities, or asking ‘what is the appropriate sample size to predict the general effect of the treatment’.
Based on above results, it was confirmed that MDS-based data modeling activities not only address the difficulties encountered in elementary school students' variability modeling, but also promote the expression of ASR and AIR, key elements of statistical literacy. In particular, the introduction of new models such as the hat plot and the empirical sampling distribution and the three stages of MDS have improved how learners structure signals and noise in data. The restructuring of signals and noise was used as the basis for informal inference, allowing students to understand the roles that signals and noise in data play in statistical inference. In addition, in the judging the appropriateness of inference, AIR could be expressed by imagining a primitive model that serves as the basis for inference and proposing an alternative model. This study confirmed that ASR and AIR, the two elements of statistical literacy, are expressed interdependently, and that appropriate teacher questions and tasks are important so that the two elements can be expressed together to understand the data structure in depth. Still, this study is an exploratory case study involving a small number of students, and caution should be taken when generalizing these results. In the future, it is necessary to further explore how the elements of statistical literacy are expressed in the context of various grades and other tasks.
The results of this study contribute to the continuous exploration of the pedagogical potential of data modeling for elementary school students. Significantly, they show how MDS-based data modeling activities can support elementary school students' variability modeling and foster statistical literacy education.