Even in a society where artificial intelligence is advancing rapidly, human evaluative judgment continues to be widely used in critical decision-making across various sectors. Within education, evaluative judgment has long been regarded as a core thin...
Even in a society where artificial intelligence is advancing rapidly, human evaluative judgment continues to be widely used in critical decision-making across various sectors. Within education, evaluative judgment has long been regarded as a core thinking skill that learners should develop. Nevertheless, systematic inquiry into measuring evaluative judgment competence is scarce. Although recent peer-assessment research has highlighted the need for such measurement in learning contexts, attempts to analyze evaluative judgment in a multifaceted manner remain limited.
In this respect, the present study sought to explore a range of methods for conducting multifaceted analyses of rating data from peer assessment and, through this effort, to lay a theoretical foundation for measuring evaluative judgment. Specifically, the study pursued three aims: first, to examine which aspects of evaluative judgment can be measured using currently available indicators, it integrated analytic indicators from expert-rater research and judgment analysis; second, to identify, as comprehensively as possible, indicators that can be used to investigate individual differences in evaluative judgment in peer-assessment research; and third, to provide empirical evidence that supports researchers’ selection and use of indicators that are aligned with their research purposes.
To achieve these aims, a literature review was conducted on person-level indicators drawn from three approaches that have examined rater expertise, cognitive biases, and judgment processes: (a) rater accuracy, (b) rater effects, and (c) the Lens Model Equation (LME). Based on interpretations in prior research, the indicators were categorized as subdimensions of evaluative judgment competence. This interpretive categorization of the indicators was organized around key assessment-design factors (raters, performances/products being rated, evaluative criteria, and rating scales). The review identified 30 applicable indicators: 14 from the rater-accuracy approach, 10 from the rater-effects approach, and 6 from the Lens-Model approach. Across approaches, the indicators were summarized into seven interpretive categories: (1) overall accuracy, (2) inconsistency, (3) discrimination accuracy of performance levels, (4) discrimination of criterion-difficulty, (5) severity, (6) rater centrality, and (7) other.
For the six primary categories (excluding other), evidence for convergent and criterion-related validity was examined using simulated data. To reflect conditions commonly observed in peer-assessment research, the number of evaluative criteria, rating design, and rater sample size were manipulated. The analyses investigated whether (a) rank-order correlations between the true evaluative-judgment parameters and the indicators, and (b) rank-order correlations among indicator pairs varied across conditions. Furthermore, using an empirical dataset rated by university students, rank-order relations among the full set of indicators were re-examined. To clarify the relationships among indicators, hierarchical cluster analysis was applied to the rank-correlation data to explore whether the theoretically proposed categories also emerged in the empirical data.
The results indicated that validity evidence was obtained for the following six aspects of evaluative judgment: overall distance-based accuracy, overall rank-order accuracy, performance-level discrimination accuracy, performance-level rank-order accuracy, criterion-difficulty discrimination, and severity. Some theoretical categories were further subdivided into two, because in the empirical data correlation-based indicators tended to form clusters distinct from variance- or distance-based indicators. Indicators targeting rater centrality and inconsistency did not yield consistent validity evidence, indicating the need for more rigorous investigation in future research. The indicator characteristics by category are summarized in a table in the Appendix.
Finally, validity evidence was markedly lower in the condition where each rater evaluated five performances (reflecting typical classroom peer assessment) than in the condition where each rater evaluated thirty performances (reflecting typical laboratory settings for theoretical investigation). However, for indicators in the criterion-difficulty discrimination category, increasing the number of evaluative criteria partially mitigated this decline. Furthermore, when rating scales were longer or when the true variance of performance quality was small, distance-based accuracy indicators showed reduced validity, suggesting sensitivity to raters’ tendencies to overestimate the dispersion of performances.
This study has significance in that it brought together analytic methods that have largely evolved independently under the common purpose of measuring evaluative judgment competence. Prior research has typically treated the three approaches separately, or only partially linked them. By focusing on interpreting the scattered indicators in terms of subdimensions of evaluative judgment competence and also empirically examining relations among indicators, the present study provides a basis for more refined future research on evaluative judgment. Practically, the proposed indicator set may reduce the methodological burden for peer-assessment researchers by presenting indicators by subdimension and by offering basic guidance for key assessment-design decisions (e.g., rater-group size, number of performances/products, assignment design, number of evaluative criteria, and rating-scale length).