RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    평가능력 측정을 위한 이론적·실증적 탐색 : 동료평가 맥락에서의 채점정확도, 채점자효과, 렌즈모형 = A Theoretical and Empirical Exploration into Measuring Evaluative Judgment Competence: Rating Accuracy, Rater Effects, and the Lens Model in Peer Assessment

    한글로보기

    https://www.riss.kr/link?id=T17451944

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Even in a society where artificial intelligence is advancing rapidly, human evaluative judgment continues to be widely used in critical decision-making across various sectors. Within education, evaluative judgment has long been regarded as a core thinking skill that learners should develop. Nevertheless, systematic inquiry into measuring evaluative judgment competence is scarce. Although recent peer-assessment research has highlighted the need for such measurement in learning contexts, attempts to analyze evaluative judgment in a multifaceted manner remain limited.
    In this respect, the present study sought to explore a range of methods for conducting multifaceted analyses of rating data from peer assessment and, through this effort, to lay a theoretical foundation for measuring evaluative judgment. Specifically, the study pursued three aims: first, to examine which aspects of evaluative judgment can be measured using currently available indicators, it integrated analytic indicators from expert-rater research and judgment analysis; second, to identify, as comprehensively as possible, indicators that can be used to investigate individual differences in evaluative judgment in peer-assessment research; and third, to provide empirical evidence that supports researchers’ selection and use of indicators that are aligned with their research purposes.
    To achieve these aims, a literature review was conducted on person-level indicators drawn from three approaches that have examined rater expertise, cognitive biases, and judgment processes: (a) rater accuracy, (b) rater effects, and (c) the Lens Model Equation (LME). Based on interpretations in prior research, the indicators were categorized as subdimensions of evaluative judgment competence. This interpretive categorization of the indicators was organized around key assessment-design factors (raters, performances/products being rated, evaluative criteria, and rating scales). The review identified 30 applicable indicators: 14 from the rater-accuracy approach, 10 from the rater-effects approach, and 6 from the Lens-Model approach. Across approaches, the indicators were summarized into seven interpretive categories: (1) overall accuracy, (2) inconsistency, (3) discrimination accuracy of performance levels, (4) discrimination of criterion-difficulty, (5) severity, (6) rater centrality, and (7) other.
    For the six primary categories (excluding other), evidence for convergent and criterion-related validity was examined using simulated data. To reflect conditions commonly observed in peer-assessment research, the number of evaluative criteria, rating design, and rater sample size were manipulated. The analyses investigated whether (a) rank-order correlations between the true evaluative-judgment parameters and the indicators, and (b) rank-order correlations among indicator pairs varied across conditions. Furthermore, using an empirical dataset rated by university students, rank-order relations among the full set of indicators were re-examined. To clarify the relationships among indicators, hierarchical cluster analysis was applied to the rank-correlation data to explore whether the theoretically proposed categories also emerged in the empirical data.
    The results indicated that validity evidence was obtained for the following six aspects of evaluative judgment: overall distance-based accuracy, overall rank-order accuracy, performance-level discrimination accuracy, performance-level rank-order accuracy, criterion-difficulty discrimination, and severity. Some theoretical categories were further subdivided into two, because in the empirical data correlation-based indicators tended to form clusters distinct from variance- or distance-based indicators. Indicators targeting rater centrality and inconsistency did not yield consistent validity evidence, indicating the need for more rigorous investigation in future research. The indicator characteristics by category are summarized in a table in the Appendix.
    Finally, validity evidence was markedly lower in the condition where each rater evaluated five performances (reflecting typical classroom peer assessment) than in the condition where each rater evaluated thirty performances (reflecting typical laboratory settings for theoretical investigation). However, for indicators in the criterion-difficulty discrimination category, increasing the number of evaluative criteria partially mitigated this decline. Furthermore, when rating scales were longer or when the true variance of performance quality was small, distance-based accuracy indicators showed reduced validity, suggesting sensitivity to raters’ tendencies to overestimate the dispersion of performances.
    This study has significance in that it brought together analytic methods that have largely evolved independently under the common purpose of measuring evaluative judgment competence. Prior research has typically treated the three approaches separately, or only partially linked them. By focusing on interpreting the scattered indicators in terms of subdimensions of evaluative judgment competence and also empirically examining relations among indicators, the present study provides a basis for more refined future research on evaluative judgment. Practically, the proposed indicator set may reduce the methodological burden for peer-assessment researchers by presenting indicators by subdimension and by offering basic guidance for key assessment-design decisions (e.g., rater-group size, number of performances/products, assignment design, number of evaluative criteria, and rating-scale length).
    번역하기

    Even in a society where artificial intelligence is advancing rapidly, human evaluative judgment continues to be widely used in critical decision-making across various sectors. Within education, evaluative judgment has long been regarded as a core thin...

    Even in a society where artificial intelligence is advancing rapidly, human evaluative judgment continues to be widely used in critical decision-making across various sectors. Within education, evaluative judgment has long been regarded as a core thinking skill that learners should develop. Nevertheless, systematic inquiry into measuring evaluative judgment competence is scarce. Although recent peer-assessment research has highlighted the need for such measurement in learning contexts, attempts to analyze evaluative judgment in a multifaceted manner remain limited.
    In this respect, the present study sought to explore a range of methods for conducting multifaceted analyses of rating data from peer assessment and, through this effort, to lay a theoretical foundation for measuring evaluative judgment. Specifically, the study pursued three aims: first, to examine which aspects of evaluative judgment can be measured using currently available indicators, it integrated analytic indicators from expert-rater research and judgment analysis; second, to identify, as comprehensively as possible, indicators that can be used to investigate individual differences in evaluative judgment in peer-assessment research; and third, to provide empirical evidence that supports researchers’ selection and use of indicators that are aligned with their research purposes.
    To achieve these aims, a literature review was conducted on person-level indicators drawn from three approaches that have examined rater expertise, cognitive biases, and judgment processes: (a) rater accuracy, (b) rater effects, and (c) the Lens Model Equation (LME). Based on interpretations in prior research, the indicators were categorized as subdimensions of evaluative judgment competence. This interpretive categorization of the indicators was organized around key assessment-design factors (raters, performances/products being rated, evaluative criteria, and rating scales). The review identified 30 applicable indicators: 14 from the rater-accuracy approach, 10 from the rater-effects approach, and 6 from the Lens-Model approach. Across approaches, the indicators were summarized into seven interpretive categories: (1) overall accuracy, (2) inconsistency, (3) discrimination accuracy of performance levels, (4) discrimination of criterion-difficulty, (5) severity, (6) rater centrality, and (7) other.
    For the six primary categories (excluding other), evidence for convergent and criterion-related validity was examined using simulated data. To reflect conditions commonly observed in peer-assessment research, the number of evaluative criteria, rating design, and rater sample size were manipulated. The analyses investigated whether (a) rank-order correlations between the true evaluative-judgment parameters and the indicators, and (b) rank-order correlations among indicator pairs varied across conditions. Furthermore, using an empirical dataset rated by university students, rank-order relations among the full set of indicators were re-examined. To clarify the relationships among indicators, hierarchical cluster analysis was applied to the rank-correlation data to explore whether the theoretically proposed categories also emerged in the empirical data.
    The results indicated that validity evidence was obtained for the following six aspects of evaluative judgment: overall distance-based accuracy, overall rank-order accuracy, performance-level discrimination accuracy, performance-level rank-order accuracy, criterion-difficulty discrimination, and severity. Some theoretical categories were further subdivided into two, because in the empirical data correlation-based indicators tended to form clusters distinct from variance- or distance-based indicators. Indicators targeting rater centrality and inconsistency did not yield consistent validity evidence, indicating the need for more rigorous investigation in future research. The indicator characteristics by category are summarized in a table in the Appendix.
    Finally, validity evidence was markedly lower in the condition where each rater evaluated five performances (reflecting typical classroom peer assessment) than in the condition where each rater evaluated thirty performances (reflecting typical laboratory settings for theoretical investigation). However, for indicators in the criterion-difficulty discrimination category, increasing the number of evaluative criteria partially mitigated this decline. Furthermore, when rating scales were longer or when the true variance of performance quality was small, distance-based accuracy indicators showed reduced validity, suggesting sensitivity to raters’ tendencies to overestimate the dispersion of performances.
    This study has significance in that it brought together analytic methods that have largely evolved independently under the common purpose of measuring evaluative judgment competence. Prior research has typically treated the three approaches separately, or only partially linked them. By focusing on interpreting the scattered indicators in terms of subdimensions of evaluative judgment competence and also empirically examining relations among indicators, the present study provides a basis for more refined future research on evaluative judgment. Practically, the proposed indicator set may reduce the methodological burden for peer-assessment researchers by presenting indicators by subdimension and by offering basic guidance for key assessment-design decisions (e.g., rater-group size, number of performances/products, assignment design, number of evaluative criteria, and rating-scale length).

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    인공지능이 급속도로 발달하는 현대 사회에도 인간의 평가능력은 사회 전반의 중요한 의사결정에 여전히 폭넓게 사용되고 있다. 교육에서도 평가는 학습자가 익혀야 할 핵심 사고력으로 간주되어 왔다. 그럼에도 평가능력 측정에 대한 본격적인 탐구를 찾아보기 어렵다. 학습 장면에서는 최근 동료평가 연구를 통해서 그 필요성이 제기되고 있지만, 아직은 평가능력을 다양하게 분석하려는 시도가 부족하다.
    본 연구는 동료평가 연구에서 채점자료를 다면적으로 분석할 수 있는 다양한 방법을 탐색하고, 이를 통해 평가능력 측정의 이론적 기초를 마련하고자 하였다. 보다 구체적으로는 다음의 세 가지를 목표로 삼았다. 첫째, 전문 채점자 연구와 판단분석 연구의 분석 지표들을 종합하여, 평가능력의 어떠한 측면들이 현재의 지표들로 측정 가능한지 확인하고자 하였다. 둘째, 동료평가 연구에서 개인차 분석에 활용 가능한 다양한 평가능력 지표를 광범위하게 탐색하고자 하였다. 셋째, 연구자의 관심에 적합한 지표를 사용하는 데 필요한 실증적 근거를 제공하고자 하였다.
    이를 위해 먼저, 평가자의 숙련도와 인지적 편향, 판단과정을 탐구해온 채점정확도, 채점자효과, 렌즈모형의 세 가지 접근법을 중심으로, 개인 수준 지표들에 대해 문헌고찰을 수행하였다. 그리고 선행 연구의 해석을 토대로 지표들을 범주화하고 이를 평가능력의 하위 측면으로 정의하였다. 지표의 해석적 범주화는 평가 설계 요인(평가자, 평가물, 채점준거, 채점척도)을 중심으로 이루어졌다. 먼저 문헌고찰 결과, 채점정확도 접근법에서 14개, 채점자효과 접근법에서 10개, 렌즈모형 접근법에서 6개 지표가 활용 가능한 것으로 나타났다. 그리고 전체 지표의 해석적 범주는 1) 총체적 정확성, 2) 비일관성, 3) 평가물 변별 정확성, 4) 채점준거 변별력, 5) 채점엄격성, 6) 중앙값편향, 7) 기타로 정리되었다.
    ‘기타’를 제외한 나머지 6가지 범주별로, 지표의 수렴타당도와 준거타당도를 시뮬레이션 자료를 통해 탐색하였다. 이때, 동료평가 연구에서 자주 사용되는 평가 조건을 반영하여, 채점준거 수와 채점설계, 평가자 집단 크기를 다르게 조작하면서, 평가능력 파라미터와 지표 간 순위상관과, 지표 쌍의 순위상관이 다르게 나타나는지 관찰하였다. 마지막으로 대학생이 실제로 평가한 자료를 통해 전체 지표 간 순위상관을 반복 탐색하였다. 전체 지표 간 관계를 보다 명료하게 확인하기 위해, 순위상관 자료에 계층적 군집분석을 적용하여, 이론적으로 가정한 범주가 실제 자료에서도 나타나는지 탐색하였다.
    분석 결과, 타당도 증거가 확보된 지표로 측정 가능한 평가능력은 다음의 6가지로 나타났다. 구체적으로는 ‘총체적 거리 정확성’, ‘총체적 서열화 정확성’, ‘평가물 변별 정확성’, ‘평가물 서열화 정확성’, ‘채점준거 난이도 변별력’, ‘채점엄격성’이었다. 이처럼 이론적 범주 중 일부가 두 가지로 세분되었는데, 이는 실제 자료에서 상관계수 형태의 지표가 분산 또는 거리 기반 지표와 구분되는 경향이 강했기 때문이다. ‘중앙값편향’과 ‘비일관성’의 지표들은 타당도 증거가 일관되게 관찰되지 않았으므로, 후속 연구를 통해 보다 엄밀한 탐색이 필요하다. 이상의 평가능력 범주별 지표와 그 특성에 대한 결과는 종합논의와 부록의 표로 요약하여 제시하였다.
    마지막으로 평가자당 5개의 평가물을 채점한 자료(실제 교실에서의 동료평가를 반영한 조건)는 30개를 채점한 자료(이론적 탐구를 위한 실험실 상황을 반영한 조건)에 비해 지표의 타당도가 현저히 떨어졌다. 다만, ‘채점준거 변별력’ 범주의 지표는 채점준거 수를 늘릴 경우, 일부분 보완될 수 있었다. 또한 채점척도가 길거나 평가물의 실제 분산이 작은 조건에서는 거리 기반 정확도 지표들이 평가자가 평가물 분산을 과대평가 경향에 민감하게 반응하여 타당도가 저하되는 한계를 보였다.
    본 연구는 거의 독립적으로 발전해 온 분석 방법들을, ‘평가능력 측정’이라는 목적 아래 포괄하였다는 점에서 학문적 의의가 있다. 선행 연구에서는 대체로 세 가지 접근법을 개별적으로 다루거나, 부분적으로만 연결하여 이해하는 시도만 확인되었다. 이처럼 산발적으로 논의되어온 방법을 평가능력의 하위 측면으로 해석하는 데에 초점을 두고 지표 간의 상관관계를 실증적으로 탐색함으로써, 향후 평가능력을 더욱 정교하게 탐구할 수 있는 이론적 기초를 마련하고자 하였다. 또한 실무적으로는 동료평가 연구자들이 방법론을 탐색하는 부담을 줄이는 데 기여할 수 있다. 평가능력을 측정하고자 하는 연구자가 필요한 분석법을 보다 효율적으로 탐색할 수 있도록, 본 연구는 ‘평가능력의 하위 측면별 분석 지표’를 제시하였다. 또한 연구 설계 요인(평가자 집단 크기, 평가물 전체 크기, 평가자에게 평가물을 배정하는 방식, 채점준거의 수, 채점척도의 길이 등)에서 참고할 수 있는 기초적인 지침을 제시하였다.
    번역하기

    인공지능이 급속도로 발달하는 현대 사회에도 인간의 평가능력은 사회 전반의 중요한 의사결정에 여전히 폭넓게 사용되고 있다. 교육에서도 평가는 학습자가 익혀야 할 핵심 사고력으로 간...

    인공지능이 급속도로 발달하는 현대 사회에도 인간의 평가능력은 사회 전반의 중요한 의사결정에 여전히 폭넓게 사용되고 있다. 교육에서도 평가는 학습자가 익혀야 할 핵심 사고력으로 간주되어 왔다. 그럼에도 평가능력 측정에 대한 본격적인 탐구를 찾아보기 어렵다. 학습 장면에서는 최근 동료평가 연구를 통해서 그 필요성이 제기되고 있지만, 아직은 평가능력을 다양하게 분석하려는 시도가 부족하다.
    본 연구는 동료평가 연구에서 채점자료를 다면적으로 분석할 수 있는 다양한 방법을 탐색하고, 이를 통해 평가능력 측정의 이론적 기초를 마련하고자 하였다. 보다 구체적으로는 다음의 세 가지를 목표로 삼았다. 첫째, 전문 채점자 연구와 판단분석 연구의 분석 지표들을 종합하여, 평가능력의 어떠한 측면들이 현재의 지표들로 측정 가능한지 확인하고자 하였다. 둘째, 동료평가 연구에서 개인차 분석에 활용 가능한 다양한 평가능력 지표를 광범위하게 탐색하고자 하였다. 셋째, 연구자의 관심에 적합한 지표를 사용하는 데 필요한 실증적 근거를 제공하고자 하였다.
    이를 위해 먼저, 평가자의 숙련도와 인지적 편향, 판단과정을 탐구해온 채점정확도, 채점자효과, 렌즈모형의 세 가지 접근법을 중심으로, 개인 수준 지표들에 대해 문헌고찰을 수행하였다. 그리고 선행 연구의 해석을 토대로 지표들을 범주화하고 이를 평가능력의 하위 측면으로 정의하였다. 지표의 해석적 범주화는 평가 설계 요인(평가자, 평가물, 채점준거, 채점척도)을 중심으로 이루어졌다. 먼저 문헌고찰 결과, 채점정확도 접근법에서 14개, 채점자효과 접근법에서 10개, 렌즈모형 접근법에서 6개 지표가 활용 가능한 것으로 나타났다. 그리고 전체 지표의 해석적 범주는 1) 총체적 정확성, 2) 비일관성, 3) 평가물 변별 정확성, 4) 채점준거 변별력, 5) 채점엄격성, 6) 중앙값편향, 7) 기타로 정리되었다.
    ‘기타’를 제외한 나머지 6가지 범주별로, 지표의 수렴타당도와 준거타당도를 시뮬레이션 자료를 통해 탐색하였다. 이때, 동료평가 연구에서 자주 사용되는 평가 조건을 반영하여, 채점준거 수와 채점설계, 평가자 집단 크기를 다르게 조작하면서, 평가능력 파라미터와 지표 간 순위상관과, 지표 쌍의 순위상관이 다르게 나타나는지 관찰하였다. 마지막으로 대학생이 실제로 평가한 자료를 통해 전체 지표 간 순위상관을 반복 탐색하였다. 전체 지표 간 관계를 보다 명료하게 확인하기 위해, 순위상관 자료에 계층적 군집분석을 적용하여, 이론적으로 가정한 범주가 실제 자료에서도 나타나는지 탐색하였다.
    분석 결과, 타당도 증거가 확보된 지표로 측정 가능한 평가능력은 다음의 6가지로 나타났다. 구체적으로는 ‘총체적 거리 정확성’, ‘총체적 서열화 정확성’, ‘평가물 변별 정확성’, ‘평가물 서열화 정확성’, ‘채점준거 난이도 변별력’, ‘채점엄격성’이었다. 이처럼 이론적 범주 중 일부가 두 가지로 세분되었는데, 이는 실제 자료에서 상관계수 형태의 지표가 분산 또는 거리 기반 지표와 구분되는 경향이 강했기 때문이다. ‘중앙값편향’과 ‘비일관성’의 지표들은 타당도 증거가 일관되게 관찰되지 않았으므로, 후속 연구를 통해 보다 엄밀한 탐색이 필요하다. 이상의 평가능력 범주별 지표와 그 특성에 대한 결과는 종합논의와 부록의 표로 요약하여 제시하였다.
    마지막으로 평가자당 5개의 평가물을 채점한 자료(실제 교실에서의 동료평가를 반영한 조건)는 30개를 채점한 자료(이론적 탐구를 위한 실험실 상황을 반영한 조건)에 비해 지표의 타당도가 현저히 떨어졌다. 다만, ‘채점준거 변별력’ 범주의 지표는 채점준거 수를 늘릴 경우, 일부분 보완될 수 있었다. 또한 채점척도가 길거나 평가물의 실제 분산이 작은 조건에서는 거리 기반 정확도 지표들이 평가자가 평가물 분산을 과대평가 경향에 민감하게 반응하여 타당도가 저하되는 한계를 보였다.
    본 연구는 거의 독립적으로 발전해 온 분석 방법들을, ‘평가능력 측정’이라는 목적 아래 포괄하였다는 점에서 학문적 의의가 있다. 선행 연구에서는 대체로 세 가지 접근법을 개별적으로 다루거나, 부분적으로만 연결하여 이해하는 시도만 확인되었다. 이처럼 산발적으로 논의되어온 방법을 평가능력의 하위 측면으로 해석하는 데에 초점을 두고 지표 간의 상관관계를 실증적으로 탐색함으로써, 향후 평가능력을 더욱 정교하게 탐구할 수 있는 이론적 기초를 마련하고자 하였다. 또한 실무적으로는 동료평가 연구자들이 방법론을 탐색하는 부담을 줄이는 데 기여할 수 있다. 평가능력을 측정하고자 하는 연구자가 필요한 분석법을 보다 효율적으로 탐색할 수 있도록, 본 연구는 ‘평가능력의 하위 측면별 분석 지표’를 제시하였다. 또한 연구 설계 요인(평가자 집단 크기, 평가물 전체 크기, 평가자에게 평가물을 배정하는 방식, 채점준거의 수, 채점척도의 길이 등)에서 참고할 수 있는 기초적인 지침을 제시하였다.

    더보기

    목차 (Table of Contents)

    • 제 1 장 서론 1
    • 1.1. 연구 배경 1
    • 1.2. 연구 문제 및 연구 목적 6
    • 1.3. 연구 의의 9
    • 1.4. 용어 정의 10
    • 제 1 장 서론 1
    • 1.1. 연구 배경 1
    • 1.2. 연구 문제 및 연구 목적 6
    • 1.3. 연구 의의 9
    • 1.4. 용어 정의 10
    • 1.5. 연구 범위 12
    • 1.6. 논문 구성 13
    • 제 2 장 이론적 배경 15
    • 2.1. 학습자의 평가능력과 관련 요인 17
    • 2.2. 채점자료를 분석하는 세 가지 접근법 22
    • 제 3 장 평가능력 지표에 대한 이론적 탐색 29
    • 3.1. 채점정확도 접근법 29
    • 3.2. 채점자효과 접근법 43
    • 3.3. 렌즈모형 접근법 53
    • 3.4. 세 가지 접근법의 종합 59
    • 3.5. 결론 65
    • 제 4 장 평가능력 지표에 대한 실증적 탐색 66
    • 4.1. 시뮬레이션 자료를 통한 탐색 68
    • 4.2. 실제 자료를 통한 탐색 110
    • 4.3. 논의 128
    • 제 5 장 종합 논의 140
    • 5.1. 주요 연구 결과 140
    • 5.2. 연구의 의의 143
    • 5.3. 연구의 제한점 및 후속 연구 제언 145
    • 참고문헌 147
    • 부록 A. 채점정확도 지표에 대한 문헌고찰 169
    • 부록 B. 4장 분석 결과 보충 표 177
    • 부록 C. 평가능력의 조작적 정의와 분석 지표 222
    • Abstract 226
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼