This study was to search the possible usage of the Constructive-Item in the large scale test to enhance the educational efficiency of the test. To achieve this goal, the research developed the Computer Automated Scoring(CAS) program and validated the ...
This study was to search the possible usage of the Constructive-Item in the large scale test to enhance the educational efficiency of the test. To achieve this goal, the research developed the Computer Automated Scoring(CAS) program and validated the program. While developing the program, prior research was analyzed and revised through a preliminary inspection and with the consultation of experts, created this program by using Mathematical Markup Language (MathML).
Math experts and computer experts verified the propriety of the program, and on the whole, the experts acknowledged the evaluation as valid. As compared with the agreement between the average scored by four human raters (two teachers and two math experts) and those of the computer, the range of the agreement ratio of the true scores and the scores by computers was .70∼.90.
Generally, the agreement ratio between the true scores and the scores by the computer was perfect agreement frequency, 736 of total 1070 (68.79%), and included one score difference, of which the frequency was 855 (79.91%). In comparison with the prior research, the agreement ratio was a little low, but it was improved many times through the inspection of propriety. This research obtained excellent results in view of the first tryout.
In inspection of the scoring errors and subjective degree between the human raters, first, the coefficient of correlation between raters had a significance statistically second, in inspection of the confidence of the raters' reliability were different in accordance with propensity of them; third, in inspection of score error items there were 216.
Computers occasionally happen to correct the errors caused by human raters during manual operations, but the computer automated scoring may have several problems. First of all, it may not be able to recognize the answers that examiners made, using different method to solve the problems. Secondly, it may not be able to recognize the answers presented through the same formula or different forms. The third thing is that it may not be able to check the answer that included some wrong elements or missed some essential parts of the solution. Finally, it may not be able to check the letter answers, as it is not the computer language.
The limit of computer automated scoring is that it is difficult to evaluate the grades, especially the questions with higher points, due to the variety of content of answer sheets. According to the confirmation result of X2, X2=6. = 2(df = 2), it doesn't matter statistically at the level of .01, which means the accuracy of computer automated scoring was insufficient, according to the variety of solution types.
Post-research will need to study the ways that can support environmental weak points and invest in them continuously.