Current trends in science education significantly emphasize on
fostering scientific inquiry competencies through engaging
students in scientific practices. Scientific practices refer to the
processes of investigating phenomena and solving problems
sci...
Current trends in science education significantly emphasize on
fostering scientific inquiry competencies through engaging
students in scientific practices. Scientific practices refer to the
processes of investigating phenomena and solving problems
scientifically, while scientific inquiry competencies denote the
ability to actively engage in scientific practices. As the goals of
science education are reorganized around scientific inquiry
competencies based on scientific practices, to ensure alignment
with the curriculum, it is necessary to adjust assessment toward
measuring these competencies.
Physics is a discipline centered on evidence-based modeling
and logical reasoning. However, according to previous studies,
written examinations in physics—which take the largest portion
of school assessments in physics-tend to focus on the
application of formulas or mathematical calculation skills rather
than measuring complex scientific inquiry competencies.
Furthermore, such assessments are disproportionately weighted
toward specific elements of scientific inquiry competencies, such
as ‘Interpreting Data’ and ‘Drawing Conclusions’.
Prompted by these concerns, this study aims to explore how
scientific inquiry competencies can be meaningfully assessed
within the format of written examinations without being biased
toward specific elements of scientific inquiry competencies.
This study identified the assessment systems for scientific
inquiry competencies in major international college entrance
examinations including AP Physics, A-Level Physics, and IB DP
Physics—all of which are internationally recognized for their
credibility and emphasis on scientific inquiry. This study
analyzed the inquiry types and structural characteristics of test
items, through an analysis of curriculum documents, past exam
papers, and mark schemes.
This study established an analytical framework by
synthesizing multi-dimensional scientific inquiry competencies
through a literature review and expert review. The framework is
composed of three dimensions and nine elements: the ‘Initiating
Inquiry’ dimension (Asking Questions, Formulating Hypotheses),
the ‘Procedural Knowledge’ dimension (Planning
investigations, Obtaining Data, Interpreting Data, and Drawing
Conclusions), and the ‘Higher-order Thinking’ dimension
(Applying Scientific Models, Argumentation, and Evaluating
Alternative Explanations).
This study investigated how the elements of scientific inquiry
competencies being assessed varied depending on the inquiry
types and structural characteristics of test items within the three
international examinations.
AP test items were classified as the‘basic inquiry’ types,
characterized by a ‘linear structure’ that required students to
independently solve the test items through the entire scientific
inquiry process feasible in a high school laboratory. This
structure effectively assessed ‘Procedural Knowledge’
dimension, specifically providing a detailed assessment of
Planning Investigations, Obtaining Data, Interpreting Data, and
Drawing Conclusions. Since the AP program does not include a
separate practical assessment, it appears that written
examinations assess scientific inquiry competencies by simulating
entire experimental process. Therefore, students can learn the
fundamental procedure of scientific inquiry independently and
repeatedly. However, opportunities for students to construct
alternative explanations or engage in argumentative
communication remained limited, as the conclusions of these test
items often focused on confirming established physical laws.
A-Level test items were classified as the ‘applied inquiry’
types, characterized by a ‘linear structur’e that provided
students with opportunities to experience the complex inquiry
processes of actual scientists. This structure effectively
assessed the ‘Higher-order Thinking’ dimension, specifically
Applying Scientific Models, Argumentation, and Evaluating
Alternative Explanations. Although A-Level includes a separate
practical assessment (named Practical Endorsement), it only
verifies completion and is not reflected in the final grade. To fill
this gap in the assessment system, it is inferred that written
examinations assess scientific inquiry competencies by
sequentially simulating the scientific inquiry. It seems that
A-Level test items enable students to engage in inquiry similar
to that of actual scientists. Meanwhile, to alleviate cognitive load,
A-Level test items appear to focus on Planning Investigations
within specific segments. Consequently, there were limitations in
providing students with a proactive experience of planning the
entire investigations.
IB DP items were classified as the ‘applied inquiry’ types
characterized by a ‘modular structure.’ IB DP items provide an
authentic, scenario-based context akin to real-world scientific
inquiry, with independent sub-items assessing inquiry
competencies separately. This structure effectively assessed
both the competencies of‘Higher-order Thinking’ dimension
and Formulating Hypotheses. In the IB DP, students can practice
the linear scientific inquiry process through Internal Assessment
(IA), which is quantitatively reflected in their final grades.
Consequently, it is inferred that IB DP focus on assessing the
‘Higher-order Thinking’ dimension within complex contexts
that are challenging to implement in school laboratory.
Furthermore, if a student sets an incorrect initial hypothesis, the
linear structures of AP and A-Level test items make it difficult
to assess sub-items, yet IB DP test items are designed so that
responses to previous sub-items are independent of subsequent
ones. This ‘modular’ structure facilitates the assessment of
Formulating Hypotheses. However, assessing scientific inquiry
holistically remains limited, as competencies are assessed
separately rather than sequentially.
The main findings of this study are summarized as follows.
First, the inquiry types of the test items varied depending on the
specific scientific inquiry competencies targeted to assess. AP
utilized ‘basic inquiry,’ which was conducive to the precise
measurement of the ‘Procedural Knowledge’ dimension;
however, it was limited in assessing the ‘Higher-order
Thinking’ dimension due to its structure, which leads students
toward predetermined results. In contrast, A-Level and IB DP
utilized ‘applied inquiry.’ In these examinations, the
assessment of the ‘Procedural Knowledge’ dimension was
simplified, yet the ‘Higher-order Thinking’ dimension was
effectively assessed by reflecting authentic scientific practices.
Second, the structural characteristics of the test items varied
according to the assessment system, and these differences
determined the specific scientific inquiry competencies that could
be assessed. In the absence of a separate practical assessment,
AP adopted a ‘linear’ structure that simulates the sequential
process of scientific inquiry within the written examination.
Similarly, A-Level, which conducts practical assessments on a
pass/not classified basis (Practical Endorsement), utilized this
linear structure. Conversely, IB DP, which quantitatively assesses
practical work through IA, utilized a ‘modular’ structure that
selectively assesses specific competencies in written
examinations. This modular’ structure was found to be
particularly conducive to assessing Formulating Hypotheses.
As a tool for assessing scientific inquiry competencies, written
examinations were found to possess both potential and
limitations. Regarding their potential, this study confirmed that
written examinations are effective tools that complement the
constraints of school-based laboratory instruction, providing
students with opportunities to indirectly experience and be
assessed on authentic scientific inquiry. Furthermore, by
removing the burden of physical performance, written
examinations enable an assessment that focuses specifically on
the cognitive dimensions of scientific inquiry competencies. This
suggests that when it is difficult to evaluate Higher-order
Thinking through direct practical assessments, written
examinations can serve as a viable alternative for assessing
these competencies.
Regarding the limitations of written examinations, ‘Asking
Questions’ remained an underserved area across all three
examination systems. This is primarily attributed to the inherent
characteristics of written examinations, which must prioritize
scoring objectivity and efficiency within standardized large-scale
assessments. These findings suggest that concerted efforts are
necessary to develop innovative approaches that can address this
assessment gap in the future.
These limitations could be overcome through more
sophisticated design of test items. By adopting a structure that
independently assesses specific competencies—similar to the
modular structure of the IB DP—it is possible to address the
constraints in assessing ‘Asking Questions.’ This is expected
to assess a broader range of competencies, thereby enhancing
the validity of the assessment while preserving the fairness
inherent in written examinations.