As large language models (LLMs) continue to improve in performance, efforts to apply them for practical use across various domains have intensified. However, challenges remain due to the inherent variability in LLM responses and the issue of hallucina...
As large language models (LLMs) continue to improve in performance, efforts to apply them for practical use across various domains have intensified. However, challenges remain due to the inherent variability in LLM responses and the issue of hallucinations, necessitating systematic management approaches. These problems are particularly pronounced in the medical domain, where such systems have yet to gain the full trust of domain experts. In this study, we propose a self-confidence score based selective response process to address these issues. This process combines selective responding with the metacognitive metrics of LLMs, filtering only those responses with high confidence scores. Furthermore, by iteratively applying this process, we demonstrate improvements in the coverage–accuracy trade-off. On the KorMedMCQA dataset, our method achieved an accuracy of 90.85% and a coverage of 71.11% after four iterations. Compared to the non-iterative setting, accuracy decreased by 2.2%, but coverage increased by 82.5%. These results indicate that leveraging LLMs’ metacognitive cues can improve the trade-off between coverage and accuracy. Furthermore, this approach demonstrates practical applicability in the medical domain and shows the contribution of metacognition based methods to enhance reliability.