Recent advancements in the field of medicine have undergone a revolutionary transformation, propelled by the proliferation of extensive patient datasets and the advancements in machine learning technology. Notably, research focusing on predicting the ...
Recent advancements in the field of medicine have undergone a revolutionary transformation, propelled by the proliferation of extensive patient datasets and the advancements in machine learning technology. Notably, research focusing on predicting the onset of specific diseases is prominently advancing, and among these, hepatocellular carcinoma (HCC) stands out as a noteworthy condition due to its high mortality rate relative to cancer incidence, coupled with an increasing risk. One major contributor to the occurrence of HCC is chronic hepatitis B infection, emphasizing the critical importance of early prediction. Previous studies faced challenges in constructing precise prediction models due to the diversity and high dimensionality of patient data. Furthermore, constraints imposed by security and personal information protection issues in claims data have limited the scope of research.
This study aims to overcome these limitations by leveraging data provided by the Health Insurance Review and Assessment Service to construct a machine learning algorithm predicting the occurrence of HCC among patients with chronic hepatitis B complications. The implemented algorithms include Penalized Logistic Regression, Random Forest, Extreme Gradient Boosting, and Support Vector Machine and the prediction results of each algorithm were compared. Moreover, the study utilized only diagnosed diseases from patient medical records as variables, and based on the importance of the variables of each algorithm, sought to identify the key diseases contributing to the prediction of HCC onset.