Despite the increasing importance of processed foods in our diet, a limited number of nutrients are listed on food labels and others remain missing. This study aimed to propose a novel approach to estimate certain missing nutrient values in imported p...
Despite the increasing importance of processed foods in our diet, a limited number of nutrients are listed on food labels and others remain missing. This study aimed to propose a novel approach to estimate certain missing nutrient values in imported processed foods (IP foods) in Korea by integrating food matching and AI-based prediction. This study was conducted in two stages: (1) similar-food matching using a food matching algorithm and (2) estimation of missing nutrient values using an artificial intelligence (AI) model. In Stage 1, a quantitative food matching algorithm was developed to match IP foods with foods from the USDA Global Branded Food Products Database (GBFPD) and/or Open Food Facts data (US foods). Initially, US foods were selected as candidate matches if both the food name and brand name shared at least one common word with the corresponding IP foods, and if the Jaccard similarity, cosine similarity, and edit-distance–based string similarity scores calculated from food names exceeded predefined thresholds of 30, 50, and 60, respectively. Subsequently, only US foods with similar nutrient contents among the candidate matches were retained as final matches by applying criteria based on nutrient similarity, nutrient content ratios, and outlier thresholds. Specifically, cosine similarity, Euclidean distance, and a weighted average of the two were calculated as nutrient similarity measures using energy and eight mandatory nutrients (sodium, carbohydrate, sugars, total fat, trans fatty acids, saturated fatty acids, cholesterol, and protein) for each IP-US food pair. Only food pairs with scores ≥70 for all three similarity measures were retained. In addition, only US foods whose energy, carbohydrate, protein, and total fat contents were between 80% and <120% of those of the corresponding IP foods were retained, while outlier thresholds (≥99th percentile) were applied separately to the remaining nutrients (sugars, sodium, cholesterol, saturated fat, and trans fat). Thereafter, among the seven machine learning models, a Stacking Ensemble Classifier with the highest classification accuracy was used to categorize the matched IP-US food pairs into high-, moderate-, and low-similarity food (HSF, MSF, and LSF). As a result, among a total of 35,538 IP foods included in the import declaration data from the Ministry of Food and Drug Safety of Korea, covering IP foods between January 2023 and March 2024, 6,266 IP foods (approximately 17%) were successfully matched with 51,839 US foods. Of these matched food pairs, 59% were classified as HSF, 36% as MSF, and 4% as LSF. A comparison of the nutrient contents between the final matched food pairs revealed strong positive correlations for most nutrients. In particular, carbohydrate, protein, and total fat exhibited correlation coefficients greater than 0.9 across most food groups, regardless of matching class. Furthermore, the nutrient contents of matched food pairs were highly similar. For HSF, the mean differences between paired IP-US foods were close to zero across all food groups and nutrients (energy: 2.9 kcal; protein, carbohydrate, sugars, total fat, saturated fatty acids, and trans fatty acids: 0.0–0.6 g; cholesterol and sodium: 0.8–1.7 mg). These findings demonstrate that the food matching algorithm developed in Stage 1 can accurately identify US foods that are nutritionally similar to IP foods. In Stage 2, an AI model was developed to estimate missing values for fiber, calcium, and iron using the matched food pair data. The estimation strategies differed according to matching class: missing nutrient values were borrowed from the final matched food for HSF while they were predicted using a hybrid model integrating similarity-based and deep learning approaches for MSF and LSF. Fiber, calcium, and iron were selected as the target nutrients for estimation. These nutrients are not subject to mandatory labeling requirements in Korea and therefore exhibit very high missing rates (greater than 96%) in domestic food composition databases (FDCs). In contrast, their missing rates in the GBFPD were relatively low (less than 20%), making the estimation of these nutrients both necessary and feasible. The performance of the developed food matching algorithm and the missing nutrient value estimation models was evaluated using global FDCs. The results showed that the nutrient contents of the matched evaluation foods and their corresponding foods were highly similar. In particular, for HSF and MSF, the mean differences for eight nutrients, excluding sodium, were close to zero (energy: 0.7–2.1 kcal; protein, carbohydrate, sugars, total fat, saturated fatty acids, and trans fatty acids: 0.0–0.2 g; cholesterol: 0.1–0.3 mg). A comparison of the model-estimated and true values for fiber, calcium, and iron indicated that the mean differences were close to zero across most food groups. It is noteworthy that the mean differences for fiber and iron did not exceed 0.4 g and 0.2 mg, respectively, across all food groups. Distributional comparisons of the three nutrients further showed that the estimated and true values had very similar means and medians. However, in certain food groups, the maximum values and variances of the estimated values were marginally lower than the true values, indicating that the influence of extreme high values was reduced during the model estimation process. Finally, the developed model was applied to estimate missing values for fiber, calcium, and iron in IP foods. Based on the results of Stage 1, among the IP foods matched by the food matching algorithm, approximately 56% of the 6,098 foods (after excluding outliers for the three target nutrients) were matched with HSF, and their missing nutrient values were directly borrowed from the corresponding final matched US foods. Approximately 37% were matched with MSF and 4% with LSF, for which missing values were estimated using predictions from the hybrid model. In conclusion, the missing nutrient value estimation models developed in this study demonstrated substantial accuracy and consistency in estimating missing nutrient values, even in the absence of true reference values for the target nutrients, by effectively leveraging matched food pair data obtained through the food matching algorithm. The proposed approach integrates similar-food matching with AI-based prediction to enable efficient and factual estimation of missing nutrient values in IP foods in Korea.