RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    매칭 알고리즘과 AI 기반 모델을 활용한 수입 가공식품 영양성분 결측치 추정 = Estimating Missing Nutrient Values in Imported Processed Foods Using a Matching Algorithm and AI-based Models

    한글로보기

    https://www.riss.kr/link?id=T17451449

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Despite the increasing importance of processed foods in our diet, a limited number of nutrients are listed on food labels and others remain missing. This study aimed to propose a novel approach to estimate certain missing nutrient values in imported processed foods (IP foods) in Korea by integrating food matching and AI-based prediction. This study was conducted in two stages: (1) similar-food matching using a food matching algorithm and (2) estimation of missing nutrient values using an artificial intelligence (AI) model. In Stage 1, a quantitative food matching algorithm was developed to match IP foods with foods from the USDA Global Branded Food Products Database (GBFPD) and/or Open Food Facts data (US foods). Initially, US foods were selected as candidate matches if both the food name and brand name shared at least one common word with the corresponding IP foods, and if the Jaccard similarity, cosine similarity, and edit-distance–based string similarity scores calculated from food names exceeded predefined thresholds of 30, 50, and 60, respectively. Subsequently, only US foods with similar nutrient contents among the candidate matches were retained as final matches by applying criteria based on nutrient similarity, nutrient content ratios, and outlier thresholds. Specifically, cosine similarity, Euclidean distance, and a weighted average of the two were calculated as nutrient similarity measures using energy and eight mandatory nutrients (sodium, carbohydrate, sugars, total fat, trans fatty acids, saturated fatty acids, cholesterol, and protein) for each IP-US food pair. Only food pairs with scores ≥70 for all three similarity measures were retained. In addition, only US foods whose energy, carbohydrate, protein, and total fat contents were between 80% and <120% of those of the corresponding IP foods were retained, while outlier thresholds (≥99th percentile) were applied separately to the remaining nutrients (sugars, sodium, cholesterol, saturated fat, and trans fat). Thereafter, among the seven machine learning models, a Stacking Ensemble Classifier with the highest classification accuracy was used to categorize the matched IP-US food pairs into high-, moderate-, and low-similarity food (HSF, MSF, and LSF). As a result, among a total of 35,538 IP foods included in the import declaration data from the Ministry of Food and Drug Safety of Korea, covering IP foods between January 2023 and March 2024, 6,266 IP foods (approximately 17%) were successfully matched with 51,839 US foods. Of these matched food pairs, 59% were classified as HSF, 36% as MSF, and 4% as LSF. A comparison of the nutrient contents between the final matched food pairs revealed strong positive correlations for most nutrients. In particular, carbohydrate, protein, and total fat exhibited correlation coefficients greater than 0.9 across most food groups, regardless of matching class. Furthermore, the nutrient contents of matched food pairs were highly similar. For HSF, the mean differences between paired IP-US foods were close to zero across all food groups and nutrients (energy: 2.9 kcal; protein, carbohydrate, sugars, total fat, saturated fatty acids, and trans fatty acids: 0.0–0.6 g; cholesterol and sodium: 0.8–1.7 mg). These findings demonstrate that the food matching algorithm developed in Stage 1 can accurately identify US foods that are nutritionally similar to IP foods. In Stage 2, an AI model was developed to estimate missing values for fiber, calcium, and iron using the matched food pair data. The estimation strategies differed according to matching class: missing nutrient values were borrowed from the final matched food for HSF while they were predicted using a hybrid model integrating similarity-based and deep learning approaches for MSF and LSF. Fiber, calcium, and iron were selected as the target nutrients for estimation. These nutrients are not subject to mandatory labeling requirements in Korea and therefore exhibit very high missing rates (greater than 96%) in domestic food composition databases (FDCs). In contrast, their missing rates in the GBFPD were relatively low (less than 20%), making the estimation of these nutrients both necessary and feasible. The performance of the developed food matching algorithm and the missing nutrient value estimation models was evaluated using global FDCs. The results showed that the nutrient contents of the matched evaluation foods and their corresponding foods were highly similar. In particular, for HSF and MSF, the mean differences for eight nutrients, excluding sodium, were close to zero (energy: 0.7–2.1 kcal; protein, carbohydrate, sugars, total fat, saturated fatty acids, and trans fatty acids: 0.0–0.2 g; cholesterol: 0.1–0.3 mg). A comparison of the model-estimated and true values for fiber, calcium, and iron indicated that the mean differences were close to zero across most food groups. It is noteworthy that the mean differences for fiber and iron did not exceed 0.4 g and 0.2 mg, respectively, across all food groups. Distributional comparisons of the three nutrients further showed that the estimated and true values had very similar means and medians. However, in certain food groups, the maximum values and variances of the estimated values were marginally lower than the true values, indicating that the influence of extreme high values was reduced during the model estimation process. Finally, the developed model was applied to estimate missing values for fiber, calcium, and iron in IP foods. Based on the results of Stage 1, among the IP foods matched by the food matching algorithm, approximately 56% of the 6,098 foods (after excluding outliers for the three target nutrients) were matched with HSF, and their missing nutrient values were directly borrowed from the corresponding final matched US foods. Approximately 37% were matched with MSF and 4% with LSF, for which missing values were estimated using predictions from the hybrid model. In conclusion, the missing nutrient value estimation models developed in this study demonstrated substantial accuracy and consistency in estimating missing nutrient values, even in the absence of true reference values for the target nutrients, by effectively leveraging matched food pair data obtained through the food matching algorithm. The proposed approach integrates similar-food matching with AI-based prediction to enable efficient and factual estimation of missing nutrient values in IP foods in Korea.
    번역하기

    Despite the increasing importance of processed foods in our diet, a limited number of nutrients are listed on food labels and others remain missing. This study aimed to propose a novel approach to estimate certain missing nutrient values in imported p...

    Despite the increasing importance of processed foods in our diet, a limited number of nutrients are listed on food labels and others remain missing. This study aimed to propose a novel approach to estimate certain missing nutrient values in imported processed foods (IP foods) in Korea by integrating food matching and AI-based prediction. This study was conducted in two stages: (1) similar-food matching using a food matching algorithm and (2) estimation of missing nutrient values using an artificial intelligence (AI) model. In Stage 1, a quantitative food matching algorithm was developed to match IP foods with foods from the USDA Global Branded Food Products Database (GBFPD) and/or Open Food Facts data (US foods). Initially, US foods were selected as candidate matches if both the food name and brand name shared at least one common word with the corresponding IP foods, and if the Jaccard similarity, cosine similarity, and edit-distance–based string similarity scores calculated from food names exceeded predefined thresholds of 30, 50, and 60, respectively. Subsequently, only US foods with similar nutrient contents among the candidate matches were retained as final matches by applying criteria based on nutrient similarity, nutrient content ratios, and outlier thresholds. Specifically, cosine similarity, Euclidean distance, and a weighted average of the two were calculated as nutrient similarity measures using energy and eight mandatory nutrients (sodium, carbohydrate, sugars, total fat, trans fatty acids, saturated fatty acids, cholesterol, and protein) for each IP-US food pair. Only food pairs with scores ≥70 for all three similarity measures were retained. In addition, only US foods whose energy, carbohydrate, protein, and total fat contents were between 80% and <120% of those of the corresponding IP foods were retained, while outlier thresholds (≥99th percentile) were applied separately to the remaining nutrients (sugars, sodium, cholesterol, saturated fat, and trans fat). Thereafter, among the seven machine learning models, a Stacking Ensemble Classifier with the highest classification accuracy was used to categorize the matched IP-US food pairs into high-, moderate-, and low-similarity food (HSF, MSF, and LSF). As a result, among a total of 35,538 IP foods included in the import declaration data from the Ministry of Food and Drug Safety of Korea, covering IP foods between January 2023 and March 2024, 6,266 IP foods (approximately 17%) were successfully matched with 51,839 US foods. Of these matched food pairs, 59% were classified as HSF, 36% as MSF, and 4% as LSF. A comparison of the nutrient contents between the final matched food pairs revealed strong positive correlations for most nutrients. In particular, carbohydrate, protein, and total fat exhibited correlation coefficients greater than 0.9 across most food groups, regardless of matching class. Furthermore, the nutrient contents of matched food pairs were highly similar. For HSF, the mean differences between paired IP-US foods were close to zero across all food groups and nutrients (energy: 2.9 kcal; protein, carbohydrate, sugars, total fat, saturated fatty acids, and trans fatty acids: 0.0–0.6 g; cholesterol and sodium: 0.8–1.7 mg). These findings demonstrate that the food matching algorithm developed in Stage 1 can accurately identify US foods that are nutritionally similar to IP foods. In Stage 2, an AI model was developed to estimate missing values for fiber, calcium, and iron using the matched food pair data. The estimation strategies differed according to matching class: missing nutrient values were borrowed from the final matched food for HSF while they were predicted using a hybrid model integrating similarity-based and deep learning approaches for MSF and LSF. Fiber, calcium, and iron were selected as the target nutrients for estimation. These nutrients are not subject to mandatory labeling requirements in Korea and therefore exhibit very high missing rates (greater than 96%) in domestic food composition databases (FDCs). In contrast, their missing rates in the GBFPD were relatively low (less than 20%), making the estimation of these nutrients both necessary and feasible. The performance of the developed food matching algorithm and the missing nutrient value estimation models was evaluated using global FDCs. The results showed that the nutrient contents of the matched evaluation foods and their corresponding foods were highly similar. In particular, for HSF and MSF, the mean differences for eight nutrients, excluding sodium, were close to zero (energy: 0.7–2.1 kcal; protein, carbohydrate, sugars, total fat, saturated fatty acids, and trans fatty acids: 0.0–0.2 g; cholesterol: 0.1–0.3 mg). A comparison of the model-estimated and true values for fiber, calcium, and iron indicated that the mean differences were close to zero across most food groups. It is noteworthy that the mean differences for fiber and iron did not exceed 0.4 g and 0.2 mg, respectively, across all food groups. Distributional comparisons of the three nutrients further showed that the estimated and true values had very similar means and medians. However, in certain food groups, the maximum values and variances of the estimated values were marginally lower than the true values, indicating that the influence of extreme high values was reduced during the model estimation process. Finally, the developed model was applied to estimate missing values for fiber, calcium, and iron in IP foods. Based on the results of Stage 1, among the IP foods matched by the food matching algorithm, approximately 56% of the 6,098 foods (after excluding outliers for the three target nutrients) were matched with HSF, and their missing nutrient values were directly borrowed from the corresponding final matched US foods. Approximately 37% were matched with MSF and 4% with LSF, for which missing values were estimated using predictions from the hybrid model. In conclusion, the missing nutrient value estimation models developed in this study demonstrated substantial accuracy and consistency in estimating missing nutrient values, even in the absence of true reference values for the target nutrients, by effectively leveraging matched food pair data obtained through the food matching algorithm. The proposed approach integrates similar-food matching with AI-based prediction to enable efficient and factual estimation of missing nutrient values in IP foods in Korea.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    식품영양성분 데이터베이스(DB)가 국민의 영양섭취 평가와 영양정책 수립의 핵심 기반 자료로 활용됨에 따라, 식품의 영양성분 함량 정보에 대한 중요성은 지속적으로 높아지고 있다. 그러나 가공식품의 경우 표시의무가 없는 영양성분의 함량 정보는 대부분 결측으로, 이러한 영양성분 결측치는 DB의 활용을 저해하는 주요 요인이다. 이에 본 연구에서는 우리나라에 수입되는 가공식품(이하 수입식품)의 영양성분 결측치 추정을 목적으로, 유사식품 매칭과 AI 예측을 통합한 새로운 결측치 추정 방법을 제안하였다. 본 연구는 1단계의 식품 매칭 알고리즘을 활용한 유사식품 매칭과 2단계의 유사식품 기반 AI 모델을 활용한 영양성분 결측치 추정으로 수행되었다. 1단계에서는 수입식품의 영양성분 결측치를 추정하기 위하여, 국외 가공식품 DB(USDA Global Branded Food Products Database 및 Open Food Facts)로부터 유사한 국외 가공식품(이하 국외식품)을 매칭하는 식품 매칭 알고리즘을 개발하였다. 개발한 식품 매칭 알고리즘을 이용하여 텍스트 유사도 기반의 후보 유사식품을 탐색하고 영양성분 함량 유사도 기반의 최종 유사식품을 매칭하였다. 우선 수입식품과 식품명 및 브랜드명이 모두 한 단어 이상 일치하면서, 식품명으로 산출한 자카드 유사도, 코사인 유사도, 편집거리 기반 문자열 유사도가 각각 30, 50. 60점 이상인 국외식품을 후보 매칭식품으로 탐색하였다. 이후 영양성분 함량의 유사도, 함량 비율, 이상치 기준을 적용하여 후보 유사식품 중 영양성분 함량이 유사한 식품만을 최종 매칭하였다. 이를 위해 우리나라 가공식품의 표시대상 영양성분 8종(나트륨, 탄수화물, 당류, 지방, 트랜스지방산, 포화지방산, 콜레스테롤, 단백질) 및 에너지를 대상으로, 코사인 유사도와 유클리드 거리 및 두 점수를 가중평균한 영양성분 종합 유사도를 산출하여 세 점수(코사인 유사도, 유클리드 유사도, 종합 유사도)가 모두 70점 이상인 식품만을 매칭하였다. 이후 매칭된 각 수입식품 대비 국외식품의 에너지, 탄수화물, 단백질, 지방의 함량 비율이 모두 80% 이상 120% 미만인 식품만을 매칭하였으며, 나머지 영양성분(당류, 나트륨, 콜레스테롤, 포화지방, 트랜스지방) 각각에 대한 이상치 기준(99백분위수 이상)을 적용하였다. 이후 일곱 가지 머신러닝 모델 중 분류 정확도가 가장 높은 Stacking Ensemble Classifier를 활용하여 매칭된 수입식품과 국외식품의 매칭 등급을 유사도가 높은 순서대로 유사도 상위식품(high-similarity food), 유사도 중위식품(moderate-similarity food), 유사도 하위식품(low-similarity food)으로 분류하였다. 그 결과, 연구에 활용한 식품의약품안전처의 수입식품 신고서 데이터(2023년 1월부터 2024년 3월까지 우리나라에 수입된 가공식품 정보 포함)의 수입식품 총 35,538개 중, 약 17%에 해당하는 6,266개 수입식품에 대해 51,839개의 국외식품이 최종 매칭되었으며, 이 중 59%가 유사도 상위식품, 36%가 유사도 중위식품, 4%가 유사도 하위식품으로 분류되었다. 최종 매칭된 수입식품과 국외식품의 영양성분(에너지 및 표시대상 8종) 함량을 비교한 결과, 대부분의 영양성분에서 높은 양의 상관관계가 확인되었다. 특히 탄수화물, 단백질, 지방은 대부분의 식품군에서 매칭 등급에 상관없이 0.9 이상의 높은 상관계수를 보였다. 또한 매칭된 수입식품과 국외식품의 영양성분 함량 역시 매우 유사하여, 유사도 상위식품의 경우 모든 식품군 및 영양성분에서 두 식품의 평균 함량 차이가 0에 근접하였다(에너지: 2.9 kcal, 단백질, 탄수화물, 당류, 지방, 포화지방산, 트랜스지방산: 0.0~0.6 g, 콜레스테롤, 나트륨: 0.8~1.7 mg). 이러한 결과는 1단계에서 개발한 식품 매칭 알고리즘이 우리나라 수입식품과 유사한 국외식품을 높은 정확도로 매칭할 수 있음을 입증하며, 2단계 ‘유사식품 기반 AI 모델을 활용한 결측치 추정’의 학습데이터 구축 및 예측 정확도 향상의 기반이 되었다. 이후 2단계에서는 1단계의 식품 매칭 알고리즘을 통해 확보된 매칭결과 데이터를 기반으로 결측 영양성분의 함량을 추정하는 AI 모델을 개발하였다. 매칭 등급별 특성을 고려하여 결측치 추정 방법을 차별화하였으며, 최종적으로는 개발한 유사식품 기반 영양성분 결측치 추정모델을 적용하여 우리나라 수입식품의 영양성분 결측치를 추정하였다. 결측치 추정 대상 영양성분은 식이섬유, 칼슘, 철로 선정하였는데, 이 세 영양성분은 우리나라에서 가공식품 의무표시 대상이 아니므로 결측률이 매우 높지만(96% 이상) 본 연구에서 활용한 국외 가공식품 DB에서의 결측률은 상대적으로 낮아(20% 미만), 현실적으로 결측치 추정이 필요하면서도 가능하였다. 개발한 영양성분 결측치 추정모델을 통해 매칭 등급별 특성을 고려한 최적의 결측치 추정 방법을 다음과 같이 적용하였다. 유사도 상위식품의 경우 최종 매칭된 국외식품의 영양성분 함량을 그대로 차용하여 효율성을 확보하였다. 유사도 중·하위식품의 경우에는 세 가지 예측모델—① 유사도 기반 회귀모델(Similarity Model), ② 심층학습 예측모델(DL Model), ③ 두 모델을 결합한 하이브리드 메타 예측모델(Hybrid Model)—을 개발하여 성능을 비교하고, 이 중 예측 정확도가 가장 뛰어난 하이브리드 메타 예측모델을 적용하여 영양성분 결측치를 예측하였다. 일부 국외식품(평가대상 식품)을 대상으로 본 연구에서 개발한 식품 매칭 알고리즘과 영양성분 결측치 추정모델의 성능을 평가한 결과, 매칭된 평가대상 식품과 국외식품의 영양성분 함량은 매우 유사하였다. 특히, 유사도 상위식품과 유사도 중위식품은 나트륨을 제외한 영양성분 8종의 평균 함량 차이가 모두 0에 근접하였다(에너지: 0.7~2.1 kcal, 단백질, 탄수화물, 당류, 지방, 포화지방산, 트랜스지방산: 0.0~0.2 g, 콜레스테롤: 0.1~0.3 mg). 영양성분 결측치 추정모델을 이용하여 추정한 식이섬유·칼슘·철의 결측치 추정값과 실제값을 비교한 결과, 대부분의 식품군에서 평균 함량 차이가 0에 근접하였다. 특히 식이섬유와 철의 경우 모든 식품군에서 평균 함량 차이가 각각 0.4 g, 0.2 mg 이하로 추정값과 실제값의 차이가 근소하였다. 세 영양성분 함량의 분포를 비교한 결과, 결측치 추정값과 실제값의 평균과 중앙값은 매우 유사하였으나, 일부 식품군에서는 추정값의 최대값과 분산이 실제값보다 다소 낮아, 모델의 추정 과정에서 고함량 값의 영향이 완화된 분포 특성을 보였다. 최종적으로는 개발한 모델을 적용하여 우리나라 수입식품의 식이섬유, 칼슘, 철의 결측치를 추정하였다. 1단계의 연구 결과, 식품 매칭 알고리즘을 통해 매칭된 수입식품(6,266개) 중 세 영양성분의 이상치를 제외한 수입식품(6,098개)의 약 56%에 유사도 상위식품으로 분류된 국외식품이 매칭되었으므로 최종 매칭된 국외식품의 값을 차용하였으며, 약 37%는 유사도 중위식품, 4%는 유사도 하위식품으로 분류된 국외식품이 매칭되어 하이브리드 메타 예측모델의 예측값으로 결측치를 추정하였다. 결론적으로, 본 연구에서 개발한 유사식품 기반 영양성분 결측치 추정모델은 유사식품 매칭을 통해 확보된 국외식품 데이터를 활용함으로써, 결측치 추정 대상 영양성분의 실제 함량 정보가 부재한 상황에서도 대부분의 식품군에서 실제값과 추정값의 차이가 근소하고 상관계수가 0.8 이상에 달하는 등, 상당한 정확도와 일관성으로 영양성분 결측치를 추정할 수 있음을 실증적으로 확인하였다. 본 연구는 유사식품 정보 차용의 안정성과 AI 예측의 정밀성을 결합한 새로운 영양성분 결측치 추정의 접근법을 제안함으로써, 수입식품의 영양성분 함량 정보의 품질 향상에 기여할 수 있는 AI 기반 결측치 추정의 실질적 토대를 마련하였다.
    번역하기

    식품영양성분 데이터베이스(DB)가 국민의 영양섭취 평가와 영양정책 수립의 핵심 기반 자료로 활용됨에 따라, 식품의 영양성분 함량 정보에 대한 중요성은 지속적으로 높아지고 있다. 그러...

    식품영양성분 데이터베이스(DB)가 국민의 영양섭취 평가와 영양정책 수립의 핵심 기반 자료로 활용됨에 따라, 식품의 영양성분 함량 정보에 대한 중요성은 지속적으로 높아지고 있다. 그러나 가공식품의 경우 표시의무가 없는 영양성분의 함량 정보는 대부분 결측으로, 이러한 영양성분 결측치는 DB의 활용을 저해하는 주요 요인이다. 이에 본 연구에서는 우리나라에 수입되는 가공식품(이하 수입식품)의 영양성분 결측치 추정을 목적으로, 유사식품 매칭과 AI 예측을 통합한 새로운 결측치 추정 방법을 제안하였다. 본 연구는 1단계의 식품 매칭 알고리즘을 활용한 유사식품 매칭과 2단계의 유사식품 기반 AI 모델을 활용한 영양성분 결측치 추정으로 수행되었다. 1단계에서는 수입식품의 영양성분 결측치를 추정하기 위하여, 국외 가공식품 DB(USDA Global Branded Food Products Database 및 Open Food Facts)로부터 유사한 국외 가공식품(이하 국외식품)을 매칭하는 식품 매칭 알고리즘을 개발하였다. 개발한 식품 매칭 알고리즘을 이용하여 텍스트 유사도 기반의 후보 유사식품을 탐색하고 영양성분 함량 유사도 기반의 최종 유사식품을 매칭하였다. 우선 수입식품과 식품명 및 브랜드명이 모두 한 단어 이상 일치하면서, 식품명으로 산출한 자카드 유사도, 코사인 유사도, 편집거리 기반 문자열 유사도가 각각 30, 50. 60점 이상인 국외식품을 후보 매칭식품으로 탐색하였다. 이후 영양성분 함량의 유사도, 함량 비율, 이상치 기준을 적용하여 후보 유사식품 중 영양성분 함량이 유사한 식품만을 최종 매칭하였다. 이를 위해 우리나라 가공식품의 표시대상 영양성분 8종(나트륨, 탄수화물, 당류, 지방, 트랜스지방산, 포화지방산, 콜레스테롤, 단백질) 및 에너지를 대상으로, 코사인 유사도와 유클리드 거리 및 두 점수를 가중평균한 영양성분 종합 유사도를 산출하여 세 점수(코사인 유사도, 유클리드 유사도, 종합 유사도)가 모두 70점 이상인 식품만을 매칭하였다. 이후 매칭된 각 수입식품 대비 국외식품의 에너지, 탄수화물, 단백질, 지방의 함량 비율이 모두 80% 이상 120% 미만인 식품만을 매칭하였으며, 나머지 영양성분(당류, 나트륨, 콜레스테롤, 포화지방, 트랜스지방) 각각에 대한 이상치 기준(99백분위수 이상)을 적용하였다. 이후 일곱 가지 머신러닝 모델 중 분류 정확도가 가장 높은 Stacking Ensemble Classifier를 활용하여 매칭된 수입식품과 국외식품의 매칭 등급을 유사도가 높은 순서대로 유사도 상위식품(high-similarity food), 유사도 중위식품(moderate-similarity food), 유사도 하위식품(low-similarity food)으로 분류하였다. 그 결과, 연구에 활용한 식품의약품안전처의 수입식품 신고서 데이터(2023년 1월부터 2024년 3월까지 우리나라에 수입된 가공식품 정보 포함)의 수입식품 총 35,538개 중, 약 17%에 해당하는 6,266개 수입식품에 대해 51,839개의 국외식품이 최종 매칭되었으며, 이 중 59%가 유사도 상위식품, 36%가 유사도 중위식품, 4%가 유사도 하위식품으로 분류되었다. 최종 매칭된 수입식품과 국외식품의 영양성분(에너지 및 표시대상 8종) 함량을 비교한 결과, 대부분의 영양성분에서 높은 양의 상관관계가 확인되었다. 특히 탄수화물, 단백질, 지방은 대부분의 식품군에서 매칭 등급에 상관없이 0.9 이상의 높은 상관계수를 보였다. 또한 매칭된 수입식품과 국외식품의 영양성분 함량 역시 매우 유사하여, 유사도 상위식품의 경우 모든 식품군 및 영양성분에서 두 식품의 평균 함량 차이가 0에 근접하였다(에너지: 2.9 kcal, 단백질, 탄수화물, 당류, 지방, 포화지방산, 트랜스지방산: 0.0~0.6 g, 콜레스테롤, 나트륨: 0.8~1.7 mg). 이러한 결과는 1단계에서 개발한 식품 매칭 알고리즘이 우리나라 수입식품과 유사한 국외식품을 높은 정확도로 매칭할 수 있음을 입증하며, 2단계 ‘유사식품 기반 AI 모델을 활용한 결측치 추정’의 학습데이터 구축 및 예측 정확도 향상의 기반이 되었다. 이후 2단계에서는 1단계의 식품 매칭 알고리즘을 통해 확보된 매칭결과 데이터를 기반으로 결측 영양성분의 함량을 추정하는 AI 모델을 개발하였다. 매칭 등급별 특성을 고려하여 결측치 추정 방법을 차별화하였으며, 최종적으로는 개발한 유사식품 기반 영양성분 결측치 추정모델을 적용하여 우리나라 수입식품의 영양성분 결측치를 추정하였다. 결측치 추정 대상 영양성분은 식이섬유, 칼슘, 철로 선정하였는데, 이 세 영양성분은 우리나라에서 가공식품 의무표시 대상이 아니므로 결측률이 매우 높지만(96% 이상) 본 연구에서 활용한 국외 가공식품 DB에서의 결측률은 상대적으로 낮아(20% 미만), 현실적으로 결측치 추정이 필요하면서도 가능하였다. 개발한 영양성분 결측치 추정모델을 통해 매칭 등급별 특성을 고려한 최적의 결측치 추정 방법을 다음과 같이 적용하였다. 유사도 상위식품의 경우 최종 매칭된 국외식품의 영양성분 함량을 그대로 차용하여 효율성을 확보하였다. 유사도 중·하위식품의 경우에는 세 가지 예측모델—① 유사도 기반 회귀모델(Similarity Model), ② 심층학습 예측모델(DL Model), ③ 두 모델을 결합한 하이브리드 메타 예측모델(Hybrid Model)—을 개발하여 성능을 비교하고, 이 중 예측 정확도가 가장 뛰어난 하이브리드 메타 예측모델을 적용하여 영양성분 결측치를 예측하였다. 일부 국외식품(평가대상 식품)을 대상으로 본 연구에서 개발한 식품 매칭 알고리즘과 영양성분 결측치 추정모델의 성능을 평가한 결과, 매칭된 평가대상 식품과 국외식품의 영양성분 함량은 매우 유사하였다. 특히, 유사도 상위식품과 유사도 중위식품은 나트륨을 제외한 영양성분 8종의 평균 함량 차이가 모두 0에 근접하였다(에너지: 0.7~2.1 kcal, 단백질, 탄수화물, 당류, 지방, 포화지방산, 트랜스지방산: 0.0~0.2 g, 콜레스테롤: 0.1~0.3 mg). 영양성분 결측치 추정모델을 이용하여 추정한 식이섬유·칼슘·철의 결측치 추정값과 실제값을 비교한 결과, 대부분의 식품군에서 평균 함량 차이가 0에 근접하였다. 특히 식이섬유와 철의 경우 모든 식품군에서 평균 함량 차이가 각각 0.4 g, 0.2 mg 이하로 추정값과 실제값의 차이가 근소하였다. 세 영양성분 함량의 분포를 비교한 결과, 결측치 추정값과 실제값의 평균과 중앙값은 매우 유사하였으나, 일부 식품군에서는 추정값의 최대값과 분산이 실제값보다 다소 낮아, 모델의 추정 과정에서 고함량 값의 영향이 완화된 분포 특성을 보였다. 최종적으로는 개발한 모델을 적용하여 우리나라 수입식품의 식이섬유, 칼슘, 철의 결측치를 추정하였다. 1단계의 연구 결과, 식품 매칭 알고리즘을 통해 매칭된 수입식품(6,266개) 중 세 영양성분의 이상치를 제외한 수입식품(6,098개)의 약 56%에 유사도 상위식품으로 분류된 국외식품이 매칭되었으므로 최종 매칭된 국외식품의 값을 차용하였으며, 약 37%는 유사도 중위식품, 4%는 유사도 하위식품으로 분류된 국외식품이 매칭되어 하이브리드 메타 예측모델의 예측값으로 결측치를 추정하였다. 결론적으로, 본 연구에서 개발한 유사식품 기반 영양성분 결측치 추정모델은 유사식품 매칭을 통해 확보된 국외식품 데이터를 활용함으로써, 결측치 추정 대상 영양성분의 실제 함량 정보가 부재한 상황에서도 대부분의 식품군에서 실제값과 추정값의 차이가 근소하고 상관계수가 0.8 이상에 달하는 등, 상당한 정확도와 일관성으로 영양성분 결측치를 추정할 수 있음을 실증적으로 확인하였다. 본 연구는 유사식품 정보 차용의 안정성과 AI 예측의 정밀성을 결합한 새로운 영양성분 결측치 추정의 접근법을 제안함으로써, 수입식품의 영양성분 함량 정보의 품질 향상에 기여할 수 있는 AI 기반 결측치 추정의 실질적 토대를 마련하였다.

    더보기

    목차 (Table of Contents)

    • 국문 초록 i
    • 목차 v
    • 표목차 viii
    • 그림목차 xii
    • 국문 초록 i
    • 목차 v
    • 표목차 viii
    • 그림목차 xii
    • I. 서론 1
    • 1. 연구 배경 및 필요성 1
    • 2. 연구 목적 9
    • 3. 연구 구성 9
    • II. 문헌고찰 11
    • 1. 가공식품 영양성분 DB의 구축 현황 11
    • 1) 국내 가공식품 영양성분 DB 11
    • 2) 국외 가공식품 영양성분 DB 19
    • 2. 결측치 추정 24
    • 1) 결측치 추정 24
    • 2) AI를 활용한 결측치 추정 27
    • 3. 식품영양성분 DB의 영양성분 결측치 추정 34
    • 1) 영양성분 결측치 추정 34
    • 2) 영양성분 결측치 추정 관련 연구 동향 38
    • 4. 결측치 추정 대상 영양성분 44
    • 1) 식이섬유 44
    • 2) 칼슘 46
    • 3) 철 48
    • III. 연구 방법 50
    • 1. 연구 방법의 개요 50
    • 2. 데이터 수집 및 전처리 55
    • 1) 데이터 수집 55
    • 2) 결측치 추정 대상 영양성분 선정 56
    • 3) 데이터 전처리 59
    • 3. 1단계: 식품 매칭 알고리즘을 활용한 유사식품 매칭 70
    • 1) 텍스트 유사도 기반의 후보 유사식품 탐색 72
    • 2) 영양성분 함량 유사도 기반의 최종 유사식품 매칭 77
    • 3) 머신러닝 모델을 활용한 매칭 등급 분류 81
    • 4. 2단계: 유사식품 기반 AI 모델을 활용한 결측치 추정 92
    • 1) 매칭 등급에 따른 결측치 추정 방법 탐색 93
    • 2) 매칭 등급에 따른 결측치 추정 방법 평가 102
    • 3) 매칭 등급에 따른 결측치 추정 방법 확정 105
    • 5. 식품 매칭 알고리즘과 영양성분 결측치 추정모델의 평가 및 적용 106
    • 1) 식품 매칭 알고리즘과 영양성분 결측치 추정모델의 평가 106
    • 2) 식품 매칭 알고리즘과 영양성분 결측치 추정모델의 적용 110
    • IV. 연구 결과 및 고찰 111
    • 1. 식품 매칭 알고리즘을 활용한 유사식품 매칭 111
    • 1) 머신러닝 모델의 성능 및 최종 선정 모델 111
    • 2) 수입식품과 국외식품 매칭 118
    • 2. 유사식품 기반 AI 모델을 활용한 영양성분 결측치 추정:
    • 유사도 중·하위식품 대상 세 가지 예측모델 비교 152
    • 1) 예측모델의 성능 152
    • 2) 영양성분 실제값과 예측값 160
    • 3) 최종 선정된 하이브리드 모델의 변수 중요도 170
    • 3. 식품 매칭 알고리즘과 영양성분 결측치 추정모델의 평가 및 적용 173
    • 1) 식품 매칭 알고리즘과 영양성분 결측치 추정모델의 평가: 국외식품 173
    • 2) 식품 매칭 알고리즘과 영양성분 결측치 추정모델의 적용: 수입식품 196
    • V. 결론 및 제언 203
    • 1. 요약 및 결론 203
    • 2. 제언 209
    • VI. 참고문헌 222
    • 부록 236
    • 부록 1. 수입식품과 국외식품의 기존 식품 분류 및 식품군별 식품수 237
    • 부록 2. 매칭된 수입식품과 국외식품의 매칭 등급 분류를 위한
    • 머신러닝 모델별 최적 파라미터 246
    • 부록 3. 개발한 식품 매칭 알고리즘을 이용한 수입식품-국외식품
    • 매칭 결과: 매칭 등급에 따른 식품군별 수입식품과 국외식품의
    • 영양성분 함량 차이 247
    • 부록 4. 개발한 유사식품 기반 영양성분 결측치 추정모델의
    • 평가 결과: 평가대상 국외식품의 영양성분 실제값과 결측치
    • 추정값 예시 258
    • 부록 5. 개발한 유사식품 기반 영양성분 결측치 추정모델의
    • 적용 결과: 수입식품의 영양성분 결측치 추정값 예시 264
    • Abstract 275
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼