
http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
박유선 고려대학교 정보보호대학원 2019 국내석사
공공기관의 사이버침해대응센터는 최근 폭발적으로 증가하는 로그와 이벤트 등의 위협정보를 처리하고 대응하는데 어려움이 있다. 이는 기하급수적으로 증가하는 정보를 처리하고 분석하는데 너무 많은 시간이 소요되기 때문이다. 하루에도 무수히 발생하는 보안 이벤트 공격 흔적을 빠르게 찾아내기 위해서는 고차원의 정보 저장·분석 능력이 필수적으로 요구된다. AI(Artificial Intelligence) 기술은 방대한 위협 정보의 분석 및 학습을 통해 공격의 탐지 및 예측이 가능하고, AI 기술을 통하여 효율적인 대응 전략을 수립하는데 도움이 될 것으로 예상하고 있다. 근래에 들어 기하급수적으로 증가하는 사이버 공격 경보 관련 데이터 분석의 정탐율을 높이고 진화하는 보안 위협에 효율적으로 대응하기 위한 방안으로 보안관제 부문에 기계학습(Machine Learning) 기술을 적용하려는 시도가 증가하고 있다. 기계학습 기반 보안관제시스템 운영을 위해서는 대용량의 데이터로부터 양질의 정보를 추출하고 분석해 최적의 학습 데이터를 만드는 것이 필요하다. 이를 위하여 보안전문가들이 직접 양질의 학습 데이터를 생성하고 선별하는 과정에 참여한다. 정보보호 분야의 기계학습 기반 보안 기술이 아직 초기 단계인 만큼, 기계학습 기반 학습 모델이 의미 있는 결과물을 창출하기 위해서는 기계학습 알고리즘에 적용하기 위한 학습 데이터를 선별하고, 원하는 결과를 얻기 위한 최적의 알고리즘을 선택·검증하는 것이 선결되어야 한다. 보안 담당자는 기계학습 알고리즘을 통해 걸러진 위험도가 높은 중요한 경보를 선제적으로 집중 분석함으로써 고도화된 보안 위협 대응에 보다 집중하고 기계학습 알고리즘에 적용할 또 다른 학습 데이터를 생성하며, 기계학습 기반 예측 모델에 더 많은 피드백을 줄 수 있다. 본 연구를 통하여 통합보안관제시스템의 경보 데이터와 침해사고 처리 내역에 기계학습 기술을 적용해 오탐지·정상탐지 여부를 예측하고, 이상행위로 판단되는 보안경보 이벤트의 위험도를 수치화해 이를 우선순위에 따라 처리할 수 있는 가이드를 제시하고자 한다.
기계학습 모형의 설명가능성에 관한 연구 : 미국 주택담보대출 자료를 중심으로
Machine learning is an area of artificial intelligence, which is known to have superior prediction power compared to standard econometrics approach. Standard econometrics approach are widely used in the social science area, including the real estate field. On the other hand, machine learning is a kind of black box model, which can not explain the cause of the results. Recent research on XAI (eXplainable Artificial Intelligence) in the field of machine learning has raised interest in the “Explainability” of the model. Explainability of machine learning provides an opportunity to open up the possibility of various studies in social science fields (real estate, finance, etc.), which have formerly used econometric technology or machine learning in prediction power. In this study, empirical studies were conducted by applying the explainability of machine learning to default risk of mortgage loans. Recently, domestic housing prices have been on the rise due to a decrease in housing supply and low interest rates in the Seoul metropolitan area, but the economic slowdown and the risk of a fall in housing prices could increase the default risk in mortgage loans and expand the scale. In addition, it has become important for financial institutions to more accurately measure Bank for International Settlements (BIS) ratios and to be recognized by regulators through self-help efforts to accurately measure risk weights for risky assets. However, prior studies related to mortgage loans have focused on explaining the factors associated with them based on standard econometrics approach models for the risk of delinquency, default or prepayment. Therefore, this study seeks to enhance the effectiveness of the model internally by measuring more delicate default risks based on machine learning, and it is also expected that this model can be used as an efficient regulatory compliance mechanism for default risks. In this study, it is analyzed mortgage data from Freddie Mac in the U.S. to derive a model of mortgage defaults based on machine learning (artificial neural network and random forest), and to explain the factors of default risk in the model as Partial Dependence Plot (PDP), marginal effects, and Shapley Additive Explanations (SHAP). In addition, by comparing the predictive power of standard econometrics approach and machine learning models, the machine learning model proved that not only the explanatory power but also the predictive power is better than the standard econometrics approach. First, results of comparing logistic regression as one of the standard econometrics techniques, and artificial neural network and random forest models as machine learning technique models, generally showed similar aspects in the coefficients of the logistic regression model, PDP and marginal effects, while the results were somewhat different in random forest. In the case of delinquency, both of artificial neural networks and random forest models showed that the months of delinquency (-), the total amount of overdue payments (-) and the credit rating (+) were in conflict with common sense, which is one of the interesting aspects of this study. It also sought to identify new potential variables through PDP, marginal effects, and SHAP analysis on datasets that added derived variables to existing independent variables. Variables related to housing price change rate in logistic regression and artificial neural network, and those related to the capital ratio in random forest, were judged to be meaningful. To compare predictive power, logistic regression was used as a econometrics technique and neural network and random forest were used as a machine learning technique. As a result, the machine learning model was found to be excellent in all indicators that verify predictive power such as Accuracy, F1 Score, and Area Under the Curve (AUC). By applying the machine learning-based model with the characteristics of “black box” form to mortgage default, this study verified the practical use potential of the explainability, as well as the predictability of the machine learning model. Furthermore, the machine learning model of this study is expected to serve as a guide for domestic financial institutions to manage the risks of mortgage loans on a machine learning basis. The results of this study could also be used to provide beneficial policy working-level guidelines when the government draws up policies to regulate household debt in the local real estate market. 기계학습은 인공지능의 한 분야로, 전통적 계량 기법에 비해 우수한 예측 능력을 갖는 것으로 알려져 있다. 전통적 계량 기법은 부동산 분야를 포함한 사회 과학 분야에 있어 많이 활용되는 반면, 그에 반해 기계학습은 Black Box 모형의 특징으로 결과에 대한 원인을 설명하는 능력이 부족한 단점이 존재한다. 하지만, 최근 기계학습 분야에 있어 XAI (eXplainable Artificial Intelligence)에 대한 연구를 통해 기계학습 모형의 설명가능성(Explainability)에 대한 관심이 높아지는 추세이다. 이러한 기계학습의 설명 가능성은 기존의 계량 기법 위주로 활용되거나 기계학습의 예측 능력 위주로 적용되던 사회 과학 분야(부동산, 금융 등)에 있어 다양한 연구의 가능성을 열어 주는 계기를 마련하였다. 이에 따라 본 연구에서는 기계학습의 설명가능성을 주택담보대출 채무불이행에 적용하여 실증 연구를 수행하였다. 최근 국내 주택 가격은 수도권 등 공급 감소과 저금리 기조에 따라 상승 추세에 있으나, 경기 침체 위험 및 그에 따른 주택 가격 하락에 의해 주택담보대출의 부실 위험과 규모는 확대될 수 있다. 또한, 금융 기관은 위험 자산에 대한 위험 가중치를 정확하게 산정하는 자구적 노력을 통해 BIS(Bank for International Settlements) 비율을 보다 정확하게 측정하고 이를 감독 기관으로부터 인정받는 것 또한 중요해졌다. 하지만, 기존의 주택담보대출과 관련된 연구에서 연체나 채무불이행, 또는 조기 상환 위험에 대해 전통적 계량 모형을 기반으로 관련 요인을 설명하는 데 치중해 왔다. 따라서, 기계학습 기반으로 보다 섬세한 채무불이행 위험 측정을 통해 내부 모형의 유효성을 증진시시키고, 채무불이행 위험에 대한 효율적인 규제 순응 기제로 활용될 수 있을 것으로 판단한다. 본 연구에서는 미 Freddie Mac社의 주택담보대출 자료를 이용하여 기계학습 기반(인공 신경망과 랜덤 포레스트)의 주택담보대출 채무불이행 모형을 도출하고, 모형에 나타난 채무불이행 위험의 요인을 PDP(Partail Dependence Plot), 한계 효과, SHAP(Shapley Additive Explanations) 등으로 설명하고자 하였다. 더불어, 전통적 계량 기법과 기계학습 모형의 예측력을 비교함으로써, 설명력뿐만 아니라, 예측력에 있어서도 기계학습 모형의 우수성을 설명하였다. 먼저, 계량 기법 중 하나인 로지스틱 회귀와 기계학습 기법 모형을 비교한 결과, 로지스틱 회귀 모형과 인공 신경망 모형을 통해 로지스틱 회귀 모형의 회귀 계수, PDP 및 한계 효과 등에서 두 모형이 대체적으로 비슷한 양상을 보인 반면, 랜덤 포레스트는 다소 상이한 결과를 보였다. 연체 상황 하에서는 인공 신경망과 랜덤 포레스트 두 모형 모두에서 연체 개월(-), 총 연체액(-), 신용 평점(+)가 상식과 대치되는 결과로, 본 연구 결과에서 흥미로운 부분 중 하나이다. 또한 기존의 설명 변수에 파생 변수를 추가한 데이터셋에 PDP, 한계 효과, SHAP 분석를 통해 신규 가망 변수를 파악하고자 하였다. 로지스틱 회귀와 인공 신경망에는 주택 가격 변화율 관련 변수가, 랜덤 포레스트에서는 자본 비율 관련 변수가 활용 가능성이 있는 변수로 고려되었다. 예측력 비교를 위해, 계량 기법으로는 로지스틱 회귀 모형을 사용하였고, 기계학습 기법으로는 인공 신경망과 랜덤 포레스트 모형을 사용하였다. 그 결과, 정확도(Accuracy), F1 Score, AUC (Area Under the Curve) 등 예측력을 검증하는 모든 지표에서 기계학습 모형이 우수한 것을 확인할 수 있었다. 본 연구를 통해 “Black Box 형태”의 특성을 갖는 기계학습 기반의 모형을 주택담보대출 채무불이행에 적용함으로써, 기계학습 모형의 예측력 뿐만 아니라, 설명가능성의 실무적 활용 가능성을 검증하였다. 더 나아가, 국내 금융기관이 기계학습 기반으로 주택담보대출 위험관리 시 길잡이 역할을 할 것으로 기대하며, 정책 입안 측면에서는 국내 부동산 금융시장의 가계부채에 대한 거시건전성 규제를 담당하는 정부에게도 유익한 정책적 실무지침을 제공하는 데 활용할 수 있을 것이다.
기계학습 알고리즘을 이용한 현황지목 분류에 관한 연구 : UAV 영상 활용을 중심으로
본 연구의 주된 목적은 UAV 영상과 기계학습 알고리즘을 이용하여 보다 효율적이고 신뢰성 있는 현황지목 식별방법을 제시하는 것이었다. 이러한 연구목적을 달성하기 위해 본 연구는 다음의 연구 방법을 이용하였다. 첫째, UAV 촬영 대상지를 선정하여 그 지역에 대한 영상 취득을 하였다. 둘째, 위에서 취득한 UAV 영상에 대하여 객체 기반 영상 분류를 실시하였다. 셋째, 분류된 객체와 영상을 중첩하여 현황지목 분류를 위한 Training 데이터를 추출하였고, 6개의 기계학습 알고리즘을 이용하여 Training 데이터를 분류하고 성능을 비교하였다. 넷째, 6개의 기계학습 알고리즘으로부터 얻은 래스터 자료와 법정 지목의 일치율 분석을 위해 연속지적도의 지목을 기반으로 한 래스터화 작업이 시도되었다. 다섯째, 마지막으로 Grid to Grid 방식을 이용하여 일치율 분석을 하였다. 영상처리는 Pix4D를 이용하였고, 자료 처리 및 분석은 QGIS, SAGA GIS 프로그램을 이용하였다. 자료 분석방법은 1차 분류인 현황지목 분류에 있어 6개의 기계학습 알고리즘별 성능 평가를 위해 분류의 정확도를 나타내는 Kappa 지수를 비교하였다. 둘째, 법정 지목과 현황지목의 일치도 분석을 위해 Kappa 및 Overall Accuracy를 비교하여 지목별, 기계학습 알고리즘별 정확성 검증을 실시하였다. 분석 결과 K-Nearest Neighbor의 알고리즘이 현황지목 분류에 있어서 가장 높은 신뢰도가 있는 것으로 나타났다. 또한 Decision Tree, Support Vector Machine, Random Forest, Naive Bayes 순으로 신뢰도가 높았으나 Artificial Neural Network은 낮은 신뢰도를 보여주었다. 둘째, 법정 지목과 현황지목 일치율 분석에 있어서는 Support Vector Machine이 가장 높은 신뢰도를 보여주었다. 또한 K-Nearest Neighbor, Decision Tree, Random Forest, Naive Bayes 등도 법정 지목과 현황지목 일치율 분석에 있어서 비교적 높은 신뢰도를 보여주었으나 현황지목 분류에서와 마찬가지로 Artificial Neural Network은 낮은 신뢰도를 보여주었다. 본 연구의 한계와 향후 연구 과제는 다음과 같다. 첫째, 본 연구는 기계학습 알고리즘을 이용하여 농촌지역의 현황지목만을 분석하였을 뿐이다. 따라서 동일한 알고리즘을 이용하여 도시지역의 현황지목을 분류한다면 다른 결과가 도출될 수 있을 것이다. 둘째, 만약에 UAV 영상분석에 있어서 다른 기계학습 알고리즘을 시용했다면 전혀 다른 결과를 도출했을 수도 있다. 본 연구에서는 여러 가지 기계학습 알고리즘 중에서 단지 6개의 알고리즘만을 사용하였다. 따라서 Sharkrf나 Sharkkm 와 같은 다른 기계학습 알고리즘을 이용한다면 본 연구의 결과와는 다른 결과를 도출할 수도 있을 것이다. 셋째, 과수원의 경우 임야와 구분이 어렵다는 점이다. 임야와 과수원의 경우 유사한 특성을 지니고 있기 때문에 OBIA를 이용하여 이들 간의 차이를 구분하기는 어렵다. 따라서 이러한 한계를 극복하기 위해서는 현황지목을 분류하는 데 있어서 OBIA뿐만 아니라 또한 Pixel 기반의 분류 기법이 동시에 이용되어야만 할 것이다. The primary purpose of this study was to propose a more efficient and reliable method for identifying land use categories using UAV images and machine learning algorithms. To achieve this purpose, this author used the following research methods. First, the UAV filming site was selected to acquire the image of the area. Second, object-based image classification was performed on the UAV images acquired above. Third, we extracted the training data for classifying the current category by superimposing the classified objects and pictures, and organizing the training data using six machine learning algorithms and comparing the performances. Fourth, a rasterization task based on the classification of continuous cadastral maps was attempted to analyze the agreement rate between the raster data obtained from the six machine learning algorithms and the legal land use categories. Fifth, we examined the concordance rate using the Grid to Grid method. Pix4d was used for image processing, and QGIS and SAGA GIS programs were utilized for data processing and analysis. The Kappa index, which represents the accuracy of classification, was compared to evaluate the performance of six machine learning algorithms. Second, Kappa and Overall Accuracy were compared to analyze the accuracy of each category and machine learning algorithm. As a result, the K-Nearest Neighbor algorithm has the highest reliability in the classification of the current category. Besides, decision trees, SVM, Random Forest, and Naive Bayes showed high reliability, but Artificial Neural Network showed low reliability. Second, SVM showed the highest reliability in analyzing the statutory and current status agreement rates. Also, K-Nearest Neighbor, Decision Tree, Random Forest, Naive Bayes showed relatively high reliability in the analysis of legal and existing category agreements, but artificial neural network showed low reliability. The limitations of this study and future research are as follows. First, this study only analyzes the current land use categories of rural areas using machine learning algorithms. Therefore, if we classify the existing land use categories of urban areas using the same algorithms, different results may be obtained. Second, if different machine learning algorithms were used in UAV image analysis, totally different results could be obtained. In this study, only six algorithms were used among the various machine learning algorithms. Thus, using other machine learning algorithms such as Sharkrf or Sharkkm may yield different results. Third, in the case of orchards, it is difficult to distinguish them from forestry. Because forests and orchards have similar physical characteristics, it is difficult to differentiate between them using OBIA. Therefore, to overcome these limitations, not only OBIA but also pixel-based classification techniques should be used simultaneously in classifying the current land use categories.
기계학습 알고리즘 기반 빙축열 시스템 시뮬레이션 모델 개발
라선중 성균관대학교 일반대학원 2018 국내석사
근래에 친환경 저에너지 건축의 관심이 증가하면서 건물의 생애주기 중 유지관리 측면에서 건물 에너지 절감의 필요성이 대두되고 있다. 사무용 건물의 에너지 사용량 중 32%는 HVAC 시스템에 의해 소비되며(Sane, Haugstetter, & Bortoff, 2006), 이에 따라 건물 에너지와 관련된 모든 이해관계자는 건물 에너지 해석을 통해 효율적으로 에너지 사용량을 관리해야 한다. 기존 건물의 에너지 해석을 위해 사용되는 방법은 전통적으로 제 1법칙 기반 모델이 활용되었으며, ISO 13790과 같은 규범적 모델(normative model) 또는 동적 시뮬레이션 도구(EnergyPlus, TRNSYS, IES-VE, IDA-ICE) 기반 모델을 활용한 방법으로 구분된다. ISO 13790과 같은 규범적 모델은 사칙연산과 간단한 대수식으로 건물 에너지 사용량을 평가하므로, 사용자에 주관적 개입이 필요치 않으며, 동일한 계산 결과를 산출 할 수 있고, 수식에 따른 입·출력 변수 사이의 관계를 파악할 수 있다(van Dijk, Spiekman, & de Wilde, 2005). 동적 시뮬레이션 도구를 사용하여 구축된 모델은 시간에 따라 변화하는 건물의 동적 거동을 표현할 수 있는 장점이 있다(안기언, 김영진, & 박철수, 2012). 그러나 기존 건물을 대상으로 동적 시뮬레이션 도구를 활용해 모사할 경우, 건물 전반에 걸친 전문 지식과 상당한 모델링 시간이 요구되는 단점이 있다. 또한 확률적으로 변화하는 변수(재실, 조명, 기기 스케쥴, 침기)에 대한 고려와 함께 수많은 입력 정보들을 반영해야 한다. 반면 기계학습 모델은 모델 작성에 고도의 전문 지식이 요구되지 않고 다수의 변수에 대한 추정이 불필요하다. 또한 실제 측정된 데이터를 활용하여 소수의 입력 변수만으로 모델 구축이 가능하며, 모델이 알고리즘에 의해 자율 구성되므로, 동적 시뮬레이션 모델 작성보다 더 실용적인 방법이 될 수 있다(서원준 & 박철수, 2016; Edwards, New, & Parker, 2012; Kim, Ahn, & Park, 2016; Zhang et al., 2015). 본 연구의 대상 건물은 서울특별시에 위치한 오피스 및 근린생활 건물이며, 연면적은 32,600, 지하 7층/지상 30층 규모이다. 대상 건물은 하절기(7~8월)동안 빙축열 시스템을 통해, 냉방을 실시하며 본 연구에서는 전력 사용량과 연관된 냉각탑, 냉동기, 브라인 펌프, 냉각수 펌프, 공기조화기에 대해 기계학습 모델을 작성하였다. 기계학습 모델을 구축하기 위한 과정은 기계학습 알고리즘의 파라미터 선정, 훈련·검증 기간 선정, 입·출력 변수 선정, 커스터마이징, 모델 구축으로 이루어지며, 구축된 기계학습 모델의 예측 성능, 연산 시간 등을 분석한다. 본 연구에서 커스터마이징을 통해 개발된 기계학습 모델은 기존 구축되는 기계학습 모델 대비, 정확성과 연산시간 측면에서 우수한 성능을 지닌 모델임을 알 수 있으며, 동시에 기존 건물을 대상으로 다양한 기계학습 모델을 구축할 때 발생하는 제약사항 등을 개선시킬 수 있었다. 본 연구의 결과는 다양한 설비 기기에 따라 적절한 기계학습 알고리즘 및 입·출력 변수 선정 등의 도움이 될 것으로 판단되며, 기계학습 모델이 기존 최적제어 및 진단에 활용되던 동적 시뮬레이션 모델의 실용적인 대안이 될 수 있음을 보여준다. 본 연구에서 최종적으로 개발된 각 기기별 기계학습 모델은 MPC(Model Predictive Control), 실시간 온라인 업데이트 모델, 각 기기별 기계학습 모델의 통합(federated modeling)에 관한 연구에 기반이 될 수 있을 것으로 판단된다.
회귀분석 및 기계학습을 활용한 연약 점성토 지반의 압축지수 평가
대규모 사회기반시설이 해안 및 연약지반에 집중적으로 건설됨에 따라, 연약지반의 압밀침하로 인한 구조물 손상의 관심이 높아지고 있다. 국내의 대표적인 연약지반은 수계를 중심으로 발달되어 있으며, 생성요인, 퇴적환경, 조성광물성분 등의 지역별 조건에 따라 지반공학적 특성이 상이하다. 이에 따라 연약지반의 압밀침하량을 정확히 산정하기 위해서는 지역특성을 고려한 정확한 지반정수를 활용하여야 한다. 하지만 공사규모에 따른 광범위한 지역에 대해 정밀한 지반조사는 시간 및 비용 등의 문제로 인해 현실적으로 불가능하며, 소수의 지반조사 결과를 바탕으로 전체 대상 지역의 압밀침하량을 예측하고 있는 실정이다. 본 연구에서는 국내 연약지반을 대상으로 기존의 지반조사 결과로부터 물리‧역학적 지반정수를 수집하여 지역별 분포특성을 파악하고 통계 및 기계학습 기법을 통해 수계별 압축지수 예측 모델을 제시하고자 하였다. 이를 위해 구축된 DB로부터 영산강, 섬진강, 낙동강 수계를 선정하여 지역별 데이터셋을 구축하였다. 또한, 구축된 데이터셋의 특성을 분석하고, 상관분석을 통해 압축지수와 유의미한 관계를 나타내는 영향인자(자연함수비, 액성한계, 단위중량, 소성지수, 초기간극비)를 선정하였으며, 단일 및 다중 인자를 활용한 압축지수 예측을 위한 회귀식을 제시하였다. 더불어, 기계학습 기법을 활용한 압축지수 예측 모델의 개발을 위해 상관분석 결과로부터 선정된 영향인자를 입력데이터로 활용하고 압축지수를 출력 데이터로 활용하여 RandomForest, XGBoost, LightGBM, Linear Regression에 적용하여 모델을 개발하였다. 모델의 평가지표로는 RMSE, MAE, R2로 선정하고, 수계별로 가장 우수한 결과가 도출된 최적의 모델을 제시하였다. 각 모델에서 도출된 영향인자의 중요도는 자연함수비와 초기간극비로 나타났다. 또한, 다중회귀식과 기계학습 모델의 비교를 통해, 기계학습 기법이 적용된 압축지수 예측 모델의 신뢰도를 확인하였다. 본 연구는 영산강, 섬진강, 낙동강 수계의 연약지반을 대상으로 통계 및 기계학습 기법이 적용된 압축지수 예측 모델을 개발하였다. 추후 다수의 데이터가 확보된다면, 본 연구에서 제안된 기계학습 모델 개발 프로토콜을 활용하여 높은 신뢰도의 압축지수 예측 모델의 개발이 가능할 것으로 기대된다. 또한, 간단한 물성실험을 통해 개략적인 압축지수의 값을 추측하고, 실험값에 대한 신뢰도를 뒷받침하는 기초자료로의 활용이 가능할 것으로 기대된다. As large-scale social infrastructure facilities are increasingly being built on coastal areas and soft grounds, the damage to structures caused by consolidation settlement of soft clayey ground is becoming a major concern. The soft grounds in Korea are developed around water systems, and the geotechnical characteristics of soft grounds vary according to local conditions such as formation factors, sedimentary environments, and composition of clay mineral components. Thus, to accurately calculate the amount of consolidation settlement of soft grounds, it is necessary to use accurate soil parameters that take into account the local characteristics. However, conducting precise geotechnical investigations for large areas based on the scale of construction projects is often impractical due to constraints such as time and cost. As a result, the current practice often involves predicting the amount of consolidation settlement of the entire target area based on a limited number of geotechnical investigation results. This study collected physical and mechanical soil parameters from existing geotechnical investigatioln results for soft ground in Korea. The study identified distribution characteristics by region and proposed a compression index prediction model for each water system through statistical and machine learning techniques. To do this, the Yeongsan River, Seomjin River, and Nakdong River water systems were selected from the database that was built, and regional data sets were constructed. Furthermore, the characteristics of the constructed dataset were analyzed. Additionally, through correlation analysis, influencing factors (such as natural water content, liquid limit, unit weight, plasticity index, and initial void ratio) that exhibit a significant relationship with the compression index were identified. Regression equations for predicting the compression index using single and multiple factors were also proposed. In addition, to develop a compression index prediction model using machine learning algorithms, we applied Random Forest, XGBoost, LightGBM, and Linear Regression. These algorithms were applied using the influencing factors selected from the correlation analysis results as input data and the compression index as output data. RMSE, MAE, and R 2 were selected as model evaluation index, and the optimal model with the best results for each water system was presented. The importance of the influencing factors derived from each model was found to be natural water content and initial void ratio. In addition, the reliability of the compression index prediction model utilizing machine learning techniques was verified through the comparison between multiple regression equations and machine learning models. This study developed a compression index prediction model using statistical and machine learning techniques for soft ground in the Yeongsan River, Seomjin River, and Nakdong River of Korea. In the future research, with the acquisition of a large volume of data, it is expected that the machine learning model development protocol proposed in this study can facilitate the development of a highly reliable compression index prediction model. Moreover, it is expected that approximate values of the compression index can be estimated through simple physical tests, and these experimental values can be used as foundational data to support their reliability.
유체-열 연성해석 기반 기계학습 건물에너지 모델 및 플랫폼 개발
This study developed and verified a machine learning building energy model based on fluid-heat interaction that operates with limited parameters provided in low-cost measurement infrastructure for the widespread application of building energy management means. In addition, a specific case was presented by developing a digital platform that can operate the building energy model presented for the contribution of the industrial aspect. To this end, a simplified hydrodynamic model that detects dynamic changes in indoor airflow in real time was developed and verified by considering the limitations of the existing building energy model consisting of heat transfer and machine learning. To this end, a simplified hydrodynamic model that detects dynamic changes in indoor airflow in real time was developed and verified by considering the limitations of the existing building energy model consisting of heat transfer and machine learning. and Based on this, a machine learning building energy gray box model based on fluid-heat transfer ductility analysis was developed to compare and analyze the improvement effect on the machine learning model. As the first step, the meaning of the building energy of indoor temperature, which is an indoor environment measurement result and an input parameter, was considered, and the appropriateness of the input parameter was reviewed using a statistical method. In this process, the indoor thermal environment was estimated based on the regression of outdoor temperature and indoor temperature, and a cooling and heating space classification model based on the coefficient of determination was developed and presented. The results of the classification model showed a strong linear relationship with Feature Importance for building energy (R=0.793), confirming that it works closely with building energy. In addition, based on the Feature Importance of characteristics reflected at this time, the significance probability of indoor temperature for building energy (p<0.05) was confirmed. Subsequently, a simplified fluid mechanics model capable of detecting dynamic changes in real-time indoor airflow was developed and verified. The limitations of existing heat transfer network models that cannot consider indoor temperature distribution were considered, and the mathematical induction process of fluid kinetic force using Reynolds Transport Theorem (RTT) was explained. In order to verify the fluid mechanics model, the indoor temperature distribution was predicted according to the installation location of a single sensor, compared and analyzed with the actual measurement results. The fluid mechanics model outputted the indoor temperature distribution based on the local measurement results, providing sophisticated prediction results (RMSE 0.132℃, R20.966) for the average temperature of the flow axis. This shows that it can compensate for the limitations of the heat transfer model, which cannot consider the indoor temperature distribution. In addition, the fluid's kinetic and shear forces were expressed in a network to explain the ductility analysis method with the heat transfer model. Finally, explained the modularization of statistics, fluid mechanics (RTT), heat transfer (RC), and machine learning (ML) models and the linkage of machine learning building energy models based on fluid-heat coupled analysis. The fluid mechanics model is derived from the Reynolds transport theorem presented above, and the heat transfer model minimizes the input parameters in a 1R1C way that can take into account the total heat resistance and capacity by reflecting the room temperature change every hour. Based on the 11 type of machine learning algorithm, a total of 33 models were developed and predicted performance was compared by dividing them into black box and gray box model groups according to the input structure. The RTT-1R1C-based model group presented in this study provides the best results with 7.4% to 43.8% predictive performance improvement according to the machine learning algorithm compared to the black box model group, showing that predictive performance for various machine learning algorithms can be improved in a limited parameter environment. Moreover, unlike improvements between black box models, RMSE and MAE simultaneously improved, improving the outlier and error sum together, which was also excellent in terms of improvement quality. RTT-1R1C-based deep learning (DNN) provided the most accurate prediction results among 33 models with 30.2% improvement (RMSE 0.148), showing 27.1% better prediction performance than the simplest multivariate linear regression (MLR, RMSE 0.203), but 14.2 times lower computational efficiency in terms of training time. Remarkably, RTT-1R1C-MRL based on fluid-heat interaction analysis provides 4.2% better predictive performance than deep learning (To-DNN, RMSE 0.212) operating outdoor temperature as input, showing that advanced algorithms are not always best. The research results show that the RTT-1R1C-ML model effectively improves the predictive performance of various machine learning algorithms in a limited parameter environment of a low-cost indoor environment measurement infrastructure. Therefore, a building energy management digital platform consisting of a fluid-heat interaction analysis-based machine learning building energy model, sensor module, and software that operates it was developed and presented as a concrete example for the widely application of economical building energy management. The results of this study are believed to be the starting point for the continuous development of policies and technologies for the spread of economic and extensive building energy management methods and energy efficiency of existing buildings. 온실가스 절감을 위한 건물에너지 효율화 정책 노력에도 불구하고 경제적 제약으로 건물에너지관리시스템의 보급은 여전히 부족하다. 효과적인 건물에너지 관리를 위한 예측은 더 정확할수록 보다 많은 입력 데이터를 요구하며, 이는 필연적으로 건물에너지관리시스템의 구축 비용 증가로 이어진다. 따라서, 건물에너지 관리 수단의 광범위한 보급과 확산을 위해서는 경제적 측정 인프라의 제한된 매개변수 환경에서 효과적으로 작동하는 건물에너지 모델이 필요하다. 본 연구는 건물에너지 관리 수단의 광범위한 보급을 목적으로 저비용 측정 인프라의 제한된 매개변수로 작동하는 유체-열 연성해석 기반 기계학습 건물에너지 모델을 개발하고 검증하였다. 또한, 산업적 측면의 기여를 위하여 제시하는 건물에너지 모델을 운용할 수 있는 실내환경 측정기반 건물에너지 관리지원 디지털 플랫폼을 개발하여 구체적인 사례를 제시하였다. 이를 위해, 기존의 열전달, 기계학습으로 구성되는 건물에너지 모델의 한계를 고찰하고, 실내 기류의 동적 변화를 실시간 감지하는 단순화된 유체역학 모델을 개발, 검증하였으며, 이를 바탕으로 유체-열 연성해석 기반 건물에너지 모델을 개발하고, 다양한 기계학습 알고리즘에 대한 건물에너지 예측 성능을 비교, 분석하였다. 먼저, 입력 매개변수인 실내온도의 건물에너지에 대한 의미를 고찰하고, 통계 방법을 활용하여 입력 매개변수의 적절성을 검토하였다. 이 과정에서 실외온도와 실내온도의 회귀를 바탕으로 실내 열 환경을 추정하고, 결정계수 기반의 냉난방 공간 분류 모델을 개발, 제시하였다. 분류 모델의 결과는 건물에너지에 대한 특성 중요도(Feature Importance)와 강한 선형 관계(R=0.793)를 보여 건물에너지와 밀접하게 작동함을 확인하였다. 또한, 입력 매개변수인 실내온도에 대하여 건물에너지에 대한 특성 중요도와 유의확률(p<0.05)를 검토하였다. 이어서 실시간 실내 기류의 동적 변화를 감지할 수 있는 단순화된 유체역학 모델을 개발, 검증하였다. 기존 열전달 모델이 실내온도 분포를 고려할 수 없는 한계를 고찰하고, 레이놀즈 수송 정리(RTT, Reynolds Transport Theorem)를 활용한 수학적 유도로 유체 운동력을 단순하게 표현하는 유체역학 모델을 제시하였다. 유체역학 모델의 검증을 위하여 단일 센서의 설치 위치에 따라 실내온도 분포를 예측하고 실측 결과와 비교, 분석하였다. 유체역학 모델은 국부적 측정 결과를 바탕으로 실내온도 분포를 출력하며, 유동 축의 평균온도에 대하여 정교한 예측 결과(RMSE 0.132 ℃, R2 0.966)를 제공하면서, 실내온도 분포를 고려할 수 없는 열전달 모델의 한계를 극복할 수 있음을 보여주었다. 또한, 유체의 운동력과 전단력으로 구성되는 네트워크 표현으로 열전달 모델과의 연성 해석 방법을 설명하였다. 마지막으로, 앞서 제시한 모델을 포함하여 통계, 유체역학(RTT), 열전달(RC), 기계학습(ML) 모델의 모듈화 및 연동으로 유체-열 연성해석 기반 기계학습 건물에너지 모델을 제시하였다. 유체역학 모델은 레이놀즈 수송 정리로 도출되었으며, 열전달 모델(RC)은 변화하는 실내온도를 반영하는 총합 열 저항과 커패시턴스로 구성되는 1R1C 형태로, 입력 매개변수를 최소화하였다. 그리고 이를 바탕으로 11 유형의 기계학습 알고리즘을 바탕으로 입력 구조에 따라 블랙박스 및 그레이박스 모델 그룹으로 구분하고, 총 33개 모델을 개발하여 예측 성능을 비교, 분석하였다. 그레이박스 모델인 유체-열 연성해석 기반 모델(RTT-1R1C)은 비교 대상인 블랙박스 모델들에 비해 알고리즘에 따라 RMSE 7.4%~43.8%의 예측 성능 개선으로 최상의 결과를 제공하면서, 제한된 매개변수 환경에서 다양한 기계학습 알고리즘의 예측 성능을 개선할 수 있음을 보여주었다. 또한, 블랙박스 모델 간의 개선과 다르게 이상치(Outlier)와 오차합의 동시 개선으로 예측 품질 측면에서도 우수했다. 유체-열 연성해석 기반 딥러닝(DNN) 모델은 블랙박스 모델 그룹의 딥러닝에 비해 30.2% 개선된(RMSE 0.148) 예측 성능으로, 33개 모델 중 가장 정확한 결과를 제공하였다. 가장 단순한 다변량 선형회귀(MLR) 알고리즘을 반영한 유체-열 연성해석 기반 모델은 딥러닝에 비해 27.1% 낮은 예측 결과(RMSE 0.203)를 보였으나, 가장 빠르게 훈련되어 계산 효율 측면에서는 가장 우수했다. 그러나 괄목할만한 것은 유체-열 연성해석 기반 다변량 선형회귀 모델(RMSE 0.203)은 실외온도를 입력으로 작동하는 딥러닝 모델(RMSE 0.212)에 비해 4.2% 우수한 예측 결과를 보인 점이다. 즉, 유체-열 연성해석 기반 기계학습 건물에너지 모델은 가장 단순한 알고리즘으로도 블랙박스 딥러닝의 예측 성능 한계를 극복할 수 있음을 보여주었다. 결과적으로, 제시한 유체-열 연성해석 기반 기계학습 건물에너지 모델은 저비용 실내환경 측정 인프라가 제공하는 제한된 매개변수 환경에서 다양한 기계학습 알고리즘의 예측 성능을 효과적으로 개선하였다. 따라서, 제시하는 건물에너지 모델과 센서 모듈 그리고 이를 운용하는 소프트웨어로 구성되는 실내환경 측정기반 건물에너지 관리지원 디지털 플랫폼을 개발하여 산업적 측면의 활용을 위한 구체적 사례를 제시하였다. 본 연구 결과는 경제적이고 광범위한 건물에너지 관리 수단의 보급 확산과 기존 건축물의 에너지 효율화를 위한 정책, 기술 개발이 지속적으로 이루어지는 단초가 될 수 있을 것으로 판단된다.
In 2022, the South Korean domestic art market had its yearly turnover surpass 1 trillion won for the first time in its history. It brings the art market substantial public attention. The perception of artwork has changed as a result of this movement, moving from aesthetic value to investment commodities. It causes a variety of trading strategies for investing in art to evolve, especially in fractional ownership of artwork. This study was carried out in accordance with these advancements in order to give prospective art purchasers more precise estimates of art prices. This study differed from others in that it used machine learning algorithms, which are commonly applied in many domains, in addition to the hedonic model to improve the precision of art price predictions. In order to conduct the analysis, 11,810 art price data for paintings sold through domestic auction houses between 2015 and 2022 were gathered and preprocessed. To ensure transparency, all data included in this study came from databases that were made available to the public. In addition, factors dependent on subjective opinion were eliminated. Only the factors that are objective were used in the analysis. Most of the variables, with the exception of a few material and transaction year variables, were found to be statistically significant based on the findings of the hedonic model estimation using the least squares method. Reliability in prediction was indicated by the model’s coefficient of determination (R-squared value), which was 0.643. Furthermore, nonlinear models were used in this work as opposed to the linear equations of the hedonic model. The accuracy of art price prediction was examined using machine learning models such as artificial neural network models and decision tree models. To prove the robustness of machine learning models, a 5-fold cross-validation was carried out to prevent potential bias resulting from validation on a small dataset. To provide a comprehensive outcome, the average of five cross-validation results was obtained. Additionally, by classifying the materials into four categories (original wet, western wet, dry and pen drawing, and ohters), artificial neural network and decision tree model were examined in order to validate the robustness of each market in addition to the data aspect. In this study, the architecture of decision trees and artificial neural network models was modified with five different hyperparameters set. The most fundamental artificial neural network model using ten perceptrons in a hidden layer had the highest coefficient of determination with a value of 0.7226. The bagging tree model, which arranges small trees in parallel, performed the best when it came to the decision tree model, with a value of 0.7655. These findings demonstrated that machine learning models, specifically the decision tree model, performed better in art price prediction than the hedonic model. It was underlined that better performance is not always a given just because machine learning models are more complex. Moreover, models exhibiting low performance and models with high performance were nearly always appeared consistently, indicating that the models were reliable. Furthermore, the prediction effectiveness of machine learning was found to be significantly impacted by uniform distribution and sufficient data acquisition. Although machine learning models outperformed other prediction models, they were limited in their ability to offer a clear understanding of the specific effects of each variable on art prices. One interpretable machine learning model, the symbolic regression model, was used to get around this restriction. Among 17 equations, the optimal equation was proposed in consideration of the degree of change in error when the error, complexity, and the derivative of error with respect to the complexity. Similar to the previous analysis of the artificial neural network and decision tree model, market-specific analysis was performed using the symbolic regression model. The optimal equations for the Eastern wet, Western wet, dry and pen drawing, and others were , , , , respectively. The number of variables used to construct the formula was the same from 6 to 7, and the variables commonly used by market were online auctions, other auction companies, and 2022 year variables, confirming that the variable had an important effect when explaining the art price model. This study is significant because it derives a model with better prediction performance than the hedonic models that have been the main focus of previous studies. As such, it validates machine learning models’ applicability in art price prediction. Future applications of more sophisticated machine learning models should improve the accuracy of art price prediction. In addition, using symbolic regression, it was visually confirmed that the error decreases as the complexity of the equation increases, but the interpretation of the equation becomes difficult, even if the prediction performance decreases. In the future, it is expected that if an advanced explainable machine learning model that can explain the degree of influence variables while performing well will be applied, it will be easy to explain and can be applied immediately in practice. This study also notes its limitations, including the challenge of gathering image data for input. On the other hand, it is thought that adding visual data from artworks as variables would produce even more precise predictions. The findings of this study can stimulate further studies that utilize advanced machine learning models and improved datasets. It is anticipated that there will be an expansion of several applications, such as investment return analyses and the computation of art price indices. Additionally, with this study as a foundation, the application of machine learning models to art price prediction will improve the transparency of the art market and address information asymmetry, which will help the market expand. The art market developing in a positive direction is the ultimate goal of this study. 2022년에는 국내 미술시장의 연간 유통액이 역사상 처음으로 1조 원을 돌파하며, 미술품 시장에 대중들의 큰 관심을 불러일으키고 있다. 이러한 경향은 미술품을 단순히 미적 가치보다는 투자 수단으로 인식하는 보편적인 관념으로 이어지며, 미술품 투자를 위한 다양한 거래방식, 특히 미술품 조각 투자 등의 등장을 촉진하고 있다. 이러한 추세에 맞춰, 본 연구는 잠재적인 미술품 구매자들에게 더 정확한 미술품 가격 예측 정보를 제공하고자 진행되었다. 미술품 가격 예측의 정확성 향상을 목표로, 기존의 선행연구와는 다르게 헤도닉 모형뿐만 아니라 최근 다양한 분야에서 폭넓게 활용되고 있는 기계학습 모형을 활용하였다. 분석을 위해 2015년부터 2022년까지 국내 경매회사를 통해 거래된 회화 작품의 미술품 가격 데이터 11,810점을 수집·가공하였다. 연구에 활용된 모든 데이터는 대중들도 쉽게 접근할 수 있는 공개된 데이터를 활용하였다. 또한, 모형 설계의 신뢰를 강화하기 위해, 기존 선행연구들에서 변수들을 참고하였으며, 연구자의 주관적 판단에 의존하는 변수는 제외하고 객관적인 정보에 근거한 변수들만 채택하였다. 최소 자승법을 사용한 헤도닉 모형의 추정 결과는 재료변수와 거래 연도 변수의 몇몇을 제외하고 대다수의 변수가 통계적으로 유의미한 것으로 나타났다. 모형의 결정계수 값은 0.643으로 나타나, 신뢰할 수 있는 예측 결과를 얻었음을 확인할 수 있었다. 더 나아가, 헤도닉 모형의 선형 방정식과는 달리 비선형의 모형으로서, 기계학습의 인공신경망 모형과 결정트리 모형을 활용하여 미술품 가격 예측 성능을 분석하였다. 일부 데이터에만 검증을 수행하는 것은 편향된 결과를 초래할 수 있으므로, 이를 방지하고 기계학습 모형의 강건성을 향상시키기 위해 5겹 교차 검증을 수행하였으며, 5번의 교차 검증 결과를 평균화하여 종합한 결과를 도출하였다. 또한 데이터의 측면외에 시장별 강건성을 확인하기 위해, 재료를 기준으로 동양 습식, 서양 습식, 건식 및 펜 드로잉, 기타 시장 총 4개로 나누어 인공신경망 모형과 결정트리 모형을 분석하였다. 본 연구에서는 인공신경망 모형과 결정트리 모형의 구조를 조정하여 각각 5개의 다양한 모형을 사용하여 실증분석을 실시하였다. 결과적으로, 인공신경망모형에서는 은닉층이 1개이며 퍼셉트론이 10개인 가장 단순한 모형이 최상의 결정계수를 가지며, 그 값은 0.7226으로 나타났다. 결정트리 모형의 경우, 작은 트리를 병렬로 배치하는 배깅트리 모형이 가장 뛰어난 결과를 나타냈으며, 는 0.7655였다. 이와 같은 결과를 토대로, 헤도닉 모형 대비 기계학습 모형이 미술품 가격 예측성능에서 더 우수함을 확인할 수 있었으며, 특히 결정트리 모형의 성능이 가장 우수한 것으로 나타났다. 또한, 기계학습 모형의 복잡도가 항상 높은 성능을 나타내는 것은 아니라는 사실을 강조하였다. 더 나아가, 시장별 분석에서도 성능이 좋지 못한 모형과 성능이 좋은 모형이 거의 일관되게 나와 강건한 모형임을 확인할 수 있었다. 또한, 기계학습의 예측 성능에는 충분한 데이터 확보와 균일한 분포가 많은 영향을 미치는 것을 확인할 수 있었다. 기계학습 모형의 경우 예측 성능이 우수하지만, 각 변수가 미술품 가격에 미치는 영향을 명확하게 확인할 수 없다는 한계가 있다. 이에 따라, 설명 가능한 기계학습 모형 중 하나인 기호회귀 모형를 활용하여, 총 17개의 식을 도출하였다. 그중에서 모형의 오차와 복잡도, 복잡도를 증가하였을 때 오차의 변화 정도를 고려하여 최적의 식 를 제안하였다. 앞서 인공신경망 모형 및 결정 트리 모형 분석과 마찬가지로 기호회귀 모형을 활용하여 시장별 분석을 실시하였다. 동양 습식 시장, 서양 습식 시장, 건식 및 펜 드로잉 시장, 기타 시장에 대한 최적의 식은 각각 , , , 이다. 전식을 구성하는데 활용되는 변수의 개수는 6개에서 7개로 동일하였으며, 시장별 공통적으로 활용된 변수로는 온라인 경매, 기타 경매회사, 2022년 변수인 것으로 확인되어, 해당 변수가 미술품 가격을 설명할 때 중요한 영향을 미치는 것을 확인할 수 있었다. 본 연구는 기존 선행연구에서 주로 사용되던 헤도닉 모형에 비해 더 우수한 예측 성능을 보이는 모델을 도출한 데에 의의가 있다. 이를 통해 미술품 가격 데이터 예측에 기계학습 모형이 적합하다는 것을 확인할 수 있었다. 향후 더 발전된 기계학습 모형을 적용함으로써 미술품 가격 예측의 정확성을 높일 수 있을 것으로 기대된다. 뿐만 아니라, 설명가능한 인공지능의 한 종류인 기호회귀를 활용하여, 예측 성능은 다소 떨어지더라도 중요한 변수가 무엇인지, 식의 복잡도가 증가함에 따라 오차는 줄어들지만, 식의 해석이 어려워지는 것을 가시적으로 확인하였다. 추후에 성능이 좋으면서 변수의 영향 정도를 파악할 수 있는 발전된 설명가능한 인공지능 모형을 적용하면 해석에도 용이하여 실무에 즉각적으로 응용될 수 있는 식이 나올 것으로 기대된다. 또한, 본 연구에서는 이미지 데이터 수집 과정에 어려움을 겪어 이를 입력으로 적용하지 못했다는 한계가 있지만, 미술품의 시각 정보를 변수로 활용할 수 있다면 더 정확한 예측이 가능할 것으로 판단된다. 본 연구를 통해, 통계적으로 높은 신뢰도를 갖는 미술품 가격 예측 모형에 대한 더 많은 연구가 이루어질 것을 기대한다. 더 나아가, 미술품 가격지수의 산출 및 투자 수익률 분석과 같은 다양한 응용 분야로 확장될 수 있길 기대한다. 뿐만 아니라, 이를 바탕으로, 미술품 시장의 정보 비대칭을 보완하고 시장 투명성을 향상시켜 미술시장의 성장에 기여하길 바라며, 미술시장이 더 나은 방향으로 발전할 수 있기를 희망한다.
음향-조음 역변환(acoustic-to-articulatory inversion)이란 음향 정보(acoustic features)를 통해 조음점(articulators)을 역추적하는 방법으로 이론적, 실용적 측면에서 다양하게 활용될 수 있는 중요한 연구임에도 불구하고, 문제 자체의 비고유성(non-uniqueness)과 비선형성(non-linearity)으로 인해 해결하기 어렵다는 단점을 갖는다. 본 연구에서는 이러한 음향-조음 역변환 문제를 해결하기 위해 선행연구에서 제시한 다양한 기계학습(machine learning) 방식을 참고하여 1개의 네트워크로 구성된 단일 기계학습 모델과, 2개의 네트워크로 구성된 복합 기계학습 모델을 제시하여 음향-조음 역변환 문제 해결에 가장 적합한 기계학습 모델을 알아보고자 한다. 본 실험에서는 단일 기계학습 모델을 DNN(Deep Neural Network), CNN(Convolutional Neural Network), RNN(Recurrent Neural Network), BRNN(Bidirectional RNN) 모델로 구성하여 음향 데이터를 가장 정확하게 학습하는 모델을 파악하고, 복합 기계학습 모델은 DRNN(DNN-RNN), DBRNN(DNN-BRNN), CBRNN(CNN-BRNN) 모델로 구성함으로써 앞 단 모델에서 음향 데이터의 특징을 가장 잘 추출하는 모델을 알아보고자 한다. 모델의 학습을 위해 mngu0 코퍼스의 음향 정보를 입력값(input)으로 삼고, EMA(electromagnetic articulo-graphy) 조음 정보를 목표값(target)으로 이용하여 전체 모델을 학습하였다. 훈련이 완료된 모델에 mngu0 코퍼스에서 제공하는 테스트 음향 정보를 넣어 조음점을 예측하였으며 그에 대한 RMSE(Root Mean Square Error)를 계산하였다. 단일 기계학습에서는 음향 데이터의 시간정보를 양방향으로 동시에 훈련하는 BRNN 모델(RMSE: 1.022mm)이 단방향 만으로 훈련하는 RNN 모델(RMSE: 1.126mm)보다 우수한 성능을 거두었으며, RNN 모델은 DNN 모델(RMSE: 1.140mm)보다 DNN 모델은 CNN 모델(RMSE: 1.194mm)보다 더 나은 성능을 보여주었다. 이는 음향-조음 역변환 문제를 다룰 때, 시간 정보의 학습이 매우 중요하다는 것을 의미한다. 복합 기계학습 모델에서는 CBRNN 모델(RMSE: 0.975mm)이 DBRNN 모델(RMSE: 1.011mm)보다 더 나은 성능을 보였는데 이는 CNN 모델의 음향 데이터 특징 추출 능력이 DNN 모델보다 탁월하다는 것을 의미하며, CNN 모델 자체의 부족한 시계열 음향 특징(temporal acoustic features)의 학습능력을 시계열 정보 학습에 뛰어난 BRNN 모델을 통해 극복하였다고 볼 수 있다.
전세특성을 고려한 공동주택 가격추정에 관한 연구 : -회귀분석과 기계학습 기법 비교를 중심으로-
이다솜 서울시립대학교 도시과학대학원 2022 국내석사
2020년은 평년과 비교하여 가파른 부동산 가격상승이 발생한 시기였다. 전세보증금을 이용한 매수 형태인 갭투자가 크게 증가하였고 갭투자는 매매유형의 하나로 자리 잡았다. 공동주택 가격 또는 가격형성요인에 관한 선행연구는 매매시장과 전세시장을 분리해서 고려하거나 매매시장과 전세시장의 지수분석을 통한 선후관계 입증이 주를 이루었다. 본 연구는 공동주택 가격추정에 있어 가격형성요인으로 전세특성을 고려하여 과연 어떤 요인들이 공동주택 매매가격에 영향을 미치는지를 회귀분석과 기계학습 기법을 통해 살펴보고, 그에 대한 분석과 함께 도출된 결과를 바탕으로 정책적 대안을 제시하는 데 목적을 두고 있다. 본 연구는 대전광역시 유성구 노은지구의 아파트 35개 단지를 대상으로 실거래가격 및 전세가격 분석을 통해 진행하였다. 연구의 방법으로는 이론적 배경으로 매매시장과 전세시장의 개념과 특성, 매매시장과 전세시장을 바라보는 두 가지 이론에 대해 알아보고 방법론으로서 기계학습 기법에 대한 개념 및 이론을 살펴보았다. 관련 이론을 바탕으로 기존 선행연구들에서 밝혀진 한계점과 문제점들을 고려하여 변수를 설정하였다. 이후 분석의 방법 및 연구모형을 정하고 가설을 설정하였으며, 실증분석을 위해 먼저 기초통계량 및 각 변수 분석을 실시하였고 통계기반 회귀분석 기법인 헤도닉 모형 분석과 기계학습 모형 분석을 병행하였다. 회귀분석의 헤도닉 모형 분석을 통해 독립변수 중 전세특성 변수의 방향성에 대해 검증하였고 기계학습 모형 분석을 통해 각 모형별 예측 정확도를 도출하였다. 이후 순열특성중요도(Permutation feature importance) 분석을 통해 공동주택 가격형성요인의 중요도 분석을 실시하였다. 분석 결과, 공동주택 가격형성요인에서 전세특성 변수인 전세거래량과 전세금액이 모두 매매가격에 정(+)의 영향력을 미치는 것으로 나타났다. 각 모형별 예측 정확도는 기계학습 기법인 랜덤포레스트 모형이 가장 높게 나왔고 그다음으로 딥러닝 모형, 회귀분석 모형 순으로 예측 정확도가 높게 나왔다. 각 모형에서 전세특성을 고려 한 경우, 전세특성을 고려하지 않은 경우에 비해 예측 정확도가 모두 상승하여 전세특성이 가격 예측에 있어 중요한 요인으로 작용함이 입증되었다. 가격형성요인 중요도 분석 결과 랜덤포레스트와 회귀분석에서 전세특성 중 전세가격이 가장 중요한 요인으로 도출되었으며 MLP 모형에서도 상대적으로 중요한 요인으로 도출되었다. 본 연구를 통한 정책적 시사점 및 제언은 다음과 같다. 첫째, 시장 참여자 및 부동산 관련 이해 당사자에게 가격형성요인의 개념의 확장을 통한 공동주택 가격형성요인의 올바른 이해를 가능하게 하고자 한다. 둘째, 주택 안정화 정책의 수립에 있어 전세시장과 매매시장을 연계시켜 고려하는 것에 대한 시사점을 제공하여 이를 바탕으로 주택 안정화 정책에 있어서 매매시장에 관한 동시적이고효율성 있는 정책 방향을 제시한다. 마지막으로, 부동산 관련 데이터를 종합적으로 고려할 수 있는 기계학습 모형을 바탕으로 주택시장을 분석하고, 공시제도 및 주택 통계지표 작성 및 정책 수립에 있어 적용하는 경우 보다 효율적이고 예측 정확도 높은 주택 정책 수행이 가능할 것이다. 본 연구는 대전광역시 노은지구라는 세부시장을 대상으로 연구를 진행하여, 이를 다른 지역에 적용하는 경우 지역의 특성 및 시장의 특성이 상이하여 분석 결과가 다르게 나타날 수 있다는 한계점이 있다. 또한 본 연구에서 사용한 회귀분석과 기계학습 기법에 따른 순열특성중요도 분석결과는 분석 방법의 차이로 인해 각 순서의 비교 및 우열의 판단이 어려워 이에 대한 종합적 결론을 내리기가 어렵다는 한계점이 존재한다. 다만 기계학습 기법의 높은 예측력과 적용의 용이성에도 불구하고 존재하는 ‘블랙박스’ 모형이라는 한계점은 기계학습 기법의 발전과 더불어 추후 해결될 것으로 기대되며, 주택 시장 분석 및 주택 정책 활용에 있어 변수 간의 관계를 직접적으로 해석할 수 있는 헤도닉 모형도 상호보완 적으로 함께 분석하여 활용할 필요가 있다. There was a high increase of real eastate price in 2020. Deposit of ‘Jeonse’ made people buy a house easily, which is defined as ‘gap investment’, and it became one of the common trading types for buying house. Previous studies on the determinants of apartment prices have separately considered conception of ‘Jeonse’ market and housing market. The privious studies were focused on proving the priority between ‘Jeonse’ market and housing market by analyzing index chart. The purposes of this study are to figure out which factors including ‘Jeonse’ factors determine the apartment prices and offer political advice. The study was conducted by analyzing housing prices and ‘Jeonse’ prices of 35 apartment complexes in Noen district of Daejeon city. First, for the theoretical backgrounds, the concepts of housing market and Jeonse market, two different theory of housing market and Jeonse market, machine learning concepts and theories were reviewed. Second, based on the relevant theories, variables were set in consideration of the limitations and problems revealed in previous studies. Next, for the analysis methods, research model and hypothesis were established. Basic statistics and variable analysis were performed by using Hedonic model and additional analysis were conducted by applying machine learning model with permutation feature importance method to figure out the importance of ‘Jeonse’ factors. From the results of the analysis, ‘Jeonse’ determinants on apartment prices has a positive effect(+) to the apartment prices and the result showed prediction accuracy of apartments prices are in the order of random forest model and MLP model as machine learning model and Hedonic model. Also, the prediction accuracy of apartments prices in all three models was increased by adding the ‘Jeonse’ factors. Permutation feature importance showed that the most important factors in determinations of apartment price was ‘Jeonse’ price factor. The following implications and political advises are drawn from the results. First, It is intended to enable market participants and real estate-related stakeholders to have a correct understanding of the determinants on apartment prices by expanding the concept of price-forming factors with ‘Jeonse’ market. Second, in establishing housing stabilization policy, it is necessary to consider the connection between the ‘Jeonse’ market and the housing market. Third, since the characteristics of ‘Jeonse’ market have a significant impact on the apartment sales market, it is necessary to consider not only the movements of the apartment sales price but also the movements of the ‘Jeonse’ market when establishing policies related to the housing market. Finally, in the case of machine learning, various variables can be considered simultaneously and complexly, regardless of the correlation of variables, and a large amount of data can be applied comprehensively, which has the advantage of realizing high prediction accuracy and it will be possible to implement housing policies with high predictive accuracy and more efficiently. The limitations of this study are here. First, this study was conducted on a sub-market called Noeun District, Daejeon City, and 2020 market. When this study is applied to other regions, there is a limitation in that the analysis results might be different due to the different characteristics of the regions and sub-market. In addition, the machine learning technique used in this study has a limitation in that it is a 'black box' model, so it is difficult to express the direct relationship between variables as a formula, in comparision with a Hedonic model, despite of high prediction accuracy. This is a technical limitation of machine learning research. To overcome this, in this study, the importance of each factor was analyzed by applying the permutation feature importance method, and the direction of each variable to the housing price was confirmed by the Hedonic model. Despite of the high predictive power and easy access of machine learning techniques, the limitations of existing 'black box' models are expected to be resolved in the future with the development of machine learning techniques.
역학기반 기계학습을 활용한 철근콘크리트 보의 전단강도 예측 모델 개발
제 목 : 역학기반 기계학습을 활용한 철근콘크리트 보의 전단 강도 예측 모델 개발 본 연구에서는 철근콘크리트 보의 전단강도를 예측하기 위해 기계학습 기법을 적 용하였고, 데이터기반 기계학습 모델과 역학기반 기계학습 모델을 생성하여 비교하 였다. 총 193개의 전단파괴가 발생한 철근콘크리트 보의 실험데이터를 수집하였으 며, 이는 전단철근을 가진 실험체이다. 9개의 서로 다른 기계학습 알고리즘이 예측 성능 비교 및 최적의 모델을 선정하기 위해 사용되었다. 데이터기반 기계학습에는 8개의 실험변수가 입력특성으로 사용되었으며, 출력값은 전단강도이다. 우수한 예측 성능을 나타냈지만, 전역 해석인 특성 중요도 결과에서 일관되지 않은 결과를 도출 하여 신뢰성과 해석 가능성이 낮은 제한점이 특정되었다. 역학기반 기계학습 모델 은 역학이론을 입력특성으로 활용하여 학습 과정에서 구조적 상관성이 효과적으로 반 영될 수 있도록 설계하였다. 입력특성이 4개로 줄었으나, 데이터기반 모델과 유사한 예 측성능을 나타냈고, 특성 중요도 결과에서 일관된 결과를 도출하였다. 또한 국소 해석을 통해 구조역학적 메커니즘과 정합되는 결과를 확인하였고, 해당 모델이 전단저항 메 커니즘의 핵심 특성을 학습 과정에서 포착하였음을 시사한다. 해당 모델의 예측과 정을 수식화하고자, 역학기반 기계학습 모델의 예측값을 종속변수로 설정하고 회귀 분 석을 수행하여 폐쇄형 방정식을 제안하였다. 제안식은 기계학습 모델에 비해서는 소폭 낮은 정확도를 보였으나, 기존 설계기준 및 이론 기반 모델에 비해서 우수한 성능을 나 타내었고, 실험값과도 높은 일치도를 보였다. 따라서 기계학습 기법에 역학을 학습하도 록 내재화시키는 방안을 활용하면 철근콘크리트 보의 설계 시 충분히 활용될 수 있음을 나타낸다.