RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    검색결과 좁혀 보기

    선택해제
    • 좁혀본 항목 보기순서

      • 원문유무
      • 음성지원유무
      • 원문제공처
        펼치기
      • 등재정보
        펼치기
      • 학술지명
        펼치기
      • 주제분류
        펼치기
      • 발행연도
        펼치기
      • 작성언어

    오늘 본 자료

    • 오늘 본 자료가 없습니다.
    더보기
    • 무료
    • 기관 내 무료
    • 유료
    • KCI등재

      결측치를 포함한 데이터의 k-평균 군집분석 방법 비교

      양대경,명재성,이승훈,송주원 한국자료분석학회 2023 Journal of the Korean Data Analysis Society Vol.25 No.6

      Cluster analysis is an unsupervised learning method to find heterogeneous clusters that capture similarities among items and separate different items into different clusters. Various cluster analysis techniques have been proposed, and the k-means clustering method, which minimizes the sum of Euclidean distances between cluster centroids and individual entities, is widely recognized as a standard cluster analysis method. When data include missing values, it is challenging to conduct cluster analysis, because it is impossible to calculate distances between centroids of clusters and incomplete items, resulting in excluding classification of these items. Techniques have been suggested to handle missing values in k-means clustering, including conducting cluster analysis after imputation of missing values or cluster analysis based on available information. In this study, we explore methods to perform k-means cluster analysis on data with missing values and evaluate performance of these methods using a simulation. The results of simulation studies indicate that conducting k-means cluster analysis after imputation yields the better performance than the one based on available information. Among the various imputation methods, k-nearest neighbors imputation performed the best.

    • KCI등재후보

      FCAnalyzer: A Functional Clustering Analysis Tool for Predicted Transcription Regulatory Elements and Gene Ontology Terms

      Kim, Sang-Bae,Ryu, Gil-Mi,Kim, Young-Jin,Heo, Jee-Yeon,Park, Chan,Oh, Berm-Seok,Kim, Hyung-Lae,Kimm, Ku-Chan,Kim, Kyu-Won,Kim, Young-Youl Korea Genome Organization 2007 Genomics & informatics Vol.5 No.1

      Numerous studies have reported that genes with similar expression patterns are co-regulated. From gene expression data, we have assumed that genes having similar expression pattern would share similar transcription factor binding sites (TFBSs). These function as the binding regions for transcription factors (TFs) and thereby regulate gene expression. In this context, various analysis tools have been developed. However, they have shortcomings in the combined analysis of expression patterns and significant TFBSs and in the functional analysis of target genes of significantly overrepresented putative regulators. In this study, we present a web-based A Functional Clustering Analysis Tool for Predicted Transcription Regulatory Elements and Gene Ontology Terms (FCAnalyzer). This system integrates microarray clustering data with similar expression patterns, and TFBS data in each cluster. FCAnalyzer is designed to perform two independent clustering procedures. The first process clusters gene expression profiles using the K-means clustering method, and the second process clusters predicted TFBSs in the upstream region of previously clustered genes using the hierarchical biclustering method for simultaneous grouping of genes and samples. This system offers retrieved information for predicted TFBSs in each cluster using $Match^{TM}$ in the TRANSFAC database. We used gene ontology term analysis for functional annotation of genes in the same cluster. We also provide the user with a combinatorial TFBS analysis of TFBS pairs. The enrichment of TFBS analysis and GO term analysis is statistically by the calculation of P values based on Fisher’s exact test, hypergeometric distribution and Bonferroni correction. FCAnalyzer is a web-based, user-friendly functional clustering analysis system that facilitates the transcriptional regulatory analysis of co-expressed genes. This system presents the analyses of clustered genes, significant TFBSs, significantly enriched TFBS combinations, their target genes and TFBS-TF pairs.

    • KCI등재

      군집분석을 통한 K리그 축구팀 플레이스타일 분류

      김종원(Jongwon Kim),최형준(Hyongjun Choi) 한국체육측정평가학회 2021 한국체육측정평가학회지 Vol.23 No.1

      본 연구는 2020 K리그 경기에서 발생한 패스 관련 분석인자들을 이용하여 군집분석을 통해 K리그 팀들의 플레이스타일을 알아보고자 하였다. 2020 K리그 모든 팀들의 전 경기(스플릿 후 경기 제외)를 대상으로 하였으며, 연구의 대상이 된 경기 수는 총 132경기였으며, 양 팀의 자료를 각각 고려하였다(n=264). K리그 프로축구연맹 ‘데이터포탈’에서 제공받은 18개의 패스 관련 분석인자들을 Microsoft Office Excel 2007을 이용하여 정리하였고, 그 후 R 3.6.2를 이용하여 자료 처리하였다. 통계적 검증을 위하여 기술통계 분석(descriptive statistics analysis)을 실시한 후, 데이터 마이닝 기법 중 하나인 k-평균 군집분석(k-means cluster analysis)과 교차분석(cross-tabulation analysis)을 실시하였다. 본 연구의 군집분석을 통하여 얻어진 군집의 수는 3개였다. 절반 이상의 팀들이 군집 1에 속하였고, 군집2(전북, 울산, 강원)와 군집3(대구, 광주, 인천)에는 각각 3팀이 속하였다. 최상위 팀인 1위 팀 전북과 2위 팀 울산이 속한 군집2는 다른 군집들과 비교해 공격 1/3지역 패스 비율, 숏 패스 비율, 전진 패스 비율을 제외한 나머지 15개의 분석인자들에서 가장 높은 평균값을 나타냈고, 군집3의 경우 가장 낮은 평균값을 보였다. 분석인자들 간의 유사성을 이용하여 군집을 나누는 방법으로 직접적인 팀의 플레이스타일을 표현하는데 한계가 있지만, 본 연구에서 사용된 분석인자들을 통해 비슷한 유형의 팀들을 군집하는데 의미가 있다. The purpose of this study was to identify the playing styles of football clubs in K-League through cluster analysis using performance indicators related to pass. All matches excepted to split matches were used for analysis and all data were provided from Korea Football League(n=264). All data were preprocessed on Microsoft Office Excel 2007 and statistical analysis was conducted on R 3.6.2. Descriptive statistical analysis was firstly used to calculate means and standard deviations for each performance indicators and then k-means cluster analysis, one of the data mining method, was conducted to identify clusters. Finally, cross-tabulation analysis was used to identify K-League teams into each cluster. Three clusters were identified and Jeonbuk, Ulsan and Gangwon was included in cluster 2 whilst Daegu, Gwangju and Incheon was included in cluster 3. The other teams were included in cluster 1. Cluster 2 had greater performance indicators related to pass rather than other clusters. Although cluster analysis, grouping performance indicators in such a way that performance indicators in the same cluster are more similar each other compared to other clusters, could not determine accurate playing styles in football, it is literally meaningful to group the similar type of teams. There needs to be a great interpretation of the characteristics of the formed clusters.

    • KCI등재

      데이터마이닝을 활용한 인구집단 구성과 경제적 불평등 평가

      민대기 아시아.유럽미래학회 2016 유라시아연구 Vol.13 No.1

      The importance of public pension system recently has been growing because of an ageing population in Korea. One of the major purposes of public pension system is to reduce economic inequalities across population groups that have been mainly defined by a single variable such as age, income, sex, etc in the literature. This paper aims to evaluate how well the population groups are practically and clearly designed in the literature so as to specify economic inequalities across different groups. For this purpose, we investigate the retired household sample obtained from KLIPS (Korean Labor and Income Panel Study) and conduct the clustering analysis to organize population groups. The clustering analysis consists of four steps: variable selection, hierarchical clustering analysis to determine the number of clusters and centroid, non-hierarchical clustering analysis (we use K-means clustering analysis) and define population groups, and describing the characteristics of population groups. The clustering analysis divides the population into six subgroups that have similar characteristics. We compare the characteristics of the population groups obtained from the clustering analysis with those designed in the literature with respect to household, income, consumption and asset. From the comparison, we observe that the assumptions on population groups in the literature are somewhat different from the findings from the clustering analysis. In particular, the period of employment and the full-time period are significantly different from the literature. The clustering analysis reveals that the period of employment is approximately 31-45 years, which is shorter than the period of employment assumed in the literature. Furthermore, few research has clearly considered the full-time period, which accounts for 14%-45% of the total period of employment. We believe that the findings in this study provides meaningful information on how to practically organize population groups as a part of designing public pension system. For example, researchers are encouraged to consider the employment periods and full-time periods obtained from the data analysis when considering the significant impact of working periods on pension payment and benefits. Population should be well grouped so as to specify the difference in economic inequalities across groups. In general, various variables are jointly involved in determining the economic inequalities, but populations have been simply grouped by a single variable such as age, gender, education level, household size, etc. in the literature. This paper employs the Generalized Entropy (GE) measure of economic inequality and its decomposition to evaluate the contributions of several variables on the overall economic inequality. For this analysis, populations are first grouped by several variables such as clusters obtained from the analysis, sex, education level, period of employment and household size. Then, we measure the inequaltiy index and its decomposition for monthly income, monthly consumption and asset. Grouping populations using the clustering analysis specifies the economic inequality more precisely provides better measure of similarity within a group and difference between groups. This result supports the effectiveness of the clustering analysis in grouping populations.

    • KCI등재

      혁신클러스터 활성화를 위한 클러스터분석(Cluster Analysis) 연구

      이원일(Lee, Won-Il) 한국산학기술학회 2012 한국산학기술학회논문지 Vol.13 No.8

      본 논문은 혁신클러스터 추진의 전략방향 설정을 위해서 클러스터 현황을 클러스터분석을 통하여 살펴보았 다. 클러스터 분석은 클러스터 내에서 일어나는 다양한 형태의 협력을 파악하고 정책적인 대안의 도출을 위해 활용되 는 실제적인 분석도구이다. 이에 본 논문에서는 광교테크노밸리의 발전단계에 따른 협력현황 분석을 위하여 산학연 협력체계 분석과 진단을 위한 클러스터분석(Cluster Analysis)을 실시하였다. 클러스터 분석결과 광교테크노밸리내 입 주기업의 산학연 협력경험은 67.3%로 매우 높은 것으로 나타났다. 대학과 기업과는 연구개발, 연구기관과는 장비활용 중심으로 협력하고 있었다. 또한, 산학연 협력의 지원정책수요는 협력기관 현황제공, 기술별 정보취득 지원 등 다양한 부문에서 협력의 수요가 존재하고 있었다. 이러한 분석결과를 토대로 혁신클러스터 구성단계를 넘어서 확장단계로 발 전하기 위해서는 첫째, 혁신클러스터 발전을 위한 새로운 비전을 조속히 제시하고 단지인근에서 입주기업의 산학연 협력의 활성화를 위한 지원을 추진해야 한다. 둘째, 혁신클러스터 단계별 발전을 위한 통합적인 지원역량 강화가 필 요하다. 마지막으로 인근의 타 혁신거점과 정책적 네트워크 구축을 통해서 타 혁신클러스터의 자원과 역량을 활용할 수 있는 지원망 구축이 필요하다. This research focused on the cluster analysis for the vitalization of the innovation cluster, Gwanggyo Technovalley. The study was performed based on both theoretical study and quantitative and qualitative study approaches. Particularly, questionnaire survey was performed for the cluster analysis of the innovation cluster. The major determinants for vitalization of the innovation cluster, Gwanggyo Technovalley can be summarized as follows; the strategy formulation for the development of the innovation cluster, the enhancement of the host institution capability and gradual enlargement of the role of the host institution. In terms of the needs of times, this study regarding the cluster analysis for the vitalization of the innovation cluster, Gwanggyo Technovalley is anticipated to be a good reference for the R&D organizations and technology cluster participants in coming years.

    • KCI등재

      군집분석을 통해 살펴본 1인 가구의 연령대별 소비지출패턴

      성영애 한국소비자학회 2013 소비자학연구 Vol.24 No.3

      The purposes of this study were to identify consumption expenditure patterns of one-person households of different age groups, find out the socio-economic characteristics of the identified clusters, and then compare the patterns of different age groups. For these purposes, one-person households were separated by three age groups : younger households(less than 35 years old), middle aged households(35-64 years old or less), and older households(over 64 years old). The consumption expenditure pattern means the way consumption categories are combined to form a way of life as a whole(Chung 1998, p.39). The consumption categories of this study employed those of the 2012 Household Income & Expenditure Survey of Statistics Korea that is the data of the study. The 12 standard consumption categories are ‘Food and non-alcoholic beverages’, ‘Alcoholic beverages and tobacco’, ‘Clothing and footwear’, ‘Housing, water, electricity and other fuels’, ‘Furnishings and household equipment’, ‘Health’, ‘Transport’, ‘Communication’, ‘Recreation and culture’, ‘Education’, ‘Restaurants and hotels’, and ‘Miscellaneous goods and services’. Cluster analysis which is a statistical method for categorizing the households into similar groups was utilized. The variables used as criteria variables are budget shares allocated to each consumption categories : i.e. the proportions of the expenditures of each consumption categories in total consumption expenditure. K-means cluster analyses were conducted using SPSS. The data of the study were the annual raw data of the 2012 Household Income & Expenditure Survey of Statistics Korea. Total sample consisted of 10,400 households of the whole country except farm and fishery households. Among them, 1,653 households were one-person households. Weighted results were presented. The main findings are as follows: Four consumption expenditure patterns are identified for each age groups. The clusters are named according to their dominant budget shares. For younger one-person households, the identified clusters are diverse activity oriented(48.6%), restaurants and hotels-dominated(25.6%), housing-dominated(21.9%) and transport-dominated(4%). Diverse activity oriented cluster consists the largest proportion of young one-person households. These households allocated their budget to diverse consumption categories and it shows their active lives compared with the households in other clusters. Of the younger one-person households, housing-dominated cluster shows the lowest economic level, that is the consumption expenditure level of this cluster is the lowest among younger one-person households. The unemployment rate is the highest and the household rate in budget deficit is also the highest. For the one-person households in the middle ages, the identified clusters are restaurants and hotels-dominated(38.2%), food-dominated(32.3%), housing-dominated(25.7%) and transport- dominated(7.1%). Except the food-dominated cluster, the clustering is similar with younger households. Even though the names are identical but the consumption expenditure structure is different. Household budget is relatively evenly allocated in various consumption categories for the one-person households in the middle ages. For the older one-person households, food-dominated(37.4%), housing-dominated(22.5%), balanced(22%) and health-dominated (18%) are identified as clusters. In general, the budget of the older one-person households are more likely tend to be intensively allocated to top three consumption categories for their all clusters than the other clusters of younger households. The levels of income and expenditure are the lowest for food-dominated cluster. But housing-dominated cluster and health-dominated cluster seem to be more problematic in terms of the consumption expenditure structures and the inability of solving the problems. The similar expenditure pattern which is identified across all ages is ho...

    • KCI등재

      차원축소를 통한 결측자료의 군집분석

      송주원 한국자료분석학회 2020 Journal of the Korean Data Analysis Society Vol.22 No.2

      Cluster analysis classify similar observations into the same cluster and different observations into different clusters. When data include many variables, reduced dimension clustering methods have been suggested instead of the standard clustering methods. The joint analysis of dimension reduction and clustering is known to perform better than tandem analysis that sequentially conducts dimension reduction and clustering. On the other hand, most data include missing values. When cluster analysis is conducted with incomplete data, incomplete observations can not be classified into any group. To avoid this problem, it is common to impute missing values before conducting cluster analysis. In this study, we suggest a method for combining dimension reduction k-means clustering and missing data imputation. The suggested method has an advantage to accurate classify observations through imputation using cluster information. A simulation is conducted to evaluate performance of the suggested method and compare the result with the one based on tandem analysis. The suggested method using an appropriate dimension reduction k-means clustering showed lower misclassification rates than tandem analysis. 군집분석은 유사한 특성들을 지닌 관측값들을 같은 군집으로, 다른 특성들을 지닌 관측값들은 서로 다른 군집으로 분류하는 분석 기법이다. 많은 변수를 포함한 고차원 자료에서는 일반적인 군집분석 대신 차원축소를 통하여 군집분석을 실시하는 방법들이 제안되어 왔다. 주성분 분석을 통해 차원을 축소한 후 축소된 차원에서 군집분석을 실시하는 직렬분석 방법보다 차원축소와 군집분석을 결합하여 동시에 실시하는 방법들이 더 우수한 성능을 보인다는 것이 알려져 있다. 한편, 대부분의 자료는 결측값을 포함하고 있는데 결측값이 포함된 자료에 대하여 군집분석을 실시하는 경우 불완전하게 관측된 자료들은 어느 군집으로도 분류되지 않는 문제가 발생한다. 따라서 군집분석을 실시하기 전에 먼저 결측값 대체를 실시하는 것이 일반적이다. 본 연구에서는 고차원 결측자료에 대하여 차원축소를 통한 k-평균 군집분석을 실시할 때 결측값 대체를 결합하여 실시하는 방법을 제안한다. 이 방법은 군집 정보를 이용한 결측값 대체를 통해 정확한 차원축소를 통한 군집분석이 가능하게 하는 장점을 지닌다. 제안된 방법은 모의실험을 통해 성능을 평가하였고 결측값을 대체한 후 대체된 자료에 대하여 차원축소를 통한 군집분석을 실시하는 직렬식 분석방법과 비교하였다. 제안된 방법은 적절한 차원축소를 통한 k-평균 군집분석을 실시한다면 직렬식 분석보다 오분류율이 낮게 나타났다.

    • KCI등재

      2022 FIBA 남자농구 아시아컵 경기기록을 통한 선수들의 군집분석

      예원진,이성노,유덕수 한국체육과학회 2023 한국체육과학회지 Vol.32 No.5

      The purpose of this study is to determine whether there is a difference in performance level among participating players and the participation of star players in the 30th Men's Basketball Asia Cup in 2022, based on box score data provided on the official Asian Cup website. Clustering was performed using k-means cluster analysis, one of the machine learning techniques. The subject of this study was the game records of 189 players out of 16 teams participating in the tournament obtained through the official records of the 2022 Men's Basketball Asian Cup tournament, and differences in the players' performance characteristics were compared through a total of 21 variables. This study used the statistical program Python version 3.10.1 along with the library to perform cluster analysis. All significance levels for statistical analysis were set to .05, and the results obtained were as follows. First, as a result of cluster analysis using official records provided by the 2022 Men's Basketball Asian Cup competition, players from each country at the Basketball Asian Cup could be classified into three clusters. Second, there was a statistically significant difference in performance by cluster between cluster 1, cluster 2, and cluster 3, and the post-hoc test results showed that the players in cluster 2 had the best performance. Next, it was confirmed in the order of cluster 3 > cluster 1. Third, the abnormality detection results showed that there were 9 abnormal values among the cluster analysis results. Looking at this, ‘Q. Zhou’, ‘H. EHaddadi’, ‘A. Al Dwairi’, ‘G. ‘RA’, ‘D. Chism’, ‘A. Alderazi’, ‘W. Artino’, ‘M. Bolden’, ‘W. Arakji’ 9 were found to be basketball players. Fourth, the outlier players and all remaining players are divided into points, shots made, shot attempts, 2-point shots made, 2-point shots attempted, free throws made, free throw attempts, offensive rebounds, defensive rebounds, turnovers, There was a significant difference in block and efficiency. Summarizing the results of this study, from the perspective of a leader, strategies and tactics for the next men's basketball Asia Cup tournament can be designed and implemented based on the characteristics of the players in each group. On the player side, based on the performance characteristics of players in each group, it is possible to determine the individual's performance level and the gap in performance with other excellent players in this competition.

    • KCI등재

      결측자료의 k-평균 군집분석

      송주원 한국자료분석학회 2017 Journal of the Korean Data Analysis Society Vol.19 No.2

      Cluster analysis is an analysis technique to classify observations with similar characteristics into the same cluster. The k-means cluster analysis conducts grouping of observations based on an optimization method minimizing the sum of Euclidean distances between observations and their cluster centers. In real data, missing values often occur in some variables, and when cluster analysis is conducted for missing data, it is common to exclude observations with missing values. However, in this case, missing values cannot be classified into any group, and it may cause biases in estimating cluster centers. Therefore, to include observations with missing values in cluster analysis, it is often to impute missing values and conduct cluster analysis using imputed data. A disadvantage of this imputation approach is to conduct imputation without using cluster information. In this study, we propose methods to impute missing values using cluster information. Simulation is conducted to compare performance of the suggested imputation method with the one based on imputation without using cluster information. The proposed imputation method provides better results than the one ignoring cluster information. 군집분석은 유사한 특성을 지닌 관측치들을 동일한 그룹으로 분류하는 분석 기법이다. k-평균 군집분석은 관측치들과 군집 평균의 유클리디언 거리의 합을 최소화하는 그룹을 찾는 최적화 기법을 통해 자료를 군집으로 분류한다. 실제 자료의 경우 일부 변수에서 결측이 발생하는 경우가 흔하며 결측을 포함한 자료에 대하여 군집분석을 실시하는 경우 결측이 발생한 관측치를 제거한 후 분석을 실시하는 것이 일반적이다. 하지만 이 경우 결측이 발생한 자료는 어느 군집에도 할당할 수 없고 각 그룹의 평균의 추정에 편향이 발생할 가능성이 높다. 따라서 결측치를 포함한 자료를 군집분석에 포함하기 위하여 흔히 사용되는 방법은 결측값에 대해 대체를 실시한 후 대체된 자료에 대하여 군집분석을 실시하는데 이 경우 군집 정보를 포함하지 않고 대체를 실시하는 단점을 지닌다. 따라서 본 연구에서는 결측치에 대한 대체를 실시할 때 군집 정보를 이용하여 대체하는 방법을 제안한다. 모의실험을 통해 본 연구에서 제안한 방법을 군집 정보를 포함하지 않고 대체를 실시한 후 군집분석을 실시하는 경우와 비교하였는데 본 연구에서 제안한 대체 방법이 더 나은 결과를 보였다.

    • KCI등재

      k-중앙개체 군집방법을 이용한 한국 프로농구선수의 군집화

      한수철,전수영,진서훈 한국자료분석학회 2008 Journal of the Korean Data Analysis Society Vol.10 No.6

      주어진 데이터의 개체들을 비슷한 특징을 가지는 소그룹으로 나누어 그 그룹들의 특징이나 대표성을 찾는 분석 과정을 군집분석이라고 한다. 군집분석은 크게 분리 군집방법과 계층적 군집방법으로 구분할 수 있다. 본 연구에서는 계층적 군집방법을 이용하여 군집의 수를 정하고, 분리 군집방법 중 하나인 k-중앙개체 군집방법을 적용하여 한국 프로농구선수들의 군집화를 시도해 보았다. 프로농구선수들의 데이터는 몇몇 변수들에 있어 특이치가 존재하기 쉽다. 따라서 이런 경우에는 특이치에 영향을 크게 받는 k-평균 군집방법을 적용하는 것보다는 특이치에 덜 민감한 k-중앙개체 군집방법의 활용이 좋은 결과를 줄 수 있다. k-중앙개체 군집방법의 구현을 위해 PAM(partitioning around medoids) 알고리즘을 이용하였다. 군집분석결과 3개의 군집으로 선수들을 군집화하였고, 각 군집의 특징을 파악하였다. Cluster analysis is one of statistical methods for finding groups so that objects in the same group are similar each other and objects in the different group are dissimilar. There are two distinctive techniques in cluster analysis. One is hierarchical method the other is partitioning method. In this study, we built the clusters from the data of korean professional basketball players. The hierarchical method was used for finding the proper number of clusters and k-medoids clustering which is one of partitioning method was used for building clusters. The professional basketball players data generally has outliers in several variables. Therefore, instead of applying k-means clustering for this kind of data, k-medoids clustering which is not affected a lot by outliers can give a better result than that of k-means clustering. In order to implement k-medoids clustering PAM(partitioning around medoids) algorithm was used. The resulting clusters are obtained as three distinguished clusters and the characteristics of each cluster are summarized.

    연관 검색어 추천

    이 검색어로 많이 본 자료

    활용도 높은 자료

    해외이동버튼