본 연구의 목적은 머신러닝 모델 투구 품질 평가 모델을 개발하여 그 정확도를 검증하고, 이를 기반으로 실투를 분류하여 실투와 선수 성과 지표 간의 상관관계를 분석하는 데 있다. 본 연구...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T17372331
서울 : 국민대학교 일반대학원, 2025
학위논문(석사) -- 국민대학교 일반대학원 , 스포츠애널리틱스전공 , 2026. 2
2025
한국어
세이버 메트릭스 ; 야구 분석 ; 투구 품질 ; Sabermetrics ; Baseball Analysis ; Pitch Quality
서울
vi, 65 ; 26 cm
지도교수: 이미영
I804:11014-200000948420
0
상세조회0
다운로드본 연구의 목적은 머신러닝 모델 투구 품질 평가 모델을 개발하여 그 정확도를 검증하고, 이를 기반으로 실투를 분류하여 실투와 선수 성과 지표 간의 상관관계를 분석하는 데 있다. 본 연구...
본 연구의 목적은 머신러닝 모델 투구 품질 평가 모델을 개발하여 그 정확도를 검증하고, 이를 기반으로 실투를 분류하여 실투와 선수 성과 지표 간의 상관관계를 분석하는 데 있다.
본 연구에 사용된 데이터는 2022년~2025년 MLB에서 발생한 모든 투구와 2025년 MLB 규정이닝과 규정타석을 충족한 선수의 성과 지표를 대상으로 하였다. 파이썬 라이브러리인 pybaseball을 통해 물리적·상황적 변인이 포함된 Statcast 투구 데이터와 성과 지표를 수집하였다. 투구 품질 평가 모델은 2022년~2024년의 투구의 물리적 변인과 경기 상황 변인을 학습하여, 11가지 투구 결과를 예측하도록 구성하였다. 모델의 정확도는 2025년 자료를 기반으로 검증하였으며, Top-k 정확도, F1-score, RMSE, MAE를 평가지표로 사용하였다. 학습된 모델의 예측값을 기반으로 투구의 홈런 발생확률을 산출하여, 이 값이 3% 이상일 경우 실투로 분류하였으며, 실투의 타당성을 검증하기 위해 실투와 비실투 간 대응표본 차이검증과 카이제곱 검정을 적용하였다. 실투를 분류한 뒤, 투수의 실투 비율, 타자의 실투 스트라이크 허용 비율, 스윙 비율, 인플레이 비율과 선수 성과 지표 간의 상관관계를 분석하였다.
연구 결과에 의하면, 본 연구에서 개발한 투구 품질 평가 모델의 Top-1 정확도는 50% 수준으로 다소 낮았으나, Top-5 정확도는 90% 수준으로 확인되었다. 실투와 비실투 간 대응표본 차이검증 결과, 기대 타율(xBA, p=.041), 기대 가중 출루율(xwOBA, p<.001)에서 통계적으로 유의한 차이가 있었다. 또한 카이제곱 검정결과, 비실투에 비해 실투에서 인플레이 비율, 홈런 비율, Barrel 비율이 유의하게 높은 것으로 확인되었다. 한편 상관분석 결과, 투수의 실투 비율이 수비 무관 평균자책점(FIP, r=.558), 9이닝 당 피홈런(HR/9, r=.540)과 뚜렷한 정적 상관관계를 나타냈다. 이에 반해, 타자의 실투 대응 방식과 성과 지표 간에는 낮은 상관관계를 보였다.
본 연구에서 개발한 투구 품질 평가 모델은 단일 범주 예측에 한계가 존재하였지만, 실제 결과와 근접하게 예측하는 경향성을 보였다. 정확도가 검증된 선행연구가 없어 성능 비교에 제약이 있으나, 투구 품질 성능의 기준점을 제시했다는 의의가 있다. 한편, 본 연구에서 제시한 실투 분류 기준은 장타가 발생할 위험이 높은 투구를 잘 선별하는 것으로 보인다. 특히 단순 위치 정보에 의존하거나 주관적 판단에서 벗어나, 다양한 변인을 고려해 실투를 객관적으로 분류했다는 데 의의가 있다. 또한, 투수의 실투 비율과 성과 지표 간의 뚜렷한 상관관계는, 실투 비율이 투수의 경기력을 설명하는 지표로 활용될 수 있음을 시사한다. 반면 타자의 실투 대응 방식과 경기력 간의 관계를 규명하기 위해서는, 향후 각 선수의 특성을 고려한 더 정밀한 분석이 필요할 것으로 보인다.
다국어 초록 (Multilingual Abstract)
The purpose of this study is to develop a machine learning–based Pitch Quality Evaluation Model, to validate its predictive accuracy, and to classify “mistake pitches” in order to examine the relationship between mistake pitches and player perfo...
The purpose of this study is to develop a machine learning–based Pitch Quality Evaluation Model, to validate its predictive accuracy, and to classify “mistake pitches” in order to examine the relationship between mistake pitches and player performance indicators.
This study included all pitches thrown in Major League Baseball (MLB) from 2022 to 2025, along with performance metrics for players who met the qualified innings pitched or plate appearance thresholds during the 2025 season. Statcast pitch-level data, including physical and situational variables, as well as player performance metrics, were collected using the Python library pybaseball.
The Pitch Quality Evaluation Model was trained on physical and situational variables from 2022 to 2024 to predict 11 distinct pitch outcomes. The model’s accuracy was verified using 2025 data, employing Top-k accuracy, F1-score, RMSE, and MAE as evaluation metrics. Based on the trained model’s predictions, the probability of a home run was calculated for each pitch; pitches with a probability of 3% or higher were classified as mistake pitches. To examine the validity of this classification, paired t-tests and Chi-square tests were conducted between mistake and non-mistake pitches. Subsequently, correlations were analyzed between player performance metrics and variables such as the pitcher’s mistake rate, and the batter’s mistake strike allowance rate, swing rate, and in-play rate.
he research results indicated that while the model's Top-1 accuracy was somewhat low at approximately 50%, the Top-5 accuracy reached a reliable level of 90%. The paired t-test results showed statistically significant differences in Expected Batting Average (xBA, p=.041) and Expected Weighted On-base Average (xwOBA, p<.001) between mistake and non-mistake pitches. Furthermore, Chi-square tests confirmed that In-play, Home Run, and Barrel rates were significantly higher for mistake pitches compared to non-mistake pitches. Correlation analysis revealed a distinct positive correlation between a pitcher’s mistake rate and Fielding Independent Pitching (FIP, r=.558) as well as Home Runs per 9 innings (HR/9, r=.540). In contrast, the batter’s response to mistake pitches showed a low correlation with performance indicators.
Although the developed model exhibited limitations in single-category prediction, it demonstrated a strong tendency to approximate actual outcomes. While direct performance comparisons were constrained by the lack of evidence from previous studies, this study was critically important in establishing a baseline for pitch quality evaluation. The proposed classification criteria for mistake pitches appear to effectively identify pitches with a high risk of yielding extra-base hits.
This study holds significance in objectively classifying mistake pitches by considering diverse variables, moving beyond subjective judgments or reliance solely on location data. Moreover, the distinct correlation between a pitcher’s mistake rate and performance metrics suggests that mistake rate can serve as a key indicator for explaining pitcher performance. Conversely, further precise analysis considering individual player characteristics will be necessary to clarify the relationship between a batter’s response to mistake pitches and their overall performance.
목차 (Table of Contents)