픽셀 기반 강화학습에서 데이터 증강은 과적합을 완화하여 성능을 크게 향상시키는 핵심 기법으로 알려져 있다. 특히 DrQ-v2는 단순한 random shift 증강만으로 DeepMind Control Suite의 여러 연속 제...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T17535445
부천 : 가톨릭대학교 (성심) 일반대학원, 2026
학위논문(석사) -- 가톨릭대학교 (성심) 일반대학원 , 수학과 응용수학 전공 , 2026. 8
2026
한국어
경기도
vii, 43 p ; 26 cm
지도교수: 이병준
I804:41027-200001007716
0
상세조회0
다운로드픽셀 기반 강화학습에서 데이터 증강은 과적합을 완화하여 성능을 크게 향상시키는 핵심 기법으로 알려져 있다. 특히 DrQ-v2는 단순한 random shift 증강만으로 DeepMind Control Suite의 여러 연속 제...
픽셀 기반 강화학습에서 데이터 증강은 과적합을 완화하여 성능을 크게 향상시키는 핵심 기법으로 알려져 있다. 특히 DrQ-v2는 단순한 random shift 증강만으로 DeepMind Control Suite의 여러 연속 제어 과제에서 최고 수준을 달성하였다. 그러나 random shift가 다른 증강들보다 왜 효과적인지에 대한 정량적 분석은 부재하며, 후보 증강의 우열을 알기 위해서는 각 증강에 대해 완전한 학습을 반복해야 하는 실용적 비용이 따른다. 본 연구는 이러한 공백을 좁히기 위해 증강의 좋고 나쁨을 구분할 수 있는 정량적 축을 식별하고, 이를 결합한 점수 함수를 제안하며 유효성 분석 프레임워크를 구축한다. 본 연구는 효과적인 증강이 다양성과 행동 안정성의 두 조건을 동시에 만족한다는 가설을 제시한다. 다양성은 증강이 입력에 충분한 변화를 제공하여 𝐸𝑛𝑐𝑜𝑑𝑒𝑟의 과적합을 방지하는 정도이며, 행동 안정성은 증강이 과제의 행동 구조를 보존하는 정도이다. 두 조건을 각각 픽셀 변화량 𝐶(𝑇)와 학습된 기준 에이전트의 행동 거리 𝐷(𝑇) 로 정량화하고, 이 둘을 결합한 유효성 점수 𝑆(𝑇) = √𝐶/𝐷(𝑇)를 정의한다. 𝑆(𝑇)에서 분자 √𝐶는 다양성 조건을, 분모 𝐷는 행동 안정성 조건을 각각 반영한다. 이 프레임워크를 검증하기 위해 DeepMind Control Suite의 세 환경(Cartpole Swingup, Cheetah Run, Quadruped Walk)에서 9개 증강, 3개 시드로 총 81회의 DrQ-v2 학습을 수행하였다. 시드 평균 𝑛 = 27에 대해 𝑆(𝑇)와 정규화 성능 사이의 Pearson 상관계수를 계산한 결과, 𝑟 = 0.454, 𝑝 = 0.017로 통계적으로 유의한 양의 상관이 확인되었다. 변화량 𝐶 단독은 성능과 사실상 무상관( 𝑟 = 0.039)이고 행동 안정성 1/𝐷 단독은 변화 부족 증강을 구분하지 못하는 구조적 한계를 보였으며, 통합 분석에서 두 축의 결합이 단일 축보다 높은 설명력을 보였다. 또한 환경별 분해에서 행동 차원이 증가할수록 1/𝐷 단독의 상관이 두드러지게 증가하는 것이 관찰되어, 두 축의 상대적 기여도가 환경의 수학적 특성에 따라 달라질 수 있음을 확인하였다. 본 연구는 한 번 학습된 기준 에이전트 위에서 후보 증강들을 추가 학습 없이 정량 평가할 수 있는 최소 프레임워크의 가능성을 보였다. 본 점수의 결정계수는 약 0.21로 잔차의 상당 부분이 두 축 외의 요인에 기인하나, 이는 향후 추가 축 도입과 환경 적응적 가중치 학습을 통해 채워 나갈 여지를 향후 과제로 한다. ※ 본 논문은 저자가 한국산업응용수학회(KSIAM)에 투고를 준비 중인 연구와 내용을 공유하며, 동일한 연구 결과에 기반하여 작성되었다.
다국어 초록 (Multilingual Abstract)
In pixel-based reinforcement learning, data augmentation has emerged as a key technique that mitigates overfitting and substantially improves performance. In particular, DrQ-v2 achieves state-of-the-art performance on a range of continuous control tas...
In pixel-based reinforcement learning, data augmentation has emerged as a key technique that mitigates overfitting and substantially improves performance. In particular, DrQ-v2 achieves state-of-the-art performance on a range of continuous control tasks in the DeepMind Control Suite using only a simple random shift augmentation. However, no quantitative analysis exists for why random shift is more effective than other augmentations. Identifying the relative quality of candidate augmentations requires a full training run for each, which incurs substantial practical cost. To address this gap, we identify a quantitative pair of axes along which the quality of augmentations can be distinguished, propose a score function that combines them, and construct an augmentation validity analysis framework.
We hypothesize that effective augmentation must simultaneously satisfy two conditions: diversity and action stability. Diversity refers to the extent to which an augmentation introduces sufficient variation in the input so as to prevent overfitting of the Encoder, and action stability refers to the extent to which an augmentation preserves the task-relevant action structure. We quantify these two conditions respectively as the pixel change 𝐶(𝑇) and the action distance 𝐷(𝑇) of a trained reference agent, and define the validity score 𝑆(𝑇) = √𝐶(𝑇)/𝐷(𝑇) as their combination. In 𝑆(𝑇), the numerator √𝐶(𝑇) reflects the diversity condition, while the denominator 𝐷 reflects the action stability condition. To validate the framework, we conducted a total of 81 DrQ-v2 training runs across three environments in the DeepMind Control Suite (Cartpole Swingup, Cheetah Run, Quadruped Walk), with nine augmentations and three seeds each. With seed-averaged 𝑛 = 27, the Pearson correlation between 𝑆(𝑇) and normalized performance was 𝑟 = 0.454 (𝑝 = 0.017), indicating a statistically significant positive relationship. The change magnitude 𝐶 alone was essentially uncorrelated with performance (𝑟 = 0.039), and the action stability term 1/𝐷 alone exhibited a structural limitation in that it could not distinguish augmentations with insufficient variation; the combined score consistently outperformed either single axis in explanatory power. In the per-environment decomposition, the correlation of 1/𝐷 alone was found to increase monotonically with the action dimension of the environment, suggesting that the relative contribution of the two axes may depend on the mathematical characteristics of the environment.
We demonstrate the feasibility of a minimal framework in which candidate augmentations can be quantitatively evaluated without additional training, on top of a single pre-trained reference agent. The coefficient of determination of the score is approximately 0.21, indicating that a substantial portion of the residual variance is attributable to factors outside the two axes; this leaves explicit room to be filled by introducing additional axes and learning environment-adaptive weights in future work.
목차 (Table of Contents)