
http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
Galaxies and Their Cosmic Variance in the First Billion Years
Trapp, Adam Charles ProQuest Dissertations & Theses University of Cali 2022 해외박사(DDOD)
Cosmic variance is the intrinsic scatter in the number density of galaxies due to fluctuations in the large-scale dark matter density field. We begin by presenting a flexible analytic model of cosmic variance in the high redshift Universe (z ~ 5–15). We find that cosmic variance in the luminosity function of galaxies at these times is dominated by the variance in the underlying dark matter halo population, and not by differences in halo accretion nor the specifics of our stellar feedback model. We also find that cosmic variance dominates over Poisson noise except for the brightest sources or at very high redshifts (z ≳ 12). We provide a linear approximation of cosmic variance via a public Python package galcv. We then develop a statistical framework that folds our model of cosmic variance into the measurement of the galaxy luminosity function. Through this framework, we forecast the performance of several major upcoming James Webb Space Telescope (JWST) galaxy surveys. We find that they can constrain field matter densities down to the theoretical limit imposed by Poisson noise and unambiguously identify over-dense (and under-dense) regions on transverse scales of tens of comoving Mpc. We then apply this framework to a real Hubble Space Telescope (HST) data set at z = 6–8, providing a new measurement of the luminosity function of galaxies, and for the first time, measure the underlying densities of the survey fields, including the most over/under-dense HST fields. We show that the distribution of densities is consistent with current predictions for cosmic variance. Finally, we develop the first quantitative, statistically robust framework to infer the underlying density and ionization environment of regions with elevated densities of Lyman-α emitters (LAEs). We apply this framework to an actual observation of 14 LAEs in a ~50,000 cMpc3 region at z = 6.93, obtaining a measurement of that region's density and ionization state, and a constraint on the average ionization fraction of the Universe.
Pricing variance swaps under stochastic volatility models
김시우 Graduate School, Yonsei University 2020 국내박사
분산스왑은 장외파생상품으로 기초자산의 변동성에 크게 의존한다. 따라서, 시장을 더 잘 표현할 수 있는 분산스왑의 공정행사가를 구하기 위하여 Gatheral의 이중 평균 회귀 모형, 이중 지수 Ornstein-Uhlenbeck 모형과 같은 확률변동성 모형을 도입한다. 여기서, 이중 평균 회귀 모형은 3 요소 모형로 시장 분산과 잘 부합하며 이중 지수 Ornstein-Uhlenbeck 모형은 로그변동성이 Ornstein-Uhlenbeck 프로세스를 따르는 2 요소 확률변동성 모형으로써 짧의 시간구간의 수익 분포를 표현하는데 유용하다. 또한, Heston 모형에 CIR 프로세스를 따르는 확률이자율을 추가한 Heston-CIR 모형을 도입하여 일반화된 분산스왑에 대한 확률이자율의 영향력을 확인한다. 이 논문에서 우리는 소개한 모형에서의 바닐라 분산스왑의 닫힌 또는 근사 해를 구하고 확률이자율을 포함한 일반화된 분산스왑의 해석적 해를 도출한다. 수치분석에서는 몬테-카를로 시뮬레이션을 이용하여 도출한 해의 타당성을 면밀히 살펴보고 모형 매개변수의 영향력을 분석하며 분산스왑에 대한 주어진 모형들의 효과성을 보여준다. A variance swap is an OTC contract, which highly depends on the volatilities of the underlying asset. To derive the fair strike prices of the variance swaps that capture the market better, we adopt various stochastic volatility models such as the double-meanreverting model introduced by Gatheral and the double exponential Ornstein-Uhlenbeck model. The double-mean-reverting model is a three factor model which shows the consistency with empirical characters of the market variance. The double exponential Ornstein-Uhlenbeck model is a two-factor log normal stochastic volatility model in which log volatility is given by an Ornstein-Uhlenbeck process, so advantageous capturing the high frequency underlying return distribution. Also, the hybrid Heston-CIR model, which is the Heston model with the stochastic interest rates driven by the Cox- Ingersoll-Ross (CIR) process, is taken to show the impact of the stochastic interest rates to the generalized variance swaps. In this dissertation, we obtain the closed form or approximate solutions for vanilla variance swaps under the given models. Furthermore, prices of generalized variance swaps with the stochastic interest rates are written in analytic form. Throughout numerical experiments, we scrutinize the validity of our solutions by exploiting the Monte Carlo simulation, test the impact of the model parameters, and show the efficiency of the given models to the variance swaps.
Option pricing and forecasting elasticity of variance via deep neural networks
김현균 Graduate School, Yonsei University 2023 국내박사
In this dissertation, we first aim to significantly reduce the computational time for pricing vanilla and exotic options using a flow-based generative model. Flow-based generative networks learn large-scale simulated two-dimensional random states based on two stochastic volatility models. Through these networks, we can simulate option prices for a given set of parameters and achieve a fair option price as the discounted mean of the simulated prices for the stochastic volatility models. In addition, the networks provide explicit probability density functions for these stochastic volatility models, which is possible due to the flow-based generative model’s unique benefits. Finally, we compare the network-based prices with those of the Monte-Carlo simulation in terms of accuracy and time cost to demonstrate the superior performance of the proposed method. Meanwhile, volatility forecasting is important because it can be used in many different applications across the industry including risk management, derivatives trading and optimal portfolio selection. On the other hand, machine learning tends to be more accurate in making predictions when large volumes of data are involved in the system which the financial services industry tends to encounter. In this dissertation, we show that a fractional stochastic generalization of the elasticity of variance can contain latent features of the market elasticity of variance by using an artificial recurrent neural network architecture called LSTM (Long Short Term Memory) to forecast the elasticity of variance. It is shown that the forecast only with the elasticity of variance data has no statistically significant difference from forward filling, but information on the Hurst exponent can improve the power of forecasting the elasticity of variance. 본 논문은 우선 흐름기반생성 모형을 이용하여 바닐라 및 이색옵션 가격 계산에 걸리는 시간의 유의미한 절감을 목표한다. 흐름기반생성 신경망은 대규모로 시뮬레이션된 확률 변동성 모형의 이차원 확률 변수를 학습한다. 학습된 신경망을 통해서 주어진 매개변수에 대한 옵션 가격을 시뮬레이션하고 이 값들의 평균을 통해서 옵션의 공정가격을 얻는다. 또한 흐름기반생성 모형은 독자적인 특징에 의해 확률변동성 모형의 명시적 확률밀도함수를 제공한다. 마지막으로 신경망을 이용해 얻은 옵션 가격과 몬테카를로 시뮬레이션의 결과를 정확도 및 시간 비용 측면에서 비교하여 본 논문에서 제안된 방법의 우수성을 입증한다. 다음으로, 변동성 예측은 위험 관리, 파생 상품 거래, 최적 포트폴리오 선택을 포함한 산업 전반에 걸쳐 다양하게 활용될 수 있어 중요하다. 한편 기계학습은 방대한 양의 데이터가 수반될 때 더 정교한 예측을 하며 이는 앞으로 금융 산업이 맞이하게 될 상황이다. 본 논문에서는 인공순환신경망인 장단기 메모리 모형을 이용하여 분산탄성을 예측하며, 이를 통해 분산탄성의 프랙셔널 확률적 일반화가 시장 분산탄성의 잠재 특성을 내포할 수 있음을 보인다. 분산탄성 데이터만을 이용한 예측은 직전 데이터를 이용해 결측값을 채우는 방법과 통계적으로 유의미하게 다르지 않은 반면 허스트 지수에 대한 정보는 분산탄성 예측력을 향상시킬 수 있음을 보였다.
Essays on multidimensional word-of-mouth effects on product performance
김민정 Graduate School, Korea University 2021 국내박사
Since the development of Internet technology, marketers actively examine word-of-mouth (WOM) effects. Based on previous research findings (i.e. the positive effect of WOM volume and valence & the negative effect of WOM variance), this thesis focuses on online word-of-mouth (WOM) effects on product performance. First, essay 1 suggests possible interactions between WOM characteristics (volume, valence, and variance) and proposes an integrative framework to measure the effects of WOM on product performance. A panel vector autoregressive with exogeneous variables model is constructed and applied to consumer reviews and book sales data obtained from amazon.com through web crawling. Product performance elasticities with respect to WOM volume are significantly larger than WOM volume elasticities with respect to product performance, especially in the long run. Increased WOM volume have a positive effect on product performance, however WOM valence and variance has no effect on product performance. WOM valence and variance are found to negatively interact with each other. Marketers should pay extra attention to WOM volume, which shows significantly increased short- and long-run effects on product performance. WOM volume may indirectly benefit a firm by interacting with WOM valence and variance, and marketers should attend to the two characteristics, as they negatively affect WOM volume. Essay 1 contributes to the literature by incorporating dynamic interactions between three WOM characteristics and product performance. Unlike previous findings, the effects of WOM valence and variance on product performance are not sizable when considering the interactions between WOM variables. This study also suggests that WOM volume plays a crucial role in future WOM aspects, such as valence and variance. In addition, essay 2 investigates the effect of WOM volatility, a new WOM variable that measures the over-time fluctuation in the WOM volume, on product performance. High WOM volatility can be regarded similar as the advertising pulsing strategy. An unexpected large WOM volume may be effective for awareness formation that has a positive impact on product performance. However, unlike advertising, WOM may contain various types of information including negative messages. Therefore, high WOM volatility can create information overload. In other words, consumers must spend more time and efforts to interpret a large amount of information due to high WOM volatility, which may lower preference. Moreover, more fluctuating WOM volume can make the contents of WOM less credible because consumers may suspect the firm’s manipulation of the WOM to better expose a product. As a result, these two conflicting forces of WOM volatility (i.e., positive effects on awareness and negative effects on preference and credibility) leave the direct relationship between the WOM volatility and product performance as an empirical question. Also, essay 2 investigate the moderating effects of WOM volatility on the relationship between existing WOM variables (i.e., volume, valence, and variance) and product performance. To verifying WOM volatility effects, cross-sectional regression analysis was conducted, using online WOM data in Korea movie industry. The results show that WOM volatility negatively affects the product performance only when consumers rely more on WOM information. Highly volatile WOMs weaken the positive effects of WOM volume and valence and the negative effects of WOM variance. This research suggested WOM volatility as a new characteristic of WOM. This study examined WOM volatility interacts with existing WOM variables. Managerial and academic implications are discussed.
계층과 적응적인 방법에 의한 웹 이미지의 텍스트 영역 추출
Text localization in web images remains an unsolved problem. Variance method can be used to localize and distinguish texts from the background in images. However previous variance methods work as single level and they revealed a limitation in dealing with diverse size, slant, orientation, translation and color of texts. In particular, they have difficulties in locating texts of large size or texts with severe color gradation due to specific value in mask sizes. It is still a significant challenge to detect texts as a mean of the effective web searching. We present a method of robustly localizing text blocks in complex web color images using two level variance maps and adaptive approaches. Also, as the second step to identify the candidate text blocks, we define three sets of features: Wavelet, Shapes, and LBP. Especially, we propose the adaptive mask of LBP which responses flexibly to various character sizes. Our features can be evaluated efficiently. They are suitable for tackling the diverse and somewhat unpredictable shapes of text. The two-level variance method works hierarchically. The first level variance finds the approximate locations of text blocks using horizontal and vertical color variances with the specific mask sizes to ensure localization of large size texts as well as small size ones. Then it segments text components in these blocks using local thresholds, in each of which a new mask size is determined adaptively. As a second level, the automatic and non heuristic gray variance map using the new mask size is applied to each block. By the second process, backgrounds tend to disappear in each block and localization can be accurate. Highly promising experimental results have been obtained using the method in 400 web images involving characters of different sizes, angles, positions, color gradations and formats without changing the input image sizes. The candidate text block is identified as a text or non-text block by using a MLP classifier trained on adaptive sliding window in LBP, wavelet and shape feature spaces to lower the false alarm rate.
(A) study on the mechanism of variance representation in orientation perception
정진혁 Graduate School, Yonsei University 2020 국내석사
인간의 시각 체계는 다수의 시각 자극들로부터 변산과 같은 통계 정보를 표상함으로써 복잡한 정보를 효율적으로 처리한다. 본 연구에서는 변산 표상의 계산적 기제를 알아보기 위하여 사람들의 변산 지각이 여러 변산 추정치들(예: 범위, 표준편차) 중 어떤 것과 유사한지 살펴보았다. 참가자들은 여러 방위 자극들로 구성된 배열의 변산을 판단하는 과제를 수행하였다. 실험 1과 실험 2에서는 표준편차와 범위 중 한 가지를 선택적으로 조작하였을 때 변산 지각이 어떻게 달라지는지 살펴보았다. 실험 1에서는 양끝의 극단적인 방위들을 제외한 나머지 방위들의 변산을 조작하여 표준편차를 선택적으로 변화시켰다. 실험 결과, 참가자들은 범위가 유사하더라도 표준편차가 큰 배열의 변산을 더 크다고 판단했는데, 이는 변산 지각이 범위보다는 표준편차에 의존한다는 것을 보여준다. 실험 2에서는 양끝의 극단적인 방위들을 조작하여 범위를 선택적으로 변화시켰다. 그 결과, 사람들은 표준편차가 유사하더라도 양끝 방위가 나머지 방위들과 멀리 떨어져 범위가 넓은 배열이 좁은 범위에 균등하게 분포한 방위 배열보다 변산이 작다고 판단하였다. 이는 사람들이 변산을 표상할 때 나머지와 상이한 극단적인 방위를 적게 고려하여 표준편차를 계산한다는 것을 의미한다. 마지막으로 실험 3에서는 배열 내 일부 방위들의 대비를 높여 그것들이 나머지보다 더 현저하게 지각되도록 조작하였다. 현저한 방위들은 방위 분포의 평균 혹은 끝부분에 위치했는데, 조건 간 물리적인 범위나 표준편차는 동일했다. 실험 결과, 참가자들은 현저한 방위들이 평균 방위보다는 끝부분에 가까울 때 변산을 더 크다고 판단하였다. 이는 사람들이 변산을 표상할 때 현저한 대상을 나머지보다 더 많이 고려하여 표준편차를 계산한다는 것을 의미한다. 본 연구의 결과들은 사람들이 변산을 표상할 때 모든 대상을 동등하게 고려하기보다는 상황에 따라 일부 대상을 더 많이 혹은 더 적게 고려하는 가중 표준편차를 이용한다는 것을 시사한다. When there are many visual items, people can represent the variance of them accurately and rapidly. However, how the visual system computes the variance is still unclear. To investigate this, we examined which of the variability measures such as the range, standard deviation, and weighted standard deviation could account for variance perception better. Participants were asked to watch two Gabor arrays of various orientations and judge which array had more heterogeneous orientations. In Experiment 1, we manipulated orientations except those near the extreme orientations to change the standard deviation while keeping the range constant. Results showed that even when two arrays had similar ranges, the array with a larger standard deviation was perceived as more variable, indicating that people represent the variance using the standard deviation rather than the range. In Experiment 2, we manipulated the deviance of extreme orientations to change the range of orientations while the standard deviations were kept similar across conditions. We found that even when two arrays had similar standard deviations, the array of a wider range with a few extreme orientations was perceived as more variable. It indicates that people consider extreme orientations less than others when computing the standard deviation. In Experiment 3, we increased the contrast of orientations either near the mean or the extreme orientation in the array so that they were more salient than the rest. Although the actual range and standard deviation of the orientations were constant across conditions, the perceived variance was higher when salient orientations were near the extreme orientation than when they were near the mean, indicating that people consider salient orientations more than others when computing the standard deviation. In summary, these results suggest that the visual system computes the weighted standard deviation by considering some items more or less than others to represent orientation variance.
이분산 오차항을 가진 불균형 자료에 대한 분산성분 F-검정의 검정력 추적
목적 : 변량 모형에서의 분산성분 F-검정은 집단 별로 같은 표본 수와 등분산성을 필요로 하지만 모든 연구에서 이러한 가정이 만족될 수는 없다. 또한 이분산을 갖는 불균형 자료에서의 검정통계량은 F-분포를 따르지 않으며, 이 때의 검정력은 신뢰할 수 없다. 본 연구에서는 임상연구에서 일반적으로 접할 수 있는 가정 위반 상황을 설정하고 그에 따른 실제 검정력을 계산하고자 하였다. 또한 각 집단 수와 전체 관측치 수에서 F-검정에 영향을 미치는 요인들에 따른 실제 검정력의 양상을 파악하여 그 관계를 보고자 하였다. 방법 : 오차 분산과 표본 수가 집단 별로 다른 환경에서는, 명목 유의수준을 만족하는 F-검정을 수행할 수 없다. 그러므로 본 연구에서는 Davies 알고리즘을 반복 적용하여 검정의 실제 크기를 조정한 후 분산성분 F-검정의 정확한 검정력을 계산하였다. 모의실험에서는 주어진 집단 수와 전체 관측치 수하에서 일반적인 오차 분산과 불균형 자료를 가정하고 실제 검정력을 계산하였다. 또한 그 결과를 바탕으로 검정력 추적을 실시하여 F-검정에 영향을 미치는 요인들과 실제 검정력의 관계를 파악하고 검정력을 최대화하는 상황을 확인하였다. 결과 : 모든 집단 수와 전체 관측치 수하에서 실제 검정력은 가정 위반을 고려하지 않은 일반적인 F-검정의 검정력과 큰 차이가 났다. 또한 집단 수가 같고 전체 관측치 수 역시 동일한 환경에서도 자료의 불균형성과 이분산 그리고 그 조합에 따라 실제 검정력은 다양한 값을 가졌다. 한편, F-검정에 영향을 미치는 요인들에 따른 검정력 추적의 결과, 공통 급내상관계수가 제한된 각 상황에서 검정력을 최대화시키는 가장 우선적인 요인이었다. 또한 자료가 더 균형적일수록 검정력은 증가하였으며, 다른 요인들이 고정되었을 때, Box의 편향비는 감소할수록 검정력이 증가하였다 결론 : 등분산 가정을 만족하지 못하는 불균형 자료에서 일반적인 분산성분 F-검정의 검정력은 명목 유의수준을 만족하지 못하므로 신뢰할 수 없다. 이 연구에서는 Davies 알고리즘을 이용하여 정확한 검정력을 계산하였으며, F-검정에 영향을 미치는 요인들과 실제 검정력의 관계를 파악하고 검정력을 최대화하는 상황을 확인하였다. Objectives : Usual F-test for variance components in random-effects model assumes the same sample size and equal error variance for each group, but these assumptions cannot be satisfied in all studies. Under heterogeneous error variances with unbalanced data, the test statistic of the F-test does not follow an exact F-distribution and, hence, its statistical power is unreliable. In this study, we tried to set up situations of assumption violation that are often encountered in clinical studies to examine performance of the usual F-test and developed an exact F-test which ensures the correct significance level. We also examine its power of the correct F-statistic and trace its maximum power according to an unbalancedness of group sizes and heterogeneous error variances to understand their impact on the power. Methods : After adjusting critical values of the standard F-statistic to maintain a pre-specified level of significance based on the Davies algorithm, correct powers of the F-test is calculated through a simulation for different group size as well as unequal error variances. These powers are, then, beta-regression modeled with factors that are considered to affect them, which are a degree of design imbalance, a common intraclass correlation coefficient, and a Box's bias ratio of reflecting a pattern of pairing group size and error variances. Based on the fitted model, the maximum of predicted powers are identified and traced, after transforming factors to spherical coordinates in order to examine its behavior on producing the maximum power. Results : The correct power differs significantly compared to the usual power of the standard F-test when a design is unbalanced with heterogeneous error variances for a random-effect model. We find that this difference is also depends on the way of a group size and a size of error variance are being paired. We also find that these three factors are sufficient to explain the size of the exact power of the correct F-test statistic. Our power tracing reveals that the common ICC is the most affecting factor for providing a maximum power. Conclusion : When a homoscedasticity of error variances is violated in unbalanced data, the usual F-test for variance components does not satisfy nominal level of significance. Even though an exact power of an F-statistic can be obtained, factors affecting its power and how these factors affect to obtain maximum power are not well studied. Our study shows how the maximum power of an exact F-test can be traced and also provide a designing strategy to obtain its value.
Representational form of perceptual average
Kim, Myoungah Graduate School, Yonsei University 2019 국내석사
People can accurately represent ensemble properties from a set of multiple items, such as mean size. While much questions have focused on the mechanism of ensemble representation, to our knowledge there was never a discussion about the form of mean representation. The condoned assumption seems to be that mean is represented as a single average, e.g., a single size. However, some evidences contradict this intuitive understanding, one of which is that mean estimation shows large bias in studies that use single item probe to report the mean size. The fact that mean and other various ensemble statistical properties are interrelated also suggests that mean representation is more complex than a single average. The current study explored the form of mean representation by examining how mean size estimation is influenced by the characteristic differences between two comparing ensembles, specifically depending on set size and variance. In each trial, observers were presented with a set of multiple circles. They were asked to report the mean size of the standard display by adjusting the size of a single circle or the overall size of multiple circles in the probe display. We measured percentage error from the actual mean size, as well as the variance of the response. In Experiment 1, we compared mean size estimation performance between using a single probe versus a set probe. Replicating the macro trend across studies, estimation error was greater in the single probe condition than the set probe condition. In Experiment 2, we further divided probe’s set size into four levels. Results showed that error becomes systematically smaller as set size disparity decreases between standard and probe displays. In Experiment 3, we checked if this observed set-size disparity effect was possibly due to a difference in sensory memory overlap by examining whether the results change when probe is presented on a different location. Results showed no significant difference between same and different location conditions, ruling out the sensory memory explanation. Finally, Experiment 4 manipulated size variance to see how variance congruency influenced mean size estimation. Error and response variance were always smaller when variance was congruent than when variance was incongruent. All in all, error and response variance of mean estimation were contingent on the characteristics of the probe displays. This supports an idea that mean representation is not represented as a single average, but includes ensemble of statistical properties, such as variance and numerosity. 인간은 유사한 사물들에서 평균과 같은 통계 정보를 추출하여 복잡한 시각 정보를 효율적으로 표상할 수 있다. 평균 정보의 형태에 대해선 알려진 바가 적지만 대체로 평균 정보가 단일 크기와 같은 하나의 대표값으로 표상된다고 가정하는 듯하다. 하지만 단일 원을 이용하여 평균을 보고하는 대부분의 연구에서 상대적으로 큰 오차가 나타나는 점과 평균을 포함한 다양한 통계 정보들이 서로 상관되어 있다는 점을 미루어 볼 때 평균 표상이 단일 크기 보다는 더 복잡한 형태라는 것을 유추해 볼 수 있다. 본 연구는 평균 표상이 비교하는 두 자극의 특성 차이에 따라 어떻게 달라지는 지 연구하였다. 이를 위하여 자극 화면과 검사 화면 사이의 자극 개수 차이와 자극 크기들의 변산 차이를 조작하였다. 참가자들은 자극 화면에서 짧게 제시된 다양한 크기의 원들을 본 후 검사 화면에서 나타난 원(들)의 크기를 조절하여 자극 화면에 제시된 원들의 평균 크기를 추정하였다. 실험 1에서는 응답 화면의 원이 하나인 경우와 여러 개(세트)인 경우를 비교하였다. 실험 결과 단일 원 조건이 세트 조건보다 평균 추정 오차와 오차의 변산이 큰 것으로 나타났다. 실험 2 에서는 검사 화면의 원의 개수를 네 단계로 세분화 하였다. 실험 결과, 자극 화면과 검사 화면의 자극 개수 차이가 커질수록 오차가 감소하는 것으로 나타났다. 실험 3에서는 앞서 나타난 결과가 위치 중첩으로 인해 나타난 단순 감각 기억의 유사성 이었는지 확인해보기 위해 응답 화면 자극이 나타나는 위치가 같을 때와 다를 때를 비교하였다. 실험 결과, 두 조건에 차이가 없는 것으로 나타났다. 마지막으로 실험 4에서는 자극 화면과 응답 화면 간에 변산이 일치하거나 다를 때 평균 추정이 어떻게 변하는지를 알아보았다. 실험 결과, 변산이 일치할 때 오차와 오차의 변산이 작은 것으로 나타났다. 결론적으로 평균 추정 오차와 오차의 변산은 자극 화면과 응답 화면의 유사성에 따라 달라졌다. 이는 평균 정보가 하나의 대표값으로 표상되는 것이 아니고 변산과 개수와 같은 통계 정보들이 평균 표상에 유기적으로 포함되어 있다는 것을 시사한다.