Bioinformatics is an important and fast-developing new field; it is associated typically with massive and unstructured data containing a massive, unprecedented volume of information. To make it really become useful for biologists, the data has to be ...
Bioinformatics is an important and fast-developing new field; it is associated typically with massive and unstructured data containing a massive, unprecedented volume of information. To make it really become useful for biologists, the data has to be processed through various computational techniques. Statistics, which provide mathematical techniques to collect, analyze and represent the data, is a perfect platform in which to develop computational tools for bioinformatics.
The main theme of this thesis is applying Statistical methods to Bioinformatics problems and it consists of three papers, (1) Zhong X, Goutsias J. Bayesian Analysis of SILAC Dynamics for Closed Biochemical Reaction Systems. (2) Zhong X, Parmigiani G. Model-based outliers identification in gene expression. (3) Zhong X, Marchionni L, Cope L, Iversen ES, Garrett-Mayer E, Gabrielson E, and Parmigiani G. Optimized cross-study analysis of microarray-based predictors.
The first paper of this thesis presents a Bayesian approach for the kinetic parameter estimation starting from Stable-Isotope Labeling by Amino Acids in Cell Culture (SILAC) experimental data on closed biochemical reactions systems. Here we use a set of nonlinear ordinary differential equations that track the effects of the simultaneously occurring reactions to represent the biological system, and we model the perturbation of the systems as well to mimic the experimental protocol to produce more informative experimental data. To learn the kinetic parameters involved in the model, the simultaneous perturbation stochastic approximation (SPSA) algorithm and the Metropolis-Hastings MCMC algorithm are applied to obtain posterior mean estimates. A numerical example is given to illustrate how the method can be used in practice, and we randomly generate synthetic data over a range of experimental conditions to evaluate the performance of the method. The second paper studies statistical methods for detecting cancer gene outliers. We compared POE with many other recently proposed outlier identification methods on synthetic data sets and a public skin cancer data set. The third paper focuses on cross-study analysis of microarray data from different platforms. We generalize the integrative correlation approach to provide a metric for evaluating the overall efficacy of preprocessing and cross-referencing, and explore optimal combinations of filtering and cross-referencing strategies.