Herramientas estadísticas para la identificación de asociaciones entre datos genómicos y características fenotípicas en germoplasma de durazno (Prunus persica L.)
Cargando...
Fecha
Authors
Título de la revista
ISSN de la revista
Título del volumen
Editor
Universidad Nacional de Rosario
Resumen
La amplia disponibilidad de herramientas estadísticas representa una oportunidad
para analizar datos aplicando diferentes algoritmos y poder obtener mayor información
sobre un mismo conjunto de datos. Seleccionar el análisis estadístico según la pregunta
de investigación requiere conocer acerca de la metodología que las mismas aplican
e interpretar sus resultados. La generación de datos en el ámbito del mejoramiento
genético vegetal es constante y de alta dimensión. Este gran volumen de información
conlleva el desafío de poder procesarla de manera eficiente para obtener resultados
interpretables. Los modelos de asociación en estudios de todo el genoma han sido am-
pliamente estudiados con el objetivo de encontrar señales que identifiquen los genes
que producen determinadas características de interés en las plantas. Esta expresión
del fenotipo, no solo se encuentra regida por la información genética si no también
por otros factores como las condiciones donde se desarrolla el cultivo y la estructura
genética que existe en los datos. En el capítulo 1, se introducen conceptos generales que
se desarrollan en cada uno de los capítulos subsiguientes. En el Capítulo 2 se aborda el
estudio de la eficiencia de modelos GWAS para detectar verdaderas asociaciones genoti-
po–fenotipo bajo heterocedasticidad ambiental. Se realizó un estudio de simulación y
sus resultados fueron evaluados con bases de datos públicas a través de ocho modelos
que contemplaron el análisis loci a loci y multiloci. Los resultados mostraron que la he-
terocedasticidad puede enmascarar los efectos medios del genotipo. El modelo MLMM
permitió distinguir adecuadamente la señal molecular del ruido ambiental, resultando
el más parsimonioso. El Capítulo 3 aborda la aplicación de la regresión por PLS, una
técnica multivariada que modela simultáneamente marcadores genotípicos y rasgos
fenotípicos, siendo especialmente útil frente a multicolinealidad. Además de evaluar la
capacidad del modelo para detectar asociaciones relevantes con valor agronómico, se
compararon dos algoritmos para la construcción de componentes latentes: SVD y NI-
PALS. Esta comparación permitió analizar ventajas computacionales y diferencias en la
estabilidad de las soluciones, aportando evidencia sobre la elección metodológica más
adecuada según el contexto de los datos. En el Capítulo 4 se utilizó el algoritmo NIPALS
dentro del marco PLS para evaluar asociaciones entre caracteres fenotípicos y SNP con
datos faltantes. Se evaluó la performance de PLS-NIPALS en escenarios con diferentes
niveles de datos faltantes, evaluando la estabilidad del modelo mediante validación
cruzada. Los resultados indicaron que, aunque el error de predicción aumenta con la
proporción de valores faltantes, PLS mantiene capacidad de identificar asociaciones
significativas, lo que evidencia su flexibilidad para estudios GWAS aún en condiciones
de datos incompletos. Todos los códigos generados para el presente trabajo de tesis se
encuentran disponibles en bortolottoeugenia-gif, 2025 y en el Anexo de esta tesis. En
conjunto, la tesis aporta evidencia sobre la influencia de la heterogeneidad ambiental
en GWAS, la utilidad de PLS como herramienta multivariada para la asociación genó-
mica y destaca la robustez de sus algoritmos ante escenarios con información faltante,
contribuyendo así al avance de metodologías estadísticas aplicadas al mejoramiento
genético vegetal.
The wide availability of statistical tools represents an opportunity to analyze data by applying different algorithms and to obtain more information from the same dataset. Selecting the statistical analysis according to the research question requires knowledge about the methodology they apply and the interpretation of their results. Data genera- tion in the field of plant genetic improvement is constant and high-dimensional. This large volume of information entails the challenge of processing it efficiently to obtain interpretable results. Association models in genome-wide studies have been widely studied with the aim of finding signals that identify the genes responsible for certain traits of interest in plants. This expression of the phenotype is governed not only by ge- netic information but also by other factors such as the conditions under which the crop develops and the genetic structure present in the data. In Chapter 1, general concepts that are developed in each of the subsequent chapters are introduced. Chapter 2 ad- dresses the study of the efficiency of GWAS models to detect true genotype–phenotype associations under environmental heteroscedasticity. A simulation study was carried out and its results were evaluated with public databases through eight models that considered both loci-by-loci and multilocus analyses. The results showed that heteros- cedasticity can mask the average effects of the genotype. The MLMM model adequately distinguished the molecular signal from environmental noise, proving to be the most parsimonious. Chapter 3 deals with the application of PLS regression, a multivariate technique that simultaneously models genotypic markers and phenotypic traits, being especially useful in the presence of multicollinearity. In addition to evaluating the mo- del’s ability to detect relevant associations with agronomic value, two algorithms for constructing latent components were compared: SVD and NIPALS. This comparison made it possible to analyze computational advantages and differences in the stability of the solutions, providing evidence on the most appropriate methodological choice depending on the data context. In Chapter 4, the NIPALS algorithm was used within the PLS framework to evaluate associations between phenotypic traits and SNP with missing data. The performance of PLS-NIPALS was assessed in scenarios with different levels of missing data, evaluating model stability through cross-validation. The results indicated that, although prediction error increases with the proportion of missing va- lues, PLS maintains the ability to identify significant associations, demonstrating its flexibility for GWAS studies even under incomplete data conditions. All codes generated for this thesis are available in bortolottoeugenia-gif, 2025 and in the Appendix of this thesis. Overall, the thesis provides evidence on the influence of environmental hetero- geneity in GWAS, the usefulness of PLS as a multivariate tool for genomic association, and highlights the robustness of its algorithms in scenarios with missing information, thereby contributing to the advancement of statistical methodologies applied to plant genetic improvement.
The wide availability of statistical tools represents an opportunity to analyze data by applying different algorithms and to obtain more information from the same dataset. Selecting the statistical analysis according to the research question requires knowledge about the methodology they apply and the interpretation of their results. Data genera- tion in the field of plant genetic improvement is constant and high-dimensional. This large volume of information entails the challenge of processing it efficiently to obtain interpretable results. Association models in genome-wide studies have been widely studied with the aim of finding signals that identify the genes responsible for certain traits of interest in plants. This expression of the phenotype is governed not only by ge- netic information but also by other factors such as the conditions under which the crop develops and the genetic structure present in the data. In Chapter 1, general concepts that are developed in each of the subsequent chapters are introduced. Chapter 2 ad- dresses the study of the efficiency of GWAS models to detect true genotype–phenotype associations under environmental heteroscedasticity. A simulation study was carried out and its results were evaluated with public databases through eight models that considered both loci-by-loci and multilocus analyses. The results showed that heteros- cedasticity can mask the average effects of the genotype. The MLMM model adequately distinguished the molecular signal from environmental noise, proving to be the most parsimonious. Chapter 3 deals with the application of PLS regression, a multivariate technique that simultaneously models genotypic markers and phenotypic traits, being especially useful in the presence of multicollinearity. In addition to evaluating the mo- del’s ability to detect relevant associations with agronomic value, two algorithms for constructing latent components were compared: SVD and NIPALS. This comparison made it possible to analyze computational advantages and differences in the stability of the solutions, providing evidence on the most appropriate methodological choice depending on the data context. In Chapter 4, the NIPALS algorithm was used within the PLS framework to evaluate associations between phenotypic traits and SNP with missing data. The performance of PLS-NIPALS was assessed in scenarios with different levels of missing data, evaluating model stability through cross-validation. The results indicated that, although prediction error increases with the proportion of missing va- lues, PLS maintains the ability to identify significant associations, demonstrating its flexibility for GWAS studies even under incomplete data conditions. All codes generated for this thesis are available in bortolottoeugenia-gif, 2025 and in the Appendix of this thesis. Overall, the thesis provides evidence on the influence of environmental hetero- geneity in GWAS, the usefulness of PLS as a multivariate tool for genomic association, and highlights the robustness of its algorithms in scenarios with missing information, thereby contributing to the advancement of statistical methodologies applied to plant genetic improvement.
