A nonparametric Bayesian approach for clustering bisulfate-based DNA methylation profiles.

Lin Zhang; Jia Meng; Hui Liu; Yufei Huang

doi:10.1186/1471-2164-13-s6-s20

A nonparametric Bayesian approach for clustering bisulfate-based DNA methylation profiles.

Lin Zhang^*, Jia Meng, Hui Liu, Yufei Huang

^*Corresponding author for this work

Research output: Contribution to journal › Article › peer-review

9 Citations (Scopus)

Abstract

DNA methylation occurs in the context of a CpG dinucleotide. It is an important epigenetic modification, which can be inherited through cell division. The two major types of methylation include hypomethylation and hypermethylation. Unique methylation patterns have been shown to exist in diseases including various types of cancer. DNA methylation analysis promises to become a powerful tool in cancer diagnosis, treatment and prognostication. Large-scale methylation arrays are now available for studying methylation genome-wide. The Illumina methylation platform simultaneously measures cytosine methylation at more than 1500 CpG sites associated with over 800 cancer-related genes. Cluster analysis is often used to identify DNA methylation subgroups for prognosis and diagnosis. However, due to the unique non-Gaussian characteristics, traditional clustering methods may not be appropriate for DNA and methylation data, and the determination of optimal cluster number is still problematic. A Dirichlet process beta mixture model (DPBMM) is proposed that models the DNA methylation expressions as an infinite number of beta mixture distribution. The model allows automatic learning of the relevant parameters such as the cluster mixing proportion, the parameters of beta distribution for each cluster, and especially the number of potential clusters. Since the model is high dimensional and analytically intractable, we proposed a Gibbs sampling "no-gaps" solution for computing the posterior distributions, hence the estimates of the parameters. The proposed algorithm was tested on simulated data as well as methylation data from 55 Glioblastoma multiform (GBM) brain tissue samples. To reduce the computational burden due to the high data dimensionality, a dimension reduction method is adopted. The two GBM clusters yielded by DPBMM are based on data of different number of loci (P-value < 0.1), while hierarchical clustering cannot yield statistically significant clusters.

Original language	English
Journal	BMC Genomics
Volume	13 Suppl 6
DOIs	https://doi.org/10.1186/1471-2164-13-s6-s20
Publication status	Published - 2012
Externally published	Yes

Access to Document

10.1186/1471-2164-13-s6-s20

Cite this

@article{1d48984563434958873afb64bfd97228,

title = "A nonparametric Bayesian approach for clustering bisulfate-based DNA methylation profiles.",

abstract = "DNA methylation occurs in the context of a CpG dinucleotide. It is an important epigenetic modification, which can be inherited through cell division. The two major types of methylation include hypomethylation and hypermethylation. Unique methylation patterns have been shown to exist in diseases including various types of cancer. DNA methylation analysis promises to become a powerful tool in cancer diagnosis, treatment and prognostication. Large-scale methylation arrays are now available for studying methylation genome-wide. The Illumina methylation platform simultaneously measures cytosine methylation at more than 1500 CpG sites associated with over 800 cancer-related genes. Cluster analysis is often used to identify DNA methylation subgroups for prognosis and diagnosis. However, due to the unique non-Gaussian characteristics, traditional clustering methods may not be appropriate for DNA and methylation data, and the determination of optimal cluster number is still problematic. A Dirichlet process beta mixture model (DPBMM) is proposed that models the DNA methylation expressions as an infinite number of beta mixture distribution. The model allows automatic learning of the relevant parameters such as the cluster mixing proportion, the parameters of beta distribution for each cluster, and especially the number of potential clusters. Since the model is high dimensional and analytically intractable, we proposed a Gibbs sampling {"}no-gaps{"} solution for computing the posterior distributions, hence the estimates of the parameters. The proposed algorithm was tested on simulated data as well as methylation data from 55 Glioblastoma multiform (GBM) brain tissue samples. To reduce the computational burden due to the high data dimensionality, a dimension reduction method is adopted. The two GBM clusters yielded by DPBMM are based on data of different number of loci (P-value < 0.1), while hierarchical clustering cannot yield statistically significant clusters.",

author = "Lin Zhang and Jia Meng and Hui Liu and Yufei Huang",

note = "Funding Information: Based on “Clustering DNA methylation expressions using nonparametric beta mixture model”, by Lin Zhang, Jia Meng, Hui Liu and Yufei Huang which appeared in Genomic Signal Processing and Statistics (GENSIPS), 2011 IEEE International Workshop on. {\textcopyright} 2011 IEEE [27]. The work of L. Zhang is supported by “the Fundamental Research Funds for the Central Universities” (2010QNA50). The work of H. Liu is supported by “the Fundamental Research Funds for the Central Universities” (2010QNA47). The work of Y. Huang is supported by Qatar National Research Fund (09-874-3-235). This article has been published as part of BMC Genomics Volume 13 Supplement 6, 2012: Selected articles from the IEEE International Workshop on Genomic Signal Processing and Statistics (GENSIPS) 2011. The full contents of the supplement are available online at http://www. biomedcentral.com/bmcgenomics/supplements/13/S6.",

year = "2012",

doi = "10.1186/1471-2164-13-s6-s20",

language = "English",

volume = "13 Suppl 6",

journal = "BMC Genomics",

issn = "1471-2164",

}

TY - JOUR

T1 - A nonparametric Bayesian approach for clustering bisulfate-based DNA methylation profiles.

AU - Zhang, Lin

AU - Meng, Jia

AU - Liu, Hui

AU - Huang, Yufei

N1 - Funding Information: Based on “Clustering DNA methylation expressions using nonparametric beta mixture model”, by Lin Zhang, Jia Meng, Hui Liu and Yufei Huang which appeared in Genomic Signal Processing and Statistics (GENSIPS), 2011 IEEE International Workshop on. © 2011 IEEE [27]. The work of L. Zhang is supported by “the Fundamental Research Funds for the Central Universities” (2010QNA50). The work of H. Liu is supported by “the Fundamental Research Funds for the Central Universities” (2010QNA47). The work of Y. Huang is supported by Qatar National Research Fund (09-874-3-235). This article has been published as part of BMC Genomics Volume 13 Supplement 6, 2012: Selected articles from the IEEE International Workshop on Genomic Signal Processing and Statistics (GENSIPS) 2011. The full contents of the supplement are available online at http://www. biomedcentral.com/bmcgenomics/supplements/13/S6.

PY - 2012

Y1 - 2012

N2 - DNA methylation occurs in the context of a CpG dinucleotide. It is an important epigenetic modification, which can be inherited through cell division. The two major types of methylation include hypomethylation and hypermethylation. Unique methylation patterns have been shown to exist in diseases including various types of cancer. DNA methylation analysis promises to become a powerful tool in cancer diagnosis, treatment and prognostication. Large-scale methylation arrays are now available for studying methylation genome-wide. The Illumina methylation platform simultaneously measures cytosine methylation at more than 1500 CpG sites associated with over 800 cancer-related genes. Cluster analysis is often used to identify DNA methylation subgroups for prognosis and diagnosis. However, due to the unique non-Gaussian characteristics, traditional clustering methods may not be appropriate for DNA and methylation data, and the determination of optimal cluster number is still problematic. A Dirichlet process beta mixture model (DPBMM) is proposed that models the DNA methylation expressions as an infinite number of beta mixture distribution. The model allows automatic learning of the relevant parameters such as the cluster mixing proportion, the parameters of beta distribution for each cluster, and especially the number of potential clusters. Since the model is high dimensional and analytically intractable, we proposed a Gibbs sampling "no-gaps" solution for computing the posterior distributions, hence the estimates of the parameters. The proposed algorithm was tested on simulated data as well as methylation data from 55 Glioblastoma multiform (GBM) brain tissue samples. To reduce the computational burden due to the high data dimensionality, a dimension reduction method is adopted. The two GBM clusters yielded by DPBMM are based on data of different number of loci (P-value < 0.1), while hierarchical clustering cannot yield statistically significant clusters.

AB - DNA methylation occurs in the context of a CpG dinucleotide. It is an important epigenetic modification, which can be inherited through cell division. The two major types of methylation include hypomethylation and hypermethylation. Unique methylation patterns have been shown to exist in diseases including various types of cancer. DNA methylation analysis promises to become a powerful tool in cancer diagnosis, treatment and prognostication. Large-scale methylation arrays are now available for studying methylation genome-wide. The Illumina methylation platform simultaneously measures cytosine methylation at more than 1500 CpG sites associated with over 800 cancer-related genes. Cluster analysis is often used to identify DNA methylation subgroups for prognosis and diagnosis. However, due to the unique non-Gaussian characteristics, traditional clustering methods may not be appropriate for DNA and methylation data, and the determination of optimal cluster number is still problematic. A Dirichlet process beta mixture model (DPBMM) is proposed that models the DNA methylation expressions as an infinite number of beta mixture distribution. The model allows automatic learning of the relevant parameters such as the cluster mixing proportion, the parameters of beta distribution for each cluster, and especially the number of potential clusters. Since the model is high dimensional and analytically intractable, we proposed a Gibbs sampling "no-gaps" solution for computing the posterior distributions, hence the estimates of the parameters. The proposed algorithm was tested on simulated data as well as methylation data from 55 Glioblastoma multiform (GBM) brain tissue samples. To reduce the computational burden due to the high data dimensionality, a dimension reduction method is adopted. The two GBM clusters yielded by DPBMM are based on data of different number of loci (P-value < 0.1), while hierarchical clustering cannot yield statistically significant clusters.

UR - http://www.scopus.com/inward/record.url?scp=84876088825&partnerID=8YFLogxK

U2 - 10.1186/1471-2164-13-s6-s20

DO - 10.1186/1471-2164-13-s6-s20

M3 - Article

C2 - 23134689

AN - SCOPUS:84876088825

SN - 1471-2164

VL - 13 Suppl 6

JO - BMC Genomics

JF - BMC Genomics

ER -

A nonparametric Bayesian approach for clustering bisulfate-based DNA methylation profiles.

Abstract

Access to Document

Other files and links

Cite this