Researchers have developed a computational method that can identify differences in DNA methylation without requiring scientists to assign samples to predefined biological groups first.

The method, called metilene3, can work with unlabeled samples and organize them according to recurring methylation patterns. The approach could help researchers explore complex datasets in which cell types, disease groups or other biological relationships are not fully known before analysis.

The peer-reviewed study was published on July 4 in Nature Communications. The findings were highlighted on August 25 by the Leibniz Institute on Aging–Fritz Lipmann Institute, one of the institutions involved in the research.

The advance is a research method rather than a new diagnostic test. The disease-related patterns identified in the study still require biological and, where relevant, clinical validation before they could be interpreted as established disease subtypes or used in patient care.

How metilene3 works without predefined labels

DNA methylation involves adding chemical methyl groups to particular DNA bases, most commonly cytosines at CpG sites in mammals. Researchers study these patterns because methylation can differ between cell types, developmental states and diseases.

One common goal is to identify differentially methylated regions, or DMRs, where DNA methylation differs between samples or biological conditions.

Many existing DMR-analysis methods require researchers to define sample groups before running the comparison. Metilene3 supports that conventional supervised approach but can also operate in an unsupervised mode when those labels are unavailable.

In unsupervised analysis, the method detects methylation differences and uses them to construct what the researchers call a Differentially Methylated Tree, or DMTree. The tree organizes samples according to methylation patterns associated with particular genomic regions.

This allows researchers not only to see which samples cluster together but also to examine the methylated regions contributing to those divisions.

Tests recovered known biological relationships

The researchers evaluated metilene3 with simulations and previously generated human datasets covering normal tissues, blood cells, glioblastoma, pancreatic tissue and cell-free DNA from cerebrospinal fluid.

In blood-cell data, the unsupervised analysis recovered relationships consistent with known immune-cell development. The method also identified methylated regions associated with regulatory factors involved in cell identity.

That result is important because it showed that the method could reconstruct biologically meaningful structure from methylation measurements without being given the expected cell-group labels beforehand.

Brain tumor analysis revealed candidate subgroups

The researchers also analyzed a dataset containing 64 glioblastoma and non-cancerous brain samples.

Metilene3 separated major tumor groups associated with IDH mutation status and divided IDH-mutated tumors into two additional clusters. The paper describes those clusters as phenotypically uncharacterized subgroups.

That wording is important. The analysis identified candidate groupings in the methylation data, but it did not establish two new clinically recognized forms of glioblastoma. Their biological and clinical significance remains uncertain.

The method also highlighted an unusual IDH-mutated tumor sample that clustered close to normal brain samples. The researchers reported additional gene-expression characteristics that remained consistent with a tumor sample, suggesting that the unusual methylation pattern was not simply explained by the sample behaving like normal brain tissue.

Pancreatic data produced a regulatory hypothesis

In pancreatic tissue data, metilene3 identified methylated regions enriched for DNA-binding motifs associated with transcription factors involved in pancreatic cancer biology, including NF-κB and NFAT.

The researchers proposed that the pattern could point to a regulatory relationship worth investigating experimentally. The study does not establish that this connection is a confirmed mechanism driving pancreatic cancer.

This distinction illustrates the intended role of the method: it can generate biologically plausible patterns and candidate relationships, but those findings still need separate experimental testing.

Benchmark results were strong, but they are not clinical accuracy figures

Metilene3 also performed strongly in computational benchmarks. In simulations, the method achieved high sensitivity and precision when researchers tested whether it could recover DMRs whose expected locations were known in advance.

The strongest numerical accuracy results therefore came from simulated benchmark data, not from a clinical diagnostic test.

The researchers also reported that a 35-sample human cancer whole-genome bisulfite sequencing dataset could be processed in about 10 minutes using 10 computing cores and less than 1 GB of memory.

Those results suggest that the method can handle sizeable methylation datasets efficiently, but they do not establish clinical performance or prove that every subgroup it identifies corresponds to a distinct biological condition.

What the method could be used for next

The main advantage of metilene3 is its ability to explore methylation data when researchers do not already know how every sample should be classified. That could be useful in studies involving mixed cell populations, heterogeneous tumors or other datasets containing previously unrecognized structure.

The authors also discuss extending the approach to additional methylation-sequencing technologies and single-cell datasets, although those applications bring their own technical challenges and require further development.

The metilene3 source code and test data are publicly available, allowing other researchers to test the method on independent datasets and assess whether candidate patterns can be reproduced.

For now, the confirmed advance is methodological: metilene3 provides researchers with a way to identify differentially methylated regions and explore sample structure without requiring predefined labels at the start of the analysis.