Can DNA sequence predict DNA methylation?
A deep-learning framework called Melody learned to predict methylation patterns from 10-kb DNA sequence windows across 39 normal human cell types, while a transcriptome-informed version extended the approach to unseen cell types.

Illustration: Nauka Prosto, created with AI assistance.
DNA methylation presents an interesting paradox. Different cell types carry essentially the same DNA sequence, yet their epigenetic profiles can be strikingly different. A new deep-learning framework called Melody suggests that it is nevertheless possible to predict DNA methylation at many genomic locations from sequence alone — particularly when the model examines about 10,000 bases of surrounding DNA rather than a few dozen.
DNA methylation is a chemical modification in which a small methyl group is added to cytosine. In humans, it is commonly studied at CpG sites, where cytosine is followed by guanine. Methylation patterns are associated with gene regulation, chromatin organization and cellular identity.
Yet methylation is not simply an additional regulatory layer disconnected from the underlying genome. DNA contains sequence motifs recognized by regulatory proteins, and the presence and arrangement of these motifs are associated with local methylation states. The key question is how much of the methylome can be inferred from sequence itself.
Looking across 10,000 bases
The researchers used previously published methylation maps from 39 normal human cell types to train Melody. The model receives a 10-kilobase DNA sequence window and predicts methylation levels at individual CpG sites.
That is a substantially broader genomic context than some earlier methylation-prediction systems. One of the comparison models, for example, used only 41-base-pair windows. Such a short sequence may capture a nearby regulatory motif while missing relevant signals located farther away.
The difference was measurable. Across 39 datasets, the multi-track version of Melody achieved mean Spearman correlations of 0.723 on a sampled test set and 0.645 on a held-out chromosome. The strongest competing baselines reached 0.590 and 0.584, respectively. The authors report an average relative improvement of about 17% in Spearman correlation over competing methods.
A correlation of 0.723 should not be interpreted as “72.3% accuracy.” It measures how closely the ordering of predicted methylation values follows the observed pattern across genomic positions.
Tests of sequence-window length were also informative. Windows of roughly 10 kb performed better than 2-kb windows, whereas extending the input to 20–50 kb added redundant sequence without improving performance.
What the model learned about methylation regulation
The researchers also asked whether Melody could predict the effects of genetic variants associated with changes in DNA methylation, known as methylation quantitative trait loci, or meQTLs.
The model evaluated how single DNA variants could create or disrupt sequence motifs recognized by regulatory proteins. The team also performed a systematic analysis of 282 transcription-factor motifs. Melody recovered both broadly acting regulatory patterns and motifs associated with particular cellular lineages.
This is potentially informative, but the distinction between prediction and mechanism matters. A motif that changes a model's methylation prediction is not automatically a motif proven to cause that methylation change inside a living cell. The authors explicitly describe these interpretation analyses as model-based associations rather than direct evidence of causal regulatory mechanisms in vivo.
Sequence is not the whole story
The limits of sequence-only prediction became especially clear when the researchers tested cell types that Melody had never encountered during training. Five cell types were completely held out, while the model was trained on the remaining 34.
For this task, the researchers developed Melody-G, which combines DNA sequence with information about cellular state derived from single-cell RNA-sequencing data. Across the five unseen cell types, the mean AUC increased from 0.633 for a simple mean-profile baseline to 0.697 after the second stage of Melody-G training, a 10.1% relative improvement.
The result illustrates the central biological point. DNA sequence contains a substantial amount of information about where methylation patterns are likely to occur, but it does not contain everything. Actual methylation states are also shaped by chromatin accessibility, transcription-factor occupancy, developmental history, replication timing and other aspects of cellular context.
The study is entirely computational. No new biological samples were collected, and the analyses relied on existing public datasets. Melody therefore should not be viewed as a replacement for experimental methylation measurements or as a clinically validated diagnostic technology.
The broader implication is more fundamental. Epigenetic information is not encoded in DNA as a simple deterministic second code, but neither is it independent of sequence. A substantial part of methylome organization appears predictable from the genome itself, while additional information about cellular state helps explain what sequence alone cannot.
© 2026 Nauka Prosto. Rights holder: David Cheishvili. Brief quotations are permitted with an active link to the original article. Copyright rules
