Наука Просто
Cell biologyStudy analysis4 min readSeptember 16, 2026

Non-genetic protein variants expand the human proteome

A proteogenomic analysis of 29 healthy human tissues identified 7,215 unique amino-acid substitutions. Some lacked matching RNA changes, pointing to an additional source of protein diversity.

A schematic DNA-to-RNA-to-protein pathway showing one amino acid in the protein chain replaced without a corresponding change in the DNA sequence.

Illustration: Nauka Prosto, created with AI assistance.

Non-genetic protein variants can arise even when the corresponding change is absent from the DNA sequence. A study published in Nature shows that the path from gene to finished protein can itself generate additional diversity: the same genetic instruction does not always end in exactly the same amino-acid sequence.

The researchers analysed a large proteogenomic dataset spanning 29 healthy human tissues. They compared RNA information with mass-spectrometry data, which can identify the sequences of protein fragments. After stringent filtering, they reported 13,910 confidently localized protein variants representing 7,215 unique single-amino-acid substitutions.

Those substitutions do not all have the same origin. Some can be explained by ordinary inherited genetic variation or by somatic mutations. But the analysis also uncovered protein variants for which a corresponding RNA sequence change was not detected. Those were of particular interest because they point to variation generated after the genomic sequence has already been set.

A protein can change without rewriting the DNA

The familiar flow of biological information looks straightforward: DNA is transcribed into RNA, and ribosomes use that RNA sequence to assemble a protein from amino acids.

Neither step is perfectly error-free. A change can arise while RNA is being made or during translation, when the ribosome reads codons and amino acids are delivered to the growing protein chain. A small fraction of protein molecules can therefore contain a different amino acid even though the underlying genome remains unchanged.

The authors went beyond computational detection. Selected non-genetic substitutions were experimentally validated in purified proteins.

They also examined cancer-derived cell lines exposed to amino-acid starvation. Under these conditions, the substitutions did not simply appear as random background noise. Distinct and recurring patterns emerged, supporting the idea that amino-acid availability and the translation machinery can systematically influence which alternative protein forms are produced.

The origin of every detected variant cannot be established with equal certainty. Failure to detect an alternative sequence in RNA does not completely rule out a very rare genetic or RNA event that escaped sampling. The study therefore relies on multiple lines of evidence, together with direct experimental validation of selected substitutions.

Beyond what a single DNA mutation can reach

There is another consequence of this process that makes it particularly interesting.

A point mutation changes one nucleotide in DNA. A codon can therefore reach only a limited set of other codons through a single mutation. Many possible amino-acid replacements would require two or even three DNA changes.

A translation error is not constrained in the same way. The protein-synthesis machinery can occasionally insert an amino acid that would be several genetic mutations away from the original sequence.

In that sense, non-genetic variation can temporarily explore a broader region of protein “sequence space” without rewriting the genome itself.

The researchers also identified hundreds of substituted non-genetic proteoforms that either recurred across multiple healthy individuals or mapped to annotated functional sites in proteins. Finding a substitution at a functional site does not by itself establish that it is beneficial, adaptive or even functionally important. Direct functional evidence was obtained only for selected variants.

The broader result is therefore not simply a catalogue of protein-synthesis mistakes. DNA defines the basic protein sequence, but it does not completely specify every protein molecule that exists in a cell. Between genome and proteome lies another layer of variation, generated while genetic information is being turned into protein.