top of page

Infectious RNA Dancing Structures: Highly Concentrated Biological Information

  • jonlieff
  • 5 days ago
  • 7 min read



RNA is a structural molecule whose three-dimensional architecture stores biological information, interacts with cellular proteins, catalyzes chemical reactions, and coordinates an infectious life cycle. It carries out these actions through precise three-dimensional structures capable of recognizing and binding other molecules with high specificity. What is less widely appreciated is that RNA achieves this wide range of interactions not despite its constant movement, but because of it — a restlessness some researchers call "breathing" or "dancing" molecules. RNA continually folds, unfolds, and refolds into different conformations, and each new shape opens new possibilities for interaction. This behavior parallels intrinsically disordered proteins (IDPs), which likewise shift among many structural states to generate the complex shapes needed for wide-ranging interactions.


The dramatic consequences of this structural dynamism are visible when comparing three small RNA-based pathogens — two viroids and one virus — each of which sustains a remarkably complex life cycle with only a small amount of genetic material.


Biological information, then, is encoded not only in nucleotide or amino acid sequence but also in dynamic structural landscapes. As RNA molecules dance, they sample alternative shapes that expose different functional surfaces. Intrinsically disordered proteins behave similarly, fluctuating among multiple conformations and becoming ordered only upon binding an appropriate partner. Together, these dynamic molecules form adaptable molecular networks capable of sensing, responding to, and coordinating complex biological processes.


Flexible RNAs and intrinsically disordered proteins frequently work in tandem within cells. Messenger RNAs, long non-coding RNAs, ribosomal RNAs, spliceosomal RNAs, and many viral genomes all undergo continual structural rearrangement. Likewise, many of the proteins that regulate transcription, RNA processing, signaling, and chromatin organization contain extensive intrinsically disordered regions. Their shared flexibility enables interactions that are simultaneously highly selective and readily reversible, allowing biological systems to rapidly assemble, reorganize, and disassemble molecular complexes as conditions change.



Three Tiny Pathogens


A comparison of three tiny RNA pathogens illustrates how fluctuating molecular shape — in RNA, and in one case a protein as well — can compress an extraordinary density of biological information into very small genomes. The progression from avocado sunblotch viroid, at roughly 250 nucleotides, to potato spindle tuber viroid, at roughly 360 nucleotides, to hepatitis delta virus, at roughly 1,700 nucleotides, traces increasing functional sophistication achieved through more elaborate structural information rather than through an expansion in genome size.


Together, the three organisms show how RNA became progressively more capable through improvements in architecture rather than growth in size. Each additional nucleotide created opportunities for new folding patterns, new sites capable of recognizing and binding other molecules, new regulatory interactions, and entirely new molecular behaviors. For comparison, HIV-1 carries roughly 9,500 nucleotides encoding nine genes in a single strand of RNA, and Ebola virus carries roughly 18,900 nucleotides encoding seven genes in a single strand of RNA — genomes many times larger than the pathogens discussed here.


Avocado viroid is among the smallest pathogens known, sitting near what may be the minimum informational threshold for an independent infectious RNA with a complex lifestyle. Despite its tiny size, it accomplishes everything infection requires without producing a single protein. Its circular RNA continually folds into dynamic three-dimensional conformations that expose or conceal different molecular surfaces. Embedded within its genome are hammerhead ribozymes that cleave the RNA precisely during replication, becoming active only when the molecule's fluctuating shape exposes the cleavage site on a nascent RNA copy. These constantly shifting structures influence copying, intracellular transport, chloroplast targeting, interactions with host proteins, and evasion of plant defenses. Nearly every nucleotide must serve multiple purposes at once, making this viroid one of the most information-dense biological molecules known.





Avocado viroid is perhaps the purest expression of an RNA-centered existence. It produces no proteins and no capsid. Every aspect of its biology depends on the geometry of its folded RNA. The molecule recruits host enzymes to copy itself, cleaves its own RNA using the hammerhead ribozyme, moves through plant tissues, persists for decades, and passes into future generations — all without ever translating a single amino acid. Its entire strategy rests on presenting the correct molecular surface to the correct host protein at precisely the right moment.


Because virtually every nucleotide participates simultaneously in several structural roles, the genome's information is extremely compressed. Regions that maintain secondary structure — the shape the RNA adopts after its first folding — also influence the ribozyme's ability to cleave, the recognition of cellular enzymes, movement within and between cells, and resistance to degradation. This degree of compression rivals the most efficient information-storage systems known in biology.


Potato viroid represents the next stage, with a somewhat larger genome of approximately 359 nucleotides. Unlike most viruses, none of this additional sequence is devoted to protein coding; instead, it folds into an exceptionally stable, rod-like structure built from extensive internal base pairing. Five distinct functional domains emerge along this rod, participating in replication, movement, engagement with the host cell, adaptation to host conditions, and the production of disease symptoms. Other regions direct movement through plant tissues, interaction with host RNA polymerase II, and systemic spread through the plant.


Like avocado viroid, potato viroid produces no proteins, relying entirely on information encoded in its RNA structure to govern its biology. But its modular architecture means individual regions are less often required to serve multiple functions simultaneously, in contrast to the highly compressed avocado viroid genome. It also requires no ribozyme to cleave the rolling-circle copy of its RNA, unlike avocado viroid. Even within this comparatively rigid rod, specific regions still fluctuate.




Even in a largely rigid structure, localized "dancing" — base pairs opening and reclosing, loops changing shape, helices flexing — proves critical to function. These transient movements let the host's RNA polymerase II and its associated factors recognize the viroid RNA by revealing interaction sites otherwise hidden within the rod. This flexibility allows binding partners to be exchanged during transport while the overall integrity of the genome is preserved, and it helps the RNA engage transport proteins and navigate the narrow channels connecting adjacent plant cells.



Hepatitis Delta Virus


Hepatitis delta virus carries approximately 1,680 nucleotides in a circular genome — almost five times larger than potato viroid and nearly seven times larger than avocado viroid. Its RNA folds into an elongated, rod-like structure stabilized by extensive base pairing while remaining remarkably dynamic. Between 70 and 75 percent of its nucleotides participate in base pairing, producing the classic rod seen in structural diagrams, much as in a viroid. Unlike a viroid's RNA, however, hep Ds genome is almost never naked: it is coated by many copies of its own protein, the hepatitis delta antigen, forming a ribonucleoprotein (RNP) complex.


Hep D’s continual shape-shifting produces long-range interactions that form and dissociate, loops that transiently open, and a ribozyme that repeatedly samples shapes before self-cleavage occurs. These structural fluctuations influence RNA copying, the editing event that produces its two antigen forms, and interactions with numerous host proteins. Rather than existing as a single fixed structure, the hep D’s genome behaves as a dynamic family of shapes that shift as the viral life cycle progresses.


Unlike the viroids, hep D does not rely solely on fluctuating RNA architecture. Its antigen is itself a versatile intrinsically disordered protein, with an architecture of its own that allows it to bind RNA, recruit host molecules, regulate replication, modulate immune responses, and coordinate the assembly of new viral particles. The emergence of this protein dramatically expands the range of behavior available to the RNA genome.

The first round of protein production from viral RNA uses 585 nucleotides — about 35 percent of the hep D genome — to produce a 195-amino-acid protein called the small hepatitis delta antigen. This small antigen promotes copying of the viral RNA during early infection. Later, a cell enzyme edits the viral RNA, converting a single nucleotide that had specified a stop codon into one specifying an amino acid. This edit extends the reading frame, producing a larger 642-nucleotide transcript that encodes a 214-amino-acid large antigen. This larger form suppresses copying while promoting the assembly of new hep D virus particles, stimulated by interaction with hepatitis B virus surface antigen proteins. In this way, one tiny genome generates two functionally distinct proteins through cell-mediated RNA editing rather than through the evolution of an entirely new gene.


Hep D's ecology is far more complex than that of the plant viroids. Viroids copy within a single plant species, whereas hep D must coordinate infection among itself, human liver cells, and hepatitis B virus. It depends on hepatitis B virus solely for envelope proteins while retaining complete control over the copying of its own genome — effectively outsourcing viral entry and particle release while keeping copying to itself. Remarkably, the same RNA sequence simultaneously stores instructions for both RNA folding and protein synthesis. This dual use of sequence allows hep D to combine the structural sophistication of a viroid with the regulatory capacities of a virus.





Hep D’s structural plasticity allows it to interact with many different partners, including its own RNA, RNA polymerase II, RNA-processing factors, chromatin-associated proteins, nuclear transport machinery, and components of the hepatitis B virus envelope. Rather than fitting one substrate with one precisely shaped active site, the antigen reshapes itself to accommodate different molecular partners as infection proceeds.

The antigen contains short linear interaction motifs, molecular recognition features, and transient secondary structures that become ordered only upon binding a specific partner. Regions involved in RNA binding and nuclear localization shift from disorder to precise, ordered conformations during these interactions. This flexibility allows the protein to function as a molecular coordinator rather than simply a catalyst, integrating numerous host pathways into the viral replication program.


Immune interactions are very complex for hep D, contending with the far more elaborate innate and adaptive immune systems of mammals. The virus persists despite interferon responses, RNA sensors, inflammatory signaling, adaptive immunity, and continual immune surveillance within the liver. Its highly base-paired RNA structure conceals vulnerable regions, while the antigen modulates numerous host pathways involved in replication, RNA processing, and antiviral defense.





The viral RNA supplies an adaptable structural scaffold capable of storing multiple layers of information within its folding landscape, while the intrinsically disordered antigen supplies a flexible protein interface capable of recognizing and organizing a wide range of host factors. Together they form a ribonucleoprotein complex far more versatile than its small genome would predict.


All three pathogens rely heavily on dancing and movement to produce precise RNA folding to position functional elements in exactly the spatial arrangement required for copying and interaction with host factors. Their genomes contain overlapping layers of condensed information, in which the same nucleotide contributes simultaneously to secondary structure, tertiary interaction, protein-binding surfaces, catalytic activity, replication signals, and — in the case of hep D — protein coding. An extraordinary amount of functional information is compressed into remarkably small genomes.

bottom of page