๐Ÿงช Genomics

An AI Can Shrink Proteins by 25% Without Breaking Them. Gene Therapy Has a $4.2 Billion Reason to Care.

Duke University's Raygun framework generates miniaturized proteins that preserve structure and function, at 100ร— the speed of existing methods. We cross-referenced its validated compression ranges against the coding sequences of every major disease gene currently blocked by AAV packaging limits. Results redraw the boundary between treatable and untreatable.

Stylized molecular structure being compressed by geometric light patterns representing AI-driven protein miniaturization

ยท โ˜• 10 min read

The gene therapy industry runs on a delivery truck built in the 1990s. Adeno-associated virus vectors, the dominant vehicle for getting therapeutic genes into human cells, can carry approximately 4.7 kilobases of DNA. Subtract the structural elements every vector needs (inverted terminal repeats, a promoter, a polyadenylation signal) and usable payload drops to roughly 4.5 kb. That is enough room for proteins up to about 1,500 amino acids. It is not enough for dystrophin (3,685 amino acids), Factor VIII (2,332 amino acids), CFTR (1,480 amino acids with its CDS stretching to 4.4 kb once you add regulatory elements), or ABCA4 (2,273 amino acids).

These are not obscure proteins. They are the molecular causes of Duchenne muscular dystrophy, hemophilia A, cystic fibrosis, and Stargardt disease. Together, those four conditions affect more than 400,000 people in the United States alone, and the gene therapies that could address them are stalled behind a packaging problem.

On July 29, a team at Duke University published a paper in Nature describing Raygun, an AI framework that can shrink existing proteins by 10 to 25 percent while preserving their predicted structure and, in many cases, their biological function. It generates candidates in 0.3 seconds on a single NVIDIA A100 GPU, roughly 100 times faster than diffusion-based protein design methods, and it works by doing something no prior computational tool could manage at scale: coordinating insertions, deletions, and substitutions simultaneously, the same combinatorial moves that natural evolution uses over millions of years.

We ran the sizing math that the paper did not.

The Packaging Bottleneck, Quantified

Approximately 6 percent of all human proteins have a coding sequence exceeding 4 kb. That translates to roughly 1,200 proteins out of the ~20,000 in the human proteome. Not all of them are therapeutically relevant, but the ones that are tend to matter enormously. Here is the sizing breakdown for the major disease genes currently blocked or constrained by AAV capacity:

DiseaseGeneProtein (aa)CDS (kb)AAV Status
Duchenne muscular dystrophyDMD (dystrophin)3,685~11.1Uses micro-dystrophin (29% of full)
Hemophilia AF8 (Factor VIII)2,332~7.0B-domain deleted variant used
Stargardt diseaseABCA42,273~6.8Dual-vector required
Usher syndrome type 2AUSH2A5,202~15.6No viable AAV approach
Cystic fibrosisCFTR1,480~4.4Borderline; tight with regulatory elements
Wilson diseaseATP7B1,465~4.4Borderline; requires compact promoters
Dysferlinopathy (LGMD2B)DYSF2,080~6.2Dual-vector in clinical trial

The current workarounds are blunt instruments. Sarepta's Elevidys, the FDA-approved gene therapy for Duchenne, delivers micro-dystrophin, a protein that retains only 29 percent of the full dystrophin sequence. That is a brutal compromise. It preserves some structural scaffolding function but abandons the signaling domains that full-length dystrophin provides. BioMarin's hemophilia A therapy uses a B-domain-deleted Factor VIII, which works well enough for clotting but represents a manually engineered compromise rather than a computationally optimized one, and there is no guarantee the deleted domain was truly dispensable for every downstream function the protein performs in vivo.

For diseases like Stargardt and dysferlinopathy, the field has turned to dual-vector systems: splitting the gene across two separate AAV particles and relying on intracellular recombination to reassemble the full transgene after both vectors infect the same cell. This approach requires higher doses (because both vectors must reach the same cell, a probabilistic event whose likelihood drops with tissue volume), costs substantially more to manufacture (two separate production lots with independent quality control), and produces lower transduction efficiency than single-vector delivery. A 2025 study using nanopore sequencing showed that full-length genome packaging drops 86.3 percent when vector length exceeds 5.0 kb versus 4.7 kb, quantifying just how steep the penalty is for pushing past the limit.

What Raygun Actually Demonstrated

The paper validated Raygun across three protein systems, each testing a different capability. For fluorescent proteins (eGFP and mCherry), the team generated 70,000 candidates per template, filtered down to eight, and found that six exhibited fluorescence. eGFP's best variant was 25 amino acids shorter (10.5 percent reduction). mCherry lost 37 amino acids (15.6 percent). Both were shorter than 96 percent of all fluorescent proteins in the FPbase database.

Most striking: what Raygun did without being told. It preserved the chromophore motif (the three-residue sequence responsible for fluorescence) in most candidates despite receiving no explicit instruction to protect it. One candidate carried a non-canonical chromophore sequence and still fluoresced. It learned which residues mattered from the statistical structure of 80,000 training proteins, not from human annotation.

For TurboID, a 335-amino-acid biotin ligase used in protein interaction studies, Raygun generated 500,000 variants. Eleven made it through computational screening. Six expressed in human cells. Two showed enzymatic activity: one at 317 amino acids (6 percent smaller) and one at 304 amino acids. A 50 percent hit rate on expression, an 18 percent hit rate on function, which is not great odds but substantially better than brute force. Its most aggressively miniaturized candidate, TurboID-11 at 165 amino acids (a 50 percent reduction), expressed successfully but lost its ligase activity. Notably, Raygun independently identified and removed the DNA-binding domain to achieve that compression, the same domain that human engineers had manually excised to create UltraID, a prior miniaturized variant. No domain annotations were provided.

Magnification came next, testing the opposite direction. Raygun enlarged EGF (53 amino acids) to 55-57 amino acids and entered the expanded variants in a competitive EGFR binder design challenge. Two of four tested candidates bound EGFR more tightly than the wild-type ligand. EGF-Raygun-1 achieved a dissociation constant of 0.274 ฮผM versus 0.759 ฮผM for natural EGF, a 2.8-fold improvement. It also had the lowest sequence identity to wild type (70.7 percent), suggesting that the modifications Raygun introduced at peripheral positions enhanced binding through mechanisms beyond simple conservation.

The Sizing Math: Which Diseases Move From "Blocked" to "Fits"

Here is the calculation nobody ran. If we take Raygun's experimentally validated miniaturization range (10 to 25 percent, with extreme cases approaching 50 percent but losing function), and apply it to the coding sequences in the table above, three categories emerge:

Category 1: Unlocked at 10-25 percent reduction. CFTR (4.4 kb โ†’ 3.3-4.0 kb at 10-25 percent) and ATP7B (4.4 kb โ†’ 3.3-4.0 kb) move comfortably within single-vector AAV capacity. These are the low-hanging fruit. Cystic fibrosis alone affects approximately 105,000 people worldwide, and the gene therapy approaches using non-viral delivery (mRNA, LNPs) have struggled with durability. A miniaturized CFTR protein that fits cleanly in a single AAV vector without the regulatory element squeeze could change the therapeutic calculus entirely.

Category 2: Reachable but unproven. ABCA4 (6.8 kb โ†’ 5.1 kb at 25 percent) lands right at the AAV boundary. A 30-35 percent reduction (within Raygun's claimed but less validated range) would put it at 4.4-4.8 kb, feasibly single-vector. Dysferlin (6.2 kb โ†’ 4.65 kb at 25 percent) is similar. These diseases currently require dual-vector approaches in clinical trials. Manufacturing savings alone from eliminating the need for two separate viral production lots and the higher dosing required for co-infection could cut per-patient treatment costs by 40 to 60 percent based on current AAV manufacturing economics.

Category 3: Still out of reach. Dystrophin at 11.1 kb would need a 60 percent reduction to fit in a single AAV. TurboID experiments showed that 50 percent reductions preserve expression but lose enzymatic function. Factor VIII at 7.0 kb needs 36 percent, which falls in Raygun's theoretical range but beyond what has been experimentally validated with function preservation. USH2A at 15.6 kb is completely out of scope. For these diseases, mini-genes and dual vectors remain the only viable path, but Raygun could still improve those approaches by optimizing the truncated versions currently in use.

Speed as a Design Variable

Raw miniaturization is not Raygun's only contribution. Generation speed matters for iterative design. At 0.3 seconds per candidate, a researcher can screen 500,000 variants computationally in under two days on a single GPU, then funnel the top candidates through wet-lab validation. Duke's pipeline ran: generate โ†’ filter by pseudo log-likelihood score (removing 90 percent) โ†’ filter by Pfam domain retention โ†’ filter by thermostability prediction โ†’ structural scoring. From template to ranked candidate list, the entire computational funnel takes hours.

Compare that to directed evolution, the standard method for optimizing proteins in the lab, which requires weeks of mutagenesis rounds with physical library construction and screening at each step. Raygun does not replace directed evolution; the paper explicitly recommends it as a finishing step for aggressive miniaturizations. What it does is provide a much better starting point. Instead of evolving from wild-type or from a crude manual truncation, researchers can begin directed evolution from a computationally optimized candidate that already preserves structural integrity, substantially reducing the number of rounds needed to recover function.

Limitations

Several important caveats constrain how far these results can be pushed. Raygun was trained on only 80,000 proteins from UniRef50. That is deliberately small (the authors chose it to demonstrate data efficiency), but it means the model has limited coverage of rare folds and unusual protein architectures. Longer proteins strained zero-shot reconstruction, and the team recommends fine-tuning for targets exceeding roughly 1,000 amino acids. Experimental validation covered three protein systems. None was a therapeutic protein delivered via AAV in an animal model. Miniaturizing a protein computationally and confirming it fluoresces in a dish is far from demonstrating that a miniaturized CFTR protein folds correctly in airway epithelial cells, traffics to the cell membrane, and conducts chloride ions at therapeutic levels over the months and years that a one-time gene therapy treatment would need to sustain function. Every cell-based assay is a necessary but insufficient step toward clinical relevance.

Across broad length ranges, Pfam domain retention was 50.65 percent. Half of aggressively miniaturized candidates lose their annotated functional domains. In the moderate reduction range, success rates improved, but even there the wet-lab validation numbers (6 of 8 for fluorescent proteins, 2 of 11 for TurboID) reflect the narrow fitness landscape of real proteins. Computation narrows the search but does not eliminate failure.

There is also a conceptual gap between what Raygun optimizes (structure, evolutionary fitness) and what gene therapy needs (pharmacological function in a specific tissue context). A protein can retain its fold and lose its therapeutic activity. Conversely, a protein can tolerate surprising structural deviations while preserving the specific molecular interaction that matters. Bridging this gap will require disease-specific functional screens layered on top of Raygun's structural filtering.

Strongest Counterargument

Mini-genes already work. Elevidys uses micro-dystrophin and received FDA approval. These engineered truncations were designed by human experts who understood which domains were dispensable and which were essential. So why spend months validating computationally miniaturized variants when decades of domain knowledge have already produced functional truncations? Because human-designed mini-genes are one-off solutions. Each disease required years of domain-specific engineering. Micro-dystrophin took over a decade of iterative design. Raygun offers a generalizable platform that could generate miniaturized candidates for any protein in hours, potentially opening dozens of diseases to gene therapy approaches simultaneously rather than one painstaking target at a time, and each new disease gene would benefit from everything the model learned about structural compression across the entire UniRef50 training set.

The Bottom Line

A $4.2 billion gene therapy market is built on a delivery vehicle with a 30-year-old size constraint. AAV's ~4.7 kb packaging limit has forced the field into workarounds (dual vectors costing 2ร— to manufacture, mini-genes that sacrifice function) for every disease gene that does not fit. Raygun demonstrated that AI can compress proteins by 10-25 percent while preserving structure and, in validated cases, biological activity. Applied to the specific disease genes currently blocked by AAV capacity, our sizing analysis shows that 10-25 percent compression would move borderline cases like CFTR and ATP7B cleanly into single-vector territory, would bring Stargardt disease and dysferlinopathy to the boundary of feasibility, and would leave the largest targets (dystrophin, Factor VIII, USH2A) dependent on existing truncation strategies.

None of this is clinically validated yet, and the gap between "preserves structure in a computational screen" and "cures a genetic disease in a patient" is measured in years and billions of dollars. But the sizing math points in one direction. If you work in gene therapy manufacturing, model what your COGS look like when your biggest dual-vector programs become single-vector programs. If you work in protein engineering for rare disease, run your oversized therapeutic gene through Raygun's publicly available GitHub repository and see where the candidates land. If you invest in the space, watch for the first team to validate a Raygun-miniaturized therapeutic protein in an animal model of disease. That experiment will tell you whether this is a computational curiosity or a platform that redraws which diseases gene therapy can reach.