Skip to content
Live newsroom 38 readers online
Saturday, August 29, 2026 Live Sync: Just now
BreakingU.S. Gains the Upper Hand in the Battle for Hormuz
Share Suggestions AVOID F Stage 4 (Conv: 3/5 | Size: 10%)

Precise genomic integration of large DNA fragments by donor-directed annealing using prime editing

Nature Biotechnology (2026) Cite this article Replacing large-scale fragments in human cells remains a substantial challenge. Here, we present a programmable gene replacement tool, named prime assembly (PA), which adapts prime editors to produce one or two pairs of 3′-flaps on both the genome and donor DNA. These 3′-flaps anneal to each other precisely, similar […]

By deepak · August 28, 2026 · 13 min read

Nature Biotechnology
(2026) Cite this article

Replacing large-scale fragments in human cells remains a substantial challenge. Here, we present a programmable gene replacement tool, named prime assembly (PA), which adapts prime editors to produce one or two pairs of 3′-flaps on both the genome and donor DNA. These 3′-flaps anneal to each other precisely, similar to Gibson assembly in DNA oligonucleotides, allowing megabase-scale genomic excision and/or kilobase-scale donor insertion at the gene of interest. PA accepts DNA plasmids and linear double-stranded DNA as donors, ranging from 1.0 to 6.5 kb in size. We demonstrate an efficiency of up to 57.8% in replacing endogenous sequences with a 2.9-kb donor DNA fragment in HEK293T cells, with an accuracy of >90% for integrated PA fragments. Furthermore, PA enables site-specific chimeric antigen receptor integration with up to 28.1% efficiency in primary human T cells. When PA containing a GFP donor is delivered to mice by hydrodynamic injection, an average integration efficiency of 4.3% is measured in GFP-positive hepatocytes.

Genetic disorders arise from a wide spectrum of pathogenic mutations scattered across large genomic regions1. CRISPR-associated tools enable the direct correction of endogenous gene mutations and have created novel opportunities in therapeutics2,3,4,5,6,7,8,9,10,11. Indeed, point mutations can be corrected by base editors (BEs)2,3,12, meanwhile, prime editing (PE) tools can resolve insertion and deletion (indel) mutations as small as several tens of base pairs6. Precise correction or replacement of large DNA sequences (>kb) in human cells requires technically challenging strategies, which would be useful for therapeutic applications such as restoring short tandem repeats (STRs) and site-specific integration of normal genes. Currently, CRISPR nuclease-based gene correction methods, including homology-directed repair (HDR)10,11,13,14 and homology-independent targeted integration15, or PE-assisted approaches such as primed microhomolog-assisted integration (PAINT)16 are available for large-scale DNA insertions. However, as these methods rely on generating double-strand breaks (DSBs), indels are also subsequently produced at the target site. Furthermore, DSBs often lead to unwanted large DNA deletions17,18, chromosomal aberrations19 and cell death20, limiting further applications. Thus, several methods have been developed to realize large-scale DNA insertions or deletions without generating DSBs21. For example, PE-combined recombinase-based tools, such as PE-assisted site-specific integrase gene editing (PASSIGE)4 and programmable addition through site-specific targeting elementss5, have been developed; however, these techniques require multiple steps involving the installation of a landing pad and the integration of donor DNA at the target site. Additionally, CRISPR-associated transposon7, bridge RNA-guided recombination22 and RNA retrotransposon-based systems23,24 have been used for large-scale DNA insertion in prokaryotes and eukaryotes; however, the overall editing efficiencies were relatively low in human cells. Moreover, no system has yet enabled precise gene replacement, that is, the simultaneous large-scale deletion and insertion of genes in the genome.

To address these limitations, we developed a large-scale DNA replacement tool, named prime assembly (PA), which adapts PE tools to produce 3′-flaps of single-stranded DNA (ssDNA) on both genome and donor DNA. These 3′-flaps anneal to each other precisely, similar to click chemistry in small molecules or Gibson assembly in DNA oligonucleotides, and initiate strand exchange25,26, allowing efficient megabase-scale genomic excision and/or multikilobase donor insertion at the gene of interest in human cells. We achieved a replacement efficiency of 57.8% at endogenous loci with multikilobase payloads using the PA system, highlighting the potential of this system for therapeutic genome editing. PA was also effective in primary human T cells, enabling site-specific chimeric antigen receptor (CAR) integration with up to 28.1% efficiency and functional engineering.

Previously, PE was used to create a 3′-flap with several nucleotides at the genome locus for precise duplication of megabase-scale chromosomal DNA in human cells27. Subsequently, we hypothesized that creating a 3′-flap at the target genome and a complementary 3′-flap at the donor DNA could induce robust hybridization, ultimately promoting large DNA insertion or replacement. Thus, we proposed three different strategies. The first was to generate a single 3′-flap (named A) at the target site and a complementary 3′-flap (named A′) in donor DNA using respective PE guide RNA (pegRNA) to encode a primer-binding site and a reverse transcriptase template (RTT) for each 3′-flap. We assumed that, after hybridization between the genomic flap (A) and complementary donor flap (A′), the donor-templated DNA extension from the A–A′ hybridization site would be incorporated into the genome through homology-dependent strand invasion. This strategy was named single-flap PA (SF-PA) (Fig. 1a). Secondly, most processes are similar to the generation of the SF-PA; however, to enhance the invasion of donor homology sequences, a further nick was created by additional nicking guide RNA (ngRNA) on the opposite strand near the homologous sequence in the genome. This strategy was named SF-PA with a nick (SFn-PA) (Fig. 1b). Lastly, four corresponding pegRNAs were used to generate two different 3′-flaps (named A and B) at two target sites and two complementary 3′-flaps (named A′ and B′) in the donor DNA. For this third strategy, we assumed that, after hybridization between genomic flaps and complementary donor flaps (that is, A–A′ and B–B′), the DNA extensions from each hybridization site would pair with the donor template and be incorporated into the genome, thereby replacing the intervening genomic sequences. This strategy was named dual-flap PA (DF-PA) (Fig. 1c).

a–c, Schematics of SF-PA (a), SFn-PA (b) and DF-PA (c). Complementary 3′-flaps, generated by pegRNAs, tether the genome and donor to enable donor-templated synthesis and sequence replacement. An additional nick in SFn-PA enhances strand invasion, while paired flaps in DF-PA mediate excision and insertion of large fragments. RT, reverse transcriptase. d, Schematic of PA function in HEK293T cells with CMV promoter knock-in at the AAVS1 locus using promoter-less copGFP-containing donors. Left, SF-PA and SFn-PA. Right, DF-PA. e, Flow cytometry analysis of integration efficiency of SF-PA and SFn-PA according to flap-to-homology distance (0–11,000 bp), with the corresponding SF donor control included for each distance. f, Flow cytometry analysis of integration efficiency of DF-PA according to flap-to-flap distance (100–10,000 bp), compared to the DF donor control.

To examine these three PA strategies in human cells, we established a cytomegalovirus (CMV) promoter knock-in HEK293T cell line, in which the CMV promoter and blasticidin S deaminase were located in the AAVS1 safe-harbor region (Extended Data Fig. 1). Then, a donor plasmid encoding copGFP (green fluorescent protein from the copepod Pontellina plumata) without a promoter was prepared. For the SF-PA and SFn-PA donor plasmids, one flap could be made upstream of copGFP, while the homologous sequences were located downstream of copGFP. By contrast, two different flaps could be made, both upstream and downstream of copGFP for the DF-PA donor plasmid (Fig. 1d). When the copGFP cassette (4.3 kb for SF-PA and SFn-PA; 2.9 kb for DF-PA) was precisely integrated downstream of the CMV promoter region using the PA system, flow cytometry could be used to detect green fluorescence signals. Initially, the distance between the genomic flap and 1-kb homology arm was designed to be 100 bp for SF-PA and SFn-PA, while the distance between two genomic flaps in the genome for DF-PA was intended to be 100 bp. In this experiment, all flap sizes were 30 nt long. Notably, the PA efficiencies for SF-PA, SFn-PA and DF-PA using PE7 (ref. 28) were very high in the flow cytometry results at 44.4%, 54.2% and 70.3%, respectively (Fig. 1e,f). Moreover, PA events were observed in all cases after increasing the flap-to-homology distance on the genome to approximately 11 kb for SF-PA and SFn-PA and the flap-to-flap distance to approximately 10 kb for DF-PA (Fig. 1e,f), which strongly supports the effectiveness of all three PA strategies as gene replacement tools in human cells.

Despite PA tools possessing high gene replacement efficiencies, we suspected that the frequencies of PA events were potentially overestimated in the flow cytometry results, as cells with a replacement in only one allele exhibited an estimated frequency of 100%. Thus, we conducted digital PCR (dPCR) experiments to quantify the genomic DNA level of PA efficiency (Supplementary Table 1). For DF-PA, we revisited the copGFP replacement experiments in the CMV knock-in HEK293T cells, while the PA events were also measured by dPCR, regardless of the flap-to-flap distance (Fig. 2a). The dPCR results were highly correlated with the flow cytometry results (R = 0.96); meanwhile, the overall efficiencies by dPCR were expectedly lower than those by flow cytometry (Fig. 2a).

a, Correlation between DF-PA integration efficiency measured by flow cytometry and dPCR across varying flap-to-flap distances, with a correlation coefficient (R) of 0.96. b, Schematic illustrating the optimized elements within donors for each PA type. SF-PA and SFn-PA were optimized for homology length (bp), flap length (nt) and flap G+C content (%), while DF-PA was optimized for flap length (nt) and flap G+C content (%). All donor types, containing promoter-less copGFP sequences, were compared across various payload sizes. c, Integration efficiencies of SF-PA, SFn-PA and DF-PA were measured with flap lengths ranging from 0 to 50 nt. d, Integration efficiencies of SF-PA, SFn-PA and DF-PA were measured using flap G+C contents ranging from 20% to 80%. e, Integration efficiencies of SF-PA and SFn-PA with varying homology arm lengths and flap-to-homology distances, measured by flow cytometry. f, Integration efficiencies of SF-PA and SFn-PA at the HPRT1, AAVS1 and VEGFA loci with donor payloads of 1.0–5.1 kb. The flap-to-homology distance was set to approximately 2,000 bp for all targets. g, Integration efficiencies of DF-PA at the HPRT1, AAVS1, VEGFA and HEK4 loci with donor payloads of 1.0–6.5 kb. The flap-to-flap distance was set to 100–200 bp on the basis of the best-performing pegRNA combinations for each target. h, Integration efficiencies of SF-PA, SFn-PA and DF-PA at VEGFA and HEK4 loci in multiple human cell lines (HEK293, HeLa and K562). Values in the heat map and bar graphs represent means (n = 3 independent biological replicates); error bars represent the s.d. Unless otherwise indicated, experiments were performed in HEK293T cells.

Next, we investigated the optimal conditions of the PA tools using flap lengths from 0 to 50 nucleotides (nt) and G+C contents from 20% to 80% (Fig. 2b, Extended Data Fig. 2a and Supplementary Table 2). Flap lengths of 30–50 nt and GC contents of 60–80% promoted efficient replacements with similar trends observed for SF-PA, SFn-PA and DF-PA (Fig. 2c,d). Thus, we selected a standard flap length condition of 40 nt and G+C content of 60%. We also examined the optimal homologous arm length (100 bp to 2,000 bp) for SF-PA and SFn-PA, with various flap-to-homology distances. Here, the PA efficiency was influenced by both the arm length and the flap-to-homology distance, while the highest efficiency was predominantly achieved using arm lengths of 1,000–2,000 bp (Fig. 2e). Therefore, a standard homology arm length condition of 1,000 bp was chosen for SF-PA and SFn-PA. Lastly, we examined the capacity of payload size for all PA systems at multiple endogenous sites. For SF-PA and SFn-PA with flap-to-homology distances of approximately 2,000 bp, we tested payload sizes ranging from 1.0 to 5.1 kb and SFn-PA generally showed higher efficiency than SF-PA. For instance, at the AAVS1 locus, SFn-PA achieved an average efficiency of 36.2%, whereas SF-PA achieved an efficiency of 27.6% (Fig. 2f). When an additional nick was induced in the target strand instead of the opposite strand, the PA efficiency of SFn-PA remained similar to that of SF-PA (Extended Data Fig. 2b,c), indicating that the additional nick on the opposite strand is critical for enhancing PA efficiency. Next, payloads ranging from 1.0 to 6.5 kb in size were tested at multiple endogenous sites for DF-PA with flap-to-flap distance of 100–200 bp. All cases presented PA events, with a maximum of 57.8% occurring with a 2.9-kb donor at the HEK4 locus (Fig. 2g and Supplementary Table 3). Typically, DF-PA exhibited relatively high integration efficiency at short flap-to-flap distances of 100–200 bp, regardless of payload size (Extended Data Fig. 2d).

To examine whether the PA system could function in different cell lines, we tested three PA tools in three additional cell types, including HEK293, HeLa and K562 cells, at the VEGFA and HEK4 loci. The PA efficiencies varied but PA replacement occurred for all endogenous targets in all tested cell lines (Fig. 2h). These findings indicate that PA tools enable efficient large-scale replacement at multiple endogenous loci across various human cells. In addition to DNA plasmids as PA donors, we tested various donor types, including circular ssDNA (cssDNA) and linear double-stranded DNA (dsDNA) and the results confirmed that PA is applicable to these various types of donors (Extended Data Fig. 3).

Oxford Nanopore sequencing was performed in the CMV knock-in HEK293T cell line to validate the accuracy of PA-mediated replacement at on-target sites (4.3-kb donors for SF-PA and SFn-PA and a 2.9-kb donor for DF-PA; Fig. 3a). For genomic DNA, Cas9–ribonucleoprotein (RNP) complexes were treated to cleave upstream of the PA target region (that is, on-target site) and these fragments were ligated with unique molecular identifier (UMI)-containing adaptors, amplified by PCR29 and subjected to long-read sequencing. The resulting PA-mediated precise integration efficiencies were 6.6%, 15.7% and 12.9% for SF-PA, SFn-PA and DF-PA, respectively, while imperfect integration efficiencies were 0.5%, 2.9% and 3.8% for SF-PA, SFn-PA and DF-PA, respectively (Fig. 3b). Specifically for DF-PA, additional integration events were observed, with deletion of genomic sequences between flaps and partial integration of flap-derived sequences, accounting for 6.4% (Fig. 3b and Extended Data Fig. 4). These data indicate that SF-PA and SFn-PA possess more accurate integration features, compared to DF-PA.

a, Schematic of a UMI-based nanopore sequencing strategy. Genomic DNA was dephosphorylated and cleaved with Cas9 RNPs, followed by ligation of UMI adaptors, PCR amplification and nanopore sequencing to classify integration and wild-type alleles. b, Classification of sequencing reads for SF-PA, SFn-PA and DF-PA into five categories. Reads were classified as wild type, incomplete flap generation resulting in genome deletion or donor-integrated forms. Donor integration events were further subdivided into precise integration, integration with flap sequence damage and integration with internal donor sequence damage on the basis of sequence integrity at the flap and donor regions. c, Schematic of genome-wide off-target integration assay. Genomic DNA was sheared into ~500-bp fragments, followed by end repair, dA tailing and then adaptor ligation. Residual plasmid sequences were depleted by in vitro Cas9–RNP cleavage. The subsequent fragments were amplified by two-step PCR with indexing in the final step, before next-generation sequencing (NGS) library preparation. d, Genome-wide profiling of integration events with on-target and potential off-target reads for each condition.

Next, we evaluated off-target integration events in the same experiments using PA tools. Thus, we designed an integration-searching method to identify genome-wide insertion sites using Illumina short-read sequencing, referencing the GUIDE-seq process30. In brief, genomic DNA was prepared and broken into approximately 500-bp fragments. After performing Cas9–RNP digestion to deplete any residual donor plasmids, these fragments were ligated with adaptors and amplified by PCR using a forward primer for ligation sequences and a reverse primer for integration sequences (Fig. 3c). The resulting libraries enabled genome-wide profiling of insertion events. No notable off-target integration events exceeding 0.06% of total reads were detected for any PA tool. By contrast, on-target integration at the AAVS1 locus yielded 5,386 reads for SF-PA, 5,170 reads for SFn-PA and 6,855 reads for DF-PA (Fig. 3d).

To benchmark the performance of PA against currently available integration methods, we compared PA tools to DSB-based approaches, including nonhomologous end joining (NHEJ), HDR, microhomology-mediated end joining (MMEJ)31, homology-mediated end joining (HMEJ)32 and PAINT, as well as a DSB-independent approach, evolved PASSIGE (eePASSIGE)4, in the CMV knock-in HEK293T cells (Extended Data Fig. 5a and Supplementary Table 4). All methods were evaluated using donor constructs with the identical payload size (4.3 kb). The donor plasmid encoding copGFP without a promoter was transfected into cells with each tool and integration efficiency was subsequently quantified by flow cytometry, while indel events were assessed by Illumina short-read sequencing at the Cas9 target site. Compared to all the methods we tested, PA tools achieved higher integration efficiency with much lower levels of undesired editing (Extended Data Fig. 5b). Notably, SF-PA, SFn-PA and DF-PA yielded integration efficiencies of 32.5%, 35.4% and 50.5%, respectively, while exhibiting low undesired editing frequencies of 0.6%, 0.6% and 0.4%, respectively, at the PE-target site. For DSB-based approaches, HDR, NHEJ, MMEJ, HMEJ and PAINT 3.0 exhibited integration efficiencies of 4.2%, 6.4%, 9.7%, 36.7% and 12.1%, respectively, while all of those methods were accompanied by unwanted indel formation at Cas9 target site, with rates of 65.4%, 45.0%, 45.4%, 45.0% and 64.6%, respectively. By contrast, eePASSIGE, a DSB-independent approach, exhibited an integration efficiency of 18.1% with an almost undetectable indel frequency. However, eePASSIGE mainly installed only Bxb1 recombinase-associated attB sequences without cargo integration, showing a failure rate of 78.8% (Extended Data Fig. 5c). To analyze the HMEJ-mediated integration results in more detail in the kilobase range, Oxford Nanopore sequencing was performed on the HMEJ-treated samples. Precise integration was detected in 7.7% of total reads, while impaired integration and large deletions were observed in 5.8% and 1.2%, respectively, and small indels were frequently detected in 59.3% (Extended Data Fig. 5d). These results indicate that PA provides more accurate integration outcomes than DSB-dependent HMEJ. We also assessed genome-wide off-target integration events using Illumina short-read sequencing. Off-target integration was rarely detected in the HMEJ analysis, which is likely because of the low off-target activity of the sgRNA used in this experiment. By contrast, several off-target integration sites were identified in the eePASSIGE analysis, with the most frequently occurring off-target site accounting for 4.0% of the total sequence reads (Extended Data Fig. 5e).

Source: Read the original article on www.nature.com