{"id":29242,"date":"2026-08-13T04:13:02","date_gmt":"2026-08-13T04:13:02","guid":{"rendered":"https:\/\/futureknowledge.in\/?p=29242"},"modified":"2026-08-13T04:13:02","modified_gmt":"2026-08-13T04:13:02","slug":"mechanistic-machine-learning-for-prediction-of-prime-editing-outcomes","status":"publish","type":"post","link":"https:\/\/futureknowledge.in\/?p=29242","title":{"rendered":"Mechanistic machine learning for prediction of prime editing outcomes"},"content":{"rendered":"<p>Nature Biotechnology<br \/>\n                             (2026) Cite this article<\/p>\n<p>Prime editing (PE) can make specific local changes to genomic DNA in living systems but its efficient application currently requires extensive optimization of PE guide RNA (pegRNA) sequences. Here we present OptiPrime, a machine learning model of PE efficiency based on current understanding of PE mechanisms. OptiPrime achieves state-of-the-art accuracy on PE efficiency prediction and enables prediction of nicking guide RNA (PE3) and dual pegRNA (twinPE) outcomes. We validate that OptiPrime has learned the determinants of mammalian mismatch repair (MMR) and is well suited for nominating MMR-evasive silent edits that improve PE efficiency. We demonstrate the use of OptiPrime in a variety of prospective therapeutic contexts in primary human and mouse cells. Lastly, we show that OptiPrime can be used to achieve streamlined and efficient in vivo correction of a pathogenic mutation in the brain of a mouse model of KIF1A-associated neurological disorder. We provide a webserver for OptiPrime (https:\/\/optipri.me\/) as a community resource.<\/p>\n<p>Precision gene-editing technologies have revolutionized research in the life sciences, as well as the treatment of both heritable and nonheritable diseases in the clinic1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48. Prime editing (PE) is a versatile gene-editing method that uses a programmable nickase and a PE guide RNA (pegRNA) to nick a target DNA sequence at a specified location, reverse-transcribe the edited sequence from the pegRNA onto the nicked target DNA strand and guide the cell through DNA repair processes to replace the original DNA sequence and make the resulting edit permanent on both DNA strands49 (Fig. 1a). PE can install virtually any substitution, small insertion, small deletion or combination thereof in the genome of a living cell without requiring double-stranded breaks (DSBs) or a donor DNA template (Fig. 1b). For example, PE has been used clinically to correct a 2-bp deletion in NCF1, resulting in the effective treatment of chronic granulomatous disease (CGD) in multiple individuals50.<\/p>\n<p>a, Prime editors write genetic information encoded by a pegRNA directly into the genome of a cell with a reverse transcriptase (RT). After 3\u2032-flap resolution, DNA repair processes, including MMR, can resolve the heteroduplex into either the unedited sequence or the edited sequence (Supplementary Fig. 1). b, Prime editors can effect small substitutions, insertions, deletions and combinations thereof. c, The search space of PE strategies is vast. Designing a pegRNA requires choosing (1) a PE protospacer; (2) MMR-evasive silent edits; (3) RTT length; and (4) PBS length.<\/p>\n<p>Although the basic mechanistic steps of PE are well understood (Supplementary Fig. 1), achieving efficient PE, especially in challenging cell types, can require evaluating hundreds or thousands of combinations of pegRNA design parameters to identify a pegRNA that maximizes PE efficiency51,52,53,54,55,56 (Fig. 1c). Each pegRNA contains a spacer that directs the prime editor to its genomic target site, a reverse transcriptase template (RTT) that encodes the edit to be installed and a primer-binding site (PBS) that hybridizes to the DNA primer to initiate reverse transcription. All three of these components are critical determinants of PE outcomes and many options are often plausible for each. In addition, we and others previously showed that evading cellular mismatch repair (MMR) substantially increases PE efficiency and incorporating additional silent or benign mutations\u2014of which there are typically many possibilities\u2014can substantially improve PE outcomes by causing PE intermediates to be poor substrates for MMR57,58.<\/p>\n<p>The vast number of potential pegRNAs that could support a given prime edit of interest, together with the high variability of PE efficiencies among these possible pegRNAs, complicates the use of PE for research and therapeutic applications. To address this challenge, we and others have sought to develop a machine learning (ML) model that computationally identifies pegRNA designs with the highest potential to support efficient PE. The ability of such a model to decrease the number of experiments required to identify a high-performing PE strategy is dependent on the accuracy of the model. Currently published ML models for PE efficiency include PRIDICT54,56, DeepPrime55,59 and OPED60. Although these models achieve high accuracy on test datasets that were sampled from the large datasets on which they were trained, a substantial portion of each of these datasets contains poor-performing pegRNAs that would be screened out during an initial screen. Consequently, although these models can reject low-efficiency pegRNAs in prospective design contexts, they struggle to accurately predict optimal pegRNAs amid a set of well-performing ones, as discussed above. Moreover, these existing models are not based on the biological mechanism of PE. As a result, they may not take advantage of mechanistic insights that suggest key nodes that strongly determine PE outcomes and their lack of correspondence with the molecular processes involved in PE makes it difficult to glean biological insights from their trained state and from their output.<\/p>\n<p>Here, we describe the development of OptiPrime, an ML model that generates predictions of PE efficiency by directly incorporating knowledge about the biological mechanism of PE into its mathematical structure. By design, OptiPrime uses separate ML models to predict \u2018pseudorates\u2019 of biochemical steps involved in the PE mechanism. Each pseudorate model only uses features that are biologically relevant to each process to make its prediction, ensuring a clean separation of tasks. The system of differential equations described by these rates is then integrated through time to obtain a final output prediction of editing efficiency. This approach contrasts with those used by PRIDICT54,56 and DeepPrime55,59, which incorporate all sequence features and representations directly into a single black-box predictor for PE efficiency. To train OptiPrime, we performed matched pegRNA\u2013target site screens in both MMR-deficient HEK293T cells and MMR-competent HeLa cells, resulting in a dataset of 74,769 PE efficiencies across 1,290 unique target sites that enabled us to systematically study both MMR-dependent and MMR-independent sequence determinants of editing efficiency. While traditional ML models usually do not directly include different experimental conditions as covariates and instead rely on model fine-tuning with additional data to make predictions in new contexts, the mechanistic structure of OptiPrime enabled its joint training on a dataset of 297,962 PE efficiencies collected across 40 experimental contexts in our laboratory and others. We demonstrate, through benchmarking and ablation studies, that OptiPrime offers best-in-class accuracy on PE efficiency prediction and its performance is dependent on its mechanism-based mathematical structure.<\/p>\n<p>Moreover, we show that the measurable effects of individual mechanistic steps in the PE process are consistent with their corresponding pseudorate constants in OptiPrime, these pseudorates can be used to train a model of PE3 efficiency and OptiPrime can make useful predictions about twinPE (a dual-flap, two-pegRNA application of PE61), even though the model was never trained on twinPE data. Lastly, we demonstrate the ability of OptiPrime to streamline the development of therapeutic PE applications in several prospective contexts, including the rapid development of a PE strategy to correct a pathogenic mutation in Kif1a that resulted in &gt;40% in vivo PE efficiency in the bulk brain cortex of mice. Collectively, these findings demonstrate that OptiPrime augments the potential of PE by accelerating the identification of efficient editing strategies, and that mechanistic ML models can provide insights into the dynamics of complex biomolecular processes, enabling modular, high-performance prediction of their outcomes in living systems.<\/p>\n<p>In a previous CRISPR interference screen of DNA repair and DNA metabolism genes, we discovered that MMR is a major bottleneck for PE efficiency57. This insight led us to develop the PE4 and PE5 editing strategies, in which a dominant negative variant of the MMR protein MLH1 (MLH1dn) is codelivered with the editing components to suppress MMR and favor other repair pathways that lead to desired PE outcomes57. Consistent with these results, we also found that the use of pegRNAs that install additional silent or benign mutations near the desired edit could substantially increase PE efficiencies by causing heteroduplex PE intermediates to contain sufficient mismatches between the edited and unedited strands to evade recognition by MMR proteins, which preferentially engage single-nucleotide or small mismatched regions57.<\/p>\n<p>The MMR pathway is initiated upon mismatch binding by the MutS\u03b1 or MutS\u03b2 complexes62, which together recognize small mismatches and small insertion\u2013deletion loops63,64. To comprehensively evaluate the sequence determinants of binding by MutS\u03b1 and MutS\u03b2, we designed \u2018Lib-MMR\u2019, a set of 10,000 pegRNAs designed to target 200 randomly selected exonic sites in the human genome with diverse edit types, including single-base substitutions (1,592), contiguous substitutions of 2\u20135 bases (2,792), noncontiguous substitutions of 2\u20135 bases (2,780), deletions of \u226410 bases (1,000), insertions of \u226410 bases (1,000), protospacer-adjacent motif (PAM) edits with deletions (400) and PAM edits with insertions (400). Lib-MMR also included 36 positive control pegRNAs that we previously used to edit endogenous genomic sites (Supplementary Table 1)57.<\/p>\n<p>Although Lib-MMR contains diverse edit types and is, therefore, useful for assessing the determinants of binding by the MutS\u03b1\/MutS\u03b2 complexes, we also sought to generate a dataset that would include many therapeutically relevant edits. Therefore, we designed \u2018Lib-CV\u2019, which comprises 10,406 pegRNAs that correct 944 known pathogenic variants in protein-coding regions from the ClinVar database65. In addition to a pegRNA that directly corrects each variant, we designed up to nine pegRNAs that also included benign \u2018silent edits\u2019 that would not alter the corrected protein-coding sequence (Supplementary Table 2). Importantly, we designed the pegRNAs in Lib-MMR and Lib-CV with empirically determined heuristics to support efficient PE52 (Methods) such that an ML model that performs well on the resulting dataset will be required to distinguish between several candidates that perform well, rather than between rare candidates that perform well amidst many poor-performing alternatives.<\/p>\n<p>To assay PE efficiencies in a high-throughput manner, we packaged Lib-MMR and Lib-CV into lentiviral libraries of pegRNA\u2013target site pairs. We transduced these libraries into cells at low multiplicity of infection (MOI\u2009&lt;\u20090.3) to ensure that most transduced cells were only infected with a single lentivirus, selected for cells with integrated library members and then transfected these cells with PE construct plasmids to induce editing (Fig. 2a). To thoroughly interrogate the effects of MMR on PE outcomes, we performed the screens with PE2 and PE4 in both HEK293T cells, which are partially MMR deficient, and HeLa cells, which are MMR proficient. As expected, we found that PE4 consistently outperformed PE2 across Lib-MMR. This effect was especially pronounced in HeLa cells, in which we observed a median 4.3-fold increase in editing rates with PE4 versus PE2, compared to only a median 1.5-fold increase in HEK293T cells (Fig. 2b,c and Supplementary Fig. 2a), consistent with previous observations that HEK293T cells are MMR deficient. We chose to use the ratio of PE4 editing efficiency to PE2 editing efficiency (PE4:PE2) as a proxy for the propensity of a given edit to be reverted to the unedited sequence by MMR. We found that PE4:PE2 correlated more strongly with PE2 in HeLa cells (Spearman correlation \u03c1\u2009=\u2009\u22120.793) than in HEK293T cells (\u03c1\u2009=\u2009\u22120.390), suggesting that while MMR is the primary determinant of PE2 efficiency in MMR-proficient cell types, other factors determine PE2 efficiency in MMR-deficient cell types (Supplementary Fig. 2b).<\/p>\n<p>a, Paired pegRNA\u2013target site lentiviral libraries enable high-throughput evaluation of PE2 and PE4 outcomes in HEK293T and HeLa cells. b, A scatter plot comparing editing efficiencies across Lib-MMR between PE2 (x axis) and PE4 (y axis) in HeLa cells. The gray line plots the equation y\u2009=\u2009x. PE4 consistently outperforms PE2 in HeLa cells. Brighter colors correspond to increased point density (Supplementary Fig. 2a). c\u2013h, PE4:PE2 represents the ratio of PE4 editing efficiency to PE2 editing efficiency across library members in Lib-MMR. All PE4:PE2 values are plotted on a logarithmic scale. PE4:PE2 values below 0.25 (HeLa) or 0.5 (HEK293T) are clipped to each respective minimum value. PE4:PE2 values above 1,024 (HeLa) or 4 (HEK293T) are clipped to each respective maximum value. c, Beeswarm plots comparing PE4:PE2 between HeLa cells and HEK293T cells. PE4:PE2 is higher in HeLa cells than in HEK293T cells. d, A scatter plot comparing PE4:PE2 across library members between cell types. PE4:PE2 correlates well between HEK293T and HeLa cells. Brighter colors correspond to increased point density. e, Beeswarm plots of PE4:PE2 in HeLa cells for all twelve possible single-base substitutions, sorted by median PE4:PE2 (Supplementary Fig. 2e). f\u2013h, Beeswarm plots comparing PE4:PE2 in HeLa cells across varying lengths of contiguous substitutions (f), pure insertions (g) and pure deletions (h) (Supplementary Fig. 2f\u2013h). i, A scatter plot comparing PE4:PE2 editing efficiency ratio of each edit to the ratio of PE2 editing efficiency with and without silent edits across Lib-CV in HeLa cells. Improvements in PE2 efficiency from including silent edits are positively correlated with PE4:PE2 efficiency ratios of the reference edit without any silent edits in HeLa cells. The red line is a total least squares regression line. Brighter colors correspond to increased point density.<\/p>\n<p>We found that, while PE4:PE2 values were generally lower in HEK293T cells than in HeLa cells as expected given the partial MMR deficiency of HEK293T cells, mean PE4:PE2 values correlated well between the cell types (\u03c1\u2009=\u20090.733), suggesting that trends in binding by MutS\u03b1 and MutS\u03b2 are conserved (Fig. 2d). Moreover, the PE4:PE2 ratios for single-base substitutions in Lib-MMR were highly consistent and linearly correlated with those obtained at endogenous genomic loci from our previous work (Pearson correlation r\u2009=\u20090.800; Supplementary Fig. 2c), with G-to-C and A-to-G edits resulting in the lowest median PE4:PE2 ratios and C-to-T edits showing the highest median PE4:PE2 ratios57 (Fig. 2e and Supplementary Fig. 2d). This ranking is consistent with the known propensity of the corresponding mismatched DNA intermediates to be repaired in vitro by HeLa cell extracts66. Although both A-to-G and C-to-T edits create mismatched T:G intermediates, these edits qualitatively show very different PE4:PE2 ratios, suggesting that T:G mismatches may be asymmetrically repaired by MMR or preferentially processed by an orthogonal repair pathway.<\/p>\n<p>Additionally, we observed that edits consisting of longer contiguous substitutions resulted in decreasing PE4:PE2 values, with Spearman \u03c1 values of \u20130.484 (HeLa) and \u20130.427 (HEK293T) (Fig. 2f and Supplementary Fig. 2e). These data are consistent with structural observations that MutS\u03b1 preferentially binds small mismatched regions rather than larger stretches of mismatched nucleotides63. Similarly, PE4:PE2 for both pure insertions and pure deletions decreased as the number of nucleotides inserted or removed increased, with deletion edits producing much higher PE4:PE2 than the corresponding insertion edits of the same length (Fig. 2g,h and Supplementary Fig. 2g,h). with strong negative rank correlations for both pure insertions (HeLa \u03c1\u2009=\u2009\u22120.638, HEK293T \u03c1\u2009=\u2009\u22120.624) and pure deletions (HeLa \u03c1\u2009=\u2009\u22120.632, HEK293T \u03c1\u2009=\u2009\u22120.615). These data together suggest that known structural features of MMR proteins may underlie differences between PE2 and PE4 across a broad range of editing contexts.<\/p>\n<p><em>Source: <a href='https:\/\/www.nature.com\/articles\/s41587-026-03261-7' target='_blank'>Read the original article on www.nature.com<\/a><\/em><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Nature Biotechnology (2026) Cite this article Prime editing (PE) can make specific local changes to genomic DNA in living systems but its efficient application currently requires extensive optimization of PE guide RNA (pegRNA) sequences. Here we present OptiPrime, a machine learning model of PE efficiency based on current understanding of PE mechanisms. OptiPrime achieves state-of-the-art [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":29243,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[37,3],"tags":[69,29,33],"class_list":["post-29242","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-commodities","category-technology","tag-impact-eth","tag-signal-avoid","tag-stage-stage-4"],"_links":{"self":[{"href":"https:\/\/futureknowledge.in\/index.php?rest_route=\/wp\/v2\/posts\/29242","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futureknowledge.in\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futureknowledge.in\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futureknowledge.in\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futureknowledge.in\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=29242"}],"version-history":[{"count":0,"href":"https:\/\/futureknowledge.in\/index.php?rest_route=\/wp\/v2\/posts\/29242\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futureknowledge.in\/index.php?rest_route=\/wp\/v2\/media\/29243"}],"wp:attachment":[{"href":"https:\/\/futureknowledge.in\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=29242"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futureknowledge.in\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=29242"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futureknowledge.in\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=29242"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}