Skip to content
Live newsroom 51 readers online
Monday, August 24, 2026 Live Sync: 1 minute ago
Demystifying Finance, Technology, and Global Markets for the Next Generation.
BreakingOil falls ahead of US announcement of new sanctions on Iran
Important BUY INTC Stage 2 (Conv: 5/5 | Size: 20%)

Adaptive model-guided protein evolution with sparse data optimizes compact eukaryotic genome editors

Nature Biotechnology (2026) Cite this article Efficient protein engineering is constrained by vast sequence space and limited experimental throughput, particularly for protein families that lack large mutational datasets. Here we combine Fanzor2 (Fz2) ortholog discovery, ωRNA scaffold engineering and EvoMax, a model-guided prioritization strategy for sparse-data engineering of compact eukaryotic Fz2 nucleases. EvoMax integrates iterative […]

By deepak · August 24, 2026 · 13 min read

Nature Biotechnology
(2026) Cite this article

Efficient protein engineering is constrained by vast sequence space and limited experimental throughput, particularly for protein families that lack large mutational datasets. Here we combine Fanzor2 (Fz2) ortholog discovery, ωRNA scaffold engineering and EvoMax, a model-guided prioritization strategy for sparse-data engineering of compact eukaryotic Fz2 nucleases. EvoMax integrates iterative experimental profiling with Gaussian process regression, protein language models and inverse folding to navigate complex sequence-to-fitness landscapes. Applied to eukaryotic Fz2 nucleases, this strategy yielded a high-performance variant, FanzMAX v3-hLa, achieving up to 97% editing efficiency at the best-performing endogenous locus and a mean editing efficiency of ~33% across 19 endogenous loci, outperforming the established compact genome editors enNlovFz2 and enCnCas12f1 by more than 2.6-fold. In vivo editing of hPCSK9 in humanized mice supported the translational potential of optimized Fz2 editors. Together, these results establish EvoMax as an integrated strategy for engineering compact eukaryotic Fz2 genome editors and identify FanzMAX v3-hLa as a high-efficiency programmable nuclease for mammalian genome editing.

Protein engineering underpins modern biotechnology, enabling the development of high-performance proteins for therapeutic and industrial applications, including genome-editing effectors1,2. Traditional strategies, including rational design, site-directed mutagenesis and directed evolution, have yielded important advances but remain labor intensive and dependent on detailed structural or mechanistic knowledge3. More recently, machine learning has enabled in silico prediction of mutational effects and accelerated protein optimization4,5. Methods such as EVOLVEpro, AiCE and MULTI-Evolve use active learning with protein language models (PLMs) or inverse folding (IF) approaches to guide mutagenesis6,7,8, while related efforts have addressed complementary challenges including protospacer-adjacent motif reprogramming and de novo CRISPR effector design9,10,11. Despite these advances, most existing strategies require large task-specific training datasets or exhibit uncertain transferability across related protein families. A central bottleneck in protein engineering is that most newly discovered protein families lack sufficient experimental data for artificial intelligence (AI)-driven optimization and evolution. This limitation constrains both traditional and computational approaches and highlights the need for data-efficient frameworks capable of leveraging sparse measurements to guide experimentally tractable optimization across related protein scaffolds4,7.

To address this challenge, we focused on Fanzor nucleases, a recently discovered family of compact, eukaryotic programmable nucleases that remains largely unexplored by data-driven engineering approaches compared with other CRISPR systems12,13. Among these, Fanzor2 (Fz2) nucleases, typically under 500 aa, are RNA-guided DNA endonucleases encoded in eukaryotic genomes and are well suited for single-AAV delivery14. Although certain eukaryotic Fz2 variants exhibit broader target-adjacent motif (TAM) flexibility than other compact nucleases15,16,17, achieving consistently high editing activity in bulk mammalian cells has remained challenging. Recent efforts have improved select Fz2 orthologs but have largely relied on scaffold-specific strategies and targeted enrichment18,19. These observations highlight the need for expanded discovery of functional eukaryotic Fz2 orthologs and adaptive model-guided frameworks capable of optimizing diverse scaffolds under sparse experimental constraints.

Here, we address these challenges through an integrated discovery and model-guided engineering framework. We first computationally prioritized and experimentally validated several active eukaryotic Fz2 orthologs from a pool of over 1,600 candidates, expanding the repertoire of functionally characterized Fz2 nucleases. We then developed EvoMax, a protein engineering pipeline that integrates transfer-learning Gaussian process regression (GPR), a PLM (ESM-2) and an IF model (ESM-IF) to prioritize candidate mutations7,20,21. Using this framework, we systematically optimized the effector protein in combination with parallel engineering of the ωRNA scaffold to generate FanzMAX v3-hLa. The final system achieved an approximately tenfold increase in endogenous editing efficiency, reaching up to 97% with a mean editing activity of ~33% across 19 loci in bulk mammalian cells, surpassing previously established benchmarks18,22. The optimized variant also exhibited expanded TAM compatibility, broadening the targeting scope of Fz2s. EvoMax-guided optimization further enabled functional rescue of multiple inactive Fz2 orthologs, supporting transfer of learned fitness constraints within this protein family. Lastly, we demonstrated in vivo genome editing by targeting hPCSK9 in mice using single-AAV delivery. These results demonstrate that coordinated ωRNA scaffold engineering and model-guided protein optimization can substantially improve compact eukaryotic Fz2 nuclease activity, yielding FanzMAX variants with enhanced editing efficiency and therapeutic potential.

Although several eukaryotic Fz2 systems have been reported, including NlovFz2, genome-editing activity in unselected bulk mammalian cells has generally remained modest, with high efficiencies often requiring cell-sorting enrichment18,19,23. To identify Fanzor effectors that combine high potency with compact size suitable for single-AAV delivery, we performed a systematic bioinformatic expansion of the Fz2 family (Supplementary Methods), identifying over 1,600 Fz2-like sequences spanning broad evolutionary diversity (Extended Data Fig. 1)16,23. From this dataset, we prioritized 332 eukaryotic orthologs for high-resolution taxonomic and phylogenetic analysis (Fig. 1a). To identify candidates with enhanced cleavage activity, we selected seven representative orthologs from distinct phylogenetic subclusters for functional validation. Using an in vitro transcription–translation (IVTT)-based TAM screening assay, we evaluated their DNA cleavage activity and TAM preferences (Fig. 1b)24. Among these candidates, the ortholog M7 exhibited robust cleavage activity with a 5′-CCG-3′ TAM preference (Fig. 1c). Therefore, we expanded our analysis to nine additional orthologs within the M7-related subcluster, leading us to select a previously cataloged eukaryotic Fz2 ortholog from Naegleria lovaniensis (XP_044555062.1; hereafter referred to as NaloFz2)12 for functional characterization. NaloFz2 displayed clear biochemical activity in vitro while exhibiting a 5′-CCG-3′ TAM preference (Fig. 1c). Comparative sequence analysis of the M7-related subcluster revealed substantial structural heterogeneity at the N-terminus. Whereas the catalytic RuvC domain and predicted DNA-binding regions were highly conserved, the N-terminal domain (NTD) showed variable integrity, ranging from fully preserved to partially truncated or entirely absent (Fig. 1d and Extended Data Fig. 2). This structural divergence correlated with functional capacity; detectable in vitro DNA cleavage and TAM identification were observed only among candidates retaining intact NTDs, whereas candidates with truncated or missing NTDs were uniformly inactive.

a, Phylogenetic tree of 332 eukaryotic Fz2 candidates. The bars forming the light-gray outer ring are proportional to the size of the Fz2 in aa, as annotated in the database. The evolutionary distance scale of 1 is shown. The selected orthologs are highlighted in blue. b, Design of an IVTT-based TAM screen. c, TAM sequences of two active Fz2 proteins determined by in vitro cleavage assay. d, Domain architecture and MSA of the NTDs (~100 aa) of selected Fz2 candidates (M1–M9) compared to NaloFz2. e, Schematic describing the detection of editing activity based on the fluorescence signal of EGFP or mCherry reporter activation in HEK293T cells. The CXCR4 site 1 with the 5′-ACCG TAM sequence was used as the target locus. f, Percentage of mCherry-positive cells among EGFP-positive cells following Fz2-mediated DSBs quantified by flow cytometry. Mock denotes a spacer with a randomized sequence; NTD1 and NTD2 indicate 20-aa and 30-aa truncations within the NTD of NaloFz2, respectively. Data are presented as the mean ± s.d. (n = 3 independent biological replicates).

To assess their activity in human cells, we used a fluorescence-based reporter assay in HEK293T cells in which mCherry expression is restored following Fz2-mediated double-strand breaks (DSBs) and subsequent single-strand annealing repair (Fig. 1e). Using a CXCR4 target site (site 1) containing a 5′-ACCG TAM, both NaloFz2 and M7 exhibited detectable genome editing, with NaloFz2 producing substantially higher reporter activation (Fig. 1f)25. Notably, deletion of as few as 20–30 N-terminal residues abolished NaloFz2-mediated editing, establishing the NTD as a necessary functional module in eukaryotic Fz2 proteins (Fig. 1f). In viral Fzs, such as ApmFNuc, NTDs have been proposed to function as nuclear localization signals (NLSs)23. However, computational analysis failed to identify strong, positively charged NLS motifs within the N-terminal regions of the eukaryotic candidates tested (Supplementary Fig. 1). These results suggest that the NTD of eukaryotic Fz2 orthologs may instead serve a structural scaffolding role required for productive ribonucleoprotein assembly and genome editing in mammalian cells14,23. Despite possessing an intact NTD, NaloFz2 exhibited only modest editing efficiency, indicating that additional constraints, including ωRNA architecture, limit robust activity.

We first used AlphaFold 3 and RNAfold to model the tertiary structure of the NaloFz2–ωRNA–target DNA complex and identify candidate regions for stabilization (Fig. 2a)26,27. The ωRNA scaffold (120 nt in length) consists of three major stem loop (SL) regions, of which distal elements appeared conformationally flexible and made limited protein contacts, suggesting potential sites for engineering. Guided by these predictions, we systematically modified the distal regions of SL1 (C7–U11), SL2 (U44–U46) and SL3 (U81–U90). Substitution of distal loops with GAAA tetraloops revealed that modification of SL3 produced the most pronounced improvement in editing activity. Building on this modification, further truncation of SL3 enhanced activity, whereas SL2 truncation had minimal effect. We further stabilized the truncated scaffold by replacing weak A–U, G–U or U–U base pairs within the remaining helices with G–C or C–G pairs. These combined modifications generated two minimized ωRNA variants, enωRNA v1 (SL3 Δ11 bp GAAA) and enωRNA v2 (SL3 Δ11 bp U–C GAAA), each 92 nt in length, representing approximately three quarters the length of the wild-type (WT) scaffold while exhibiting substantially increased editing activity (Fig. 2b,c). Sequence alignment across orthologs revealed substantial conservation of core ωRNA elements, with partial truncations observed in the ωRNA of M1 and NaloFz2 (Extended Data Fig. 3a). Therefore, we hypothesized that the optimized enωRNA architecture could function across related Fz2 ortholog subclusters. Application of enωRNA v2 enhanced the editing activity of M7 by 8.7-fold and enabled functional rescue of the previously inactive ortholog M9 when the enωRNA v2 scaffold was paired with replacement of its truncated NTD by that of NaloFz2 (Fig. 2d). Together, these results establish enωRNA as a modular and transferable scaffold for enhancing activity across divergent eukaryotic Fz2 orthologs.

a, Left, predicted three-dimensional structure of the NaloFz2 ribonucleoprotein highlighting SL1, SL2 and SL3 by AlphaFold 3, visualized in ChimeraX. Right, secondary structure of the 120-nt NaloFz2 WT-ωRNA, illustrating the base-pairing interactions and the pseudoknot (PK) fold. b, Functional mutational scanning of the ωRNA scaffold using the mCherry reporter assay. Reporter assays in b and d used CXCR4 site 1. Various deletions (Δ) and base-pair substitutions were tested for editing activity. Data are presented as the mean ± s.d. (n = 3 independent biological replicates). c, Secondary-structure schematic of the 92-nt enωRNA scaffolds (v1 and v2), illustrating the base-pairing interactions and PK fold. d, mCherry reporter assay quantifying editing activity in HEK293T cells using the WT-ωRNA versus the engineered enωRNA v2 across different Fz2 variants. A modified M9 variant containing the NaloFz2 NTD was tested to evaluate the requirement of NTD integrity for Fz2 activity. Data are presented as the mean ± s.d. (n = 3 independent biological replicates). The mock, WT-NaloFz2 and WT-M7 (with WT-ωRNA) control data in d are the same measurements shown in Fig. 1f. e, Screening of NaloFz2 editing activity (mCherry reporter assay using CXCR4 site 2) for various RNA-binding La protein fusions and ωRNA versions. Data are presented as the mean ± s.d. (n = 3 independent biological replicates).

In addition to internal scaffold stability, the compact architecture of Fz2 leaves the 3′ spacer region exposed to degradation by cellular exonucleases. Structural analysis indicated limited shielding of this region by NaloFz2 (Extended Data Fig. 3b). Insertion of auxiliary RNA structures did not improve activity (Extended Data Fig. 3c), prompting us to evaluate protein-based stabilization strategies28. Inspired by prime-editing strategies in which the La RNA-binding domain stabilizes 3′ RNA extensions, we tested whether a similar strategy could enhance ωRNA stability29. We evaluated fusions of La orthologs from multiple species to either the N- or C-terminus of NaloFz2. Fusion of human La (hLa) to the C-terminus of NaloFz2 consistently produced the highest editing activity under stringent assay conditions (using CXCR4 site 2), independent of whether enωRNA v1 or v2 was used (Fig. 2e).

To enable efficient mammalian genome editing with NaloFz2, we undertook a systematic protein engineering effort. However, conventional optimization strategies, including experimental directed evolution and rational mutagenesis, are poorly suited to CRISPR nucleases. Directed evolution requires iterative construction of large combinatorial libraries coupled with extensive high-throughput screening, rendering it impractical in mammalian systems (Extended Data Fig. 4)30. Rational design, while structurally informed, often fails to capture the functional complexity of newly identified effectors such as Fz2. Consistent with these limitations, our preliminary analyses indicated that rationally designed NlovFz2 variants showed a low success rate, with only ~25% of mutations improving activity and most being neutral or deleterious (Extended Data Fig. 5a).

To address these constraints, we developed EvoMax, a sequence-informed and structure-constrained framework for identifying beneficial mutations under limited experimental sampling. EvoMax follows a three-component workflow comprising data curation, empirical model selection and hierarchical multiobjective optimization (Fig. 3a). A central challenge in engineering the functionally characterized NaloFz2 is the limited availability of high-quality labeled mutational data. Therefore, we used a transfer-learning strategy and assembled a training dataset of 209 single-point mutations with measured activity values, combining an internal library of NlovFz2 variants with previously published functional datasets (Extended Data Fig. 5a)18. Given the close evolutionary and functional relationship between NlovFz2 and NaloFz2, this strategy captures conserved biochemical constraints across the Fz2 family. During model selection, we benchmarked multiple supervised learning algorithms, including random forest and XGBoost. A GPR model using BLOSUM62-derived encoding demonstrated the strongest empirical performance, achieving an R2 score of 0.693 and a root-mean-squared error (r.m.s.e.) of 0.232 (Fig. 3b,c)20. This model encodes biochemical similarity between amino acid substitutions and serves as the initial empirical component of EvoMax’s two-stage hierarchical selection framework (Extended Data Fig. 5b–e). To efficiently explore sequence space while enabling prioritization beyond the initial labeled set, EvoMax uses an adaptive scoring workflow (Eqs. 1–4). In stage 1, a high-potential shortlist comprising the top 1.5% of candidates is generated using a weighted combination of GPR-predicted activity and ESM-2 evolutionary scores (Eq. 5), with the latter providing priors derived from masked-language modeling by a 650-million-parameter PLM (Fig. 3d,e)21. In stage 2, shortlisted variants are reranked using a structure-informed component (ESM-IF) that evaluates compatibility with the AlphaFold-predicted NaloFz2 backbone, while retaining the contributions from GPR and ESM-2 to integrate empirical, evolutionary and structural information (Fig. 3f and equation (6)). This hierarchical framework prioritizes substitutions that capture shared fitness constraints within the Fz2 family, allowing efficient identification of high-performance variants. EvoMax was applied in successive optimization cycles to iteratively select and validate candidate mutations.

a, Data-to-protein-property prediction pipeline. b,c, Benchmarking by R2 (b) and r.m.s.e. (c). d, ESM-2 masked-mutant probabilities by position. e, Principal component analysis of position-level ESM-2 representations; substitutions at the same residue share the masked-position representation. f, GPR, ESM-2 and ESM-IF integration for iterative screening.

To evaluate the performance of EvoMax, we conducted three sequential rounds of protein engineering for NaloFz2, using the top-performing variant from each round as the starting point for the subsequent iteration. Across cycles, the relative weighting of empirical prediction, evolutionary priors and structural constraints was adaptively adjusted to balance local optimization with broader exploration of sequence space (Supplementary Fig. 2). In round 1, stage 1 used a 90:10 (GPR:ESM-2) configuration, whereas stage 2 emphasized empirical activity prediction (90% GPR, 5% ESM-IF, 5% ESM-2). In round 2, stage 1 used a 70:30 (GPR:ESM-2) ratio, whereas stage 2 shifted toward structural constraints (90% ESM-IF, 5% GPR, 5% ESM-2). By round 3, stage 1 adopted a 35:65 (GPR:ESM-2) configuration to promote exploration of more distal beneficial mutations, whereas stage 2 increased the evolutionary contribution of ESM-2 to 70%, supported by 25% ESM-IF and 5% GPR (Fig. 4a). Across rounds, this adaptive strategy drove progressive and substantial improvements in genome-editing efficiency, with 10–20 candidates selected per round for validation using a fluorescence reporter assay in HEK293T cells (Fig. 4b). EvoMax consistently produced improved variants across all three rounds, including in round 3 where the starting variant was already highly optimized (Fig. 4b,c). When benchmarked against EVOLVEpro and AiCE predictions on the same round 1 starting background (WT-NaloFz2), EvoMax achieved a higher hit rate for functional variants (84%) compared to EVOLVEpro (20%) and AiCE (35%) (Fig. 4c). The best-performing variants from each round, designated as FanzMAX v1, v2 and v3, exhibited stepwise improvements, culminating in FanzMAX v3, which showed a more than 13-fold increase in mCherry activation compared with WT-NaloFz2 paired with the WT-ωRNA, when combined with the optimized enωRNA v2 (Fig. 4d).

Source: Read the original article on www.nature.com