Nature Biotechnology
(2026) Cite this article
Cell surface display (CSD) elements are a major class of bioengineering modules, but systematic rules linking CSD sequence to functional potency are lacking. Here we develop DeepSCan, a suite of deep learning and artificial intelligence models to systematically map CSD sequence–function relationships to reliably capture potent elements’ conserved features and iteratively design and develop potent de novo CSD modules for mRNA antigen display. We experimentally quantify surface expression across >570 chimeric antigens, derive cell surface translocation strength labels for ~310 CSD elements and compile ~45 independent training datasets. We train three generations of DeepSCan models and evaluate their performance. Guided by these models, we computationally design 3,700 and experimentally validate approximately 120 generative CSDs, identifying 7 generative CSDs that match or exceed the cell surface translocation strength of the most potent naturally occurring CSDs. Enhanced surface displays are validated across multiple cell types and are functional in an antigen-specific CAR-T cytotoxicity assay.
This is a preview of subscription content, access via your institution
Access Nature and 54 other Nature Portfolio journals
Get Nature+, our best-value online-access subscription
Receive 12 print issues and online access
Prices may be subject to local taxes which are calculated during checkout
All primary data supporting the findings of this study are available within the Article and via Zenodo at https://doi.org/10.5281/zenodo.19179097 (ref. 62). Source data are provided with this paper.
Protein language models were implemented using the code and pretrained parameters available from their official GitHub repositories, following the provided instructions. The ESM2 code and models are available via GitHub at https://github.com/facebookresearch/esm. We used XGBoost and other classical ML models (for example, logistic regression, random forest) from the RAPIDS cuML library for accelerated training on graphics processing unit (GPU). Fine-tuned DL model checkpoints and source code will be publicly released on Hugging Face, Docker and Github (https://github.com/fangzhe3/DeepSCan)57. The GitHub repository also introduces three demo scenarios using the DeepSCan web server or Docker image. The Docker image of the Genesis Quant 8,000 model enables users to process large-scale sequence analyses using GPUs. The supplementary software contains the DeepSCan web server application (version: app_v7.0_deployed_on_AWS), Genesis model training code (version: Genesis_Quant_8k_attention_pooling2_LLRD_v2), Genesis and Omni model inference codes and additional computational analyses codes. The DeepSCan models assume Python (v.3.11.5) + PyTorch (v.2.5.1 with CUDA v.12.4)-based workflow.
Overington, J. P., Al-Lazikani, B. & Hopkins, A. L. How many drug targets are there? Nat. Rev. Drug Discov. 5, 993–996 (2006).
Article
CAS
PubMed
Google Scholar
Hegde, R. S. & Keenan, R. J. The mechanisms of integral membrane protein biogenesis. Nat. Rev. Mol. Cell Biol. 23, 107–124 (2022).
Article
CAS
PubMed
Google Scholar
Bugge, K., Lindorff-Larsen, K. & Kragelund, B. B. Understanding single-pass transmembrane receptor signaling from a structural viewpoint—what are we missing?. FEBS J. 283, 4424–4451 (2016).
Article
CAS
PubMed
Google Scholar