Nature Biotechnology
(2026) Cite this article
Microbial symbiosis drives the functional and phylogenomic diversification of life on Earth yet remains underexplored because of culturing challenges. This study used machine learning (ML) to predict symbiotic lifestyles in more than a hundred thousand microbial genomes from diverse environmental metagenome samples and reference genomes. Predictions were performed using symclatron, an ML framework developed to identify genomic signatures of symbionts. Predictions were deposited in a catalog we established called Symbiont Genomes (SymGs). The results indicate that 15–23% of uncultivated microorganisms likely engage in symbiotic relationships with other organisms, categorized as host-associated or obligate intracellular lifestyles, and are present in half of all known bacterial and archaeal phyla. We also identify genomic signatures of symbiotic lifestyles, including the loss of certain metabolic functions and the differential presence of metabolic modules that may enable host-dependent living. The symclatron software and the SymGs catalog represent valuable resources for studying symbioses, potentially facilitating future mechanistic investigations and engineering of host–microorganism associations.
The intricate and persistent relationships between microorganisms and their hosts, known as microbial symbioses, are fundamental to the functioning of life on Earth1,2,3. These dynamic relationships span a spectrum from mutualism to parasitism4 and are integral to the health and evolution of all partners involved5. On the one hand, microbial symbionts have critical roles in processes such as nutrient cycling6,7, transfer of energy in the form of ATP8, enhancement of host stress resistance9 and disease suppression and protection from pathogens10,11,12. On the other hand, microbial pathogens infect eukaryotic cells and adversely affect their host fitness13. The genetic makeup of symbionts often determines host range and interaction4,14; however, the availability of symbiont genomes is skewed toward those that are animal or human associated or amenable to laboratory cultivation15,16,17. This bias has caused researchers to overlook the true breadth of symbiont diversity in nature, potentially missing novel and globally important symbiont clades yet to be identified.
Recent advances in metagenomics have provided the tools to capture and analyze the vast array of genetic material present in environmental samples, offering unprecedented insights into the microbial world18. The ability to reconstruct population genomes, so-called metagenome-assembled genomes (MAGs) and single amplified genomes (SAGs), from environmental sequencing data has revolutionized our understanding of microbial diversity, functional capabilities and ecology19,20,21,22,23,24 and the evolution of the eukaryotic cell25,26. To date, hundreds of thousands of MAGs have been generated27,28,29. While MAG taxonomy can be linked to their predicted metabolic traits, the lifestyles of these organisms remain largely unknown. Taxonomy alone is often insufficient to predict lifestyle because of variation within related groups and the prevalence of novel organisms in MAG datasets. Analyzing genomic features related to function and dependence offers a more direct approach30. In particular, genomes of symbiotic clades have been disproportionately underrepresented in large-scale MAG studies. This bias occurs because estimates of genome completeness often fall below the medium-quality threshold for these genomes31,32,33. The concept of the symbiome encompasses the entire assemblage of colocalized and coevolving organisms, including both host and symbionts34. However, a linkage between symbiont and host can often not be established in complex metagenomic datasets. Yet, it is plausible that many MAGs and SAGs represent uncultivated microbial symbionts; therefore, a symbiont-centric view of the global microbiome data could represent a powerful approach to uncover the full extent of the genomic space of bacterial and archaeal symbionts.
Understanding microbial symbiosis has exciting implications for biotechnology. Symbiotic systems represent naturally evolved solutions for metabolic division of labor, offering blueprints for engineering synthetic microbial consortia for industrial applications35. The streamlined genomes of obligate symbionts provide insights into minimal genome design for synthetic biology applications, while their specialized metabolic capabilities suggest enzymatic tools for biotechnology36. Furthermore, symbiont–host metabolic complementarity can guide the design of engineered microbial communities for bioproduction, bioremediation and agricultural applications37,38. A comprehensive catalog of symbiont genomes would accelerate the discovery of additional biosynthetic pathways, specialized enzymes, metabolic division of labor and metabolic interactions that can be harnessed for biotechnological innovation in the engineering of host–microorganism systems.
Here, we leverage machine learning (ML) models trained on large-scale genomic data to map putative microbial symbionts across the globe. For simplicity, we henceforth refer to these putative symbionts as ‘symbionts’, recognizing that their status remains provisional, being predicted from genomic data. Operationally, in this study, we define a symbiont as an organism that maintains a spatially close and temporally prolonged association with a host organism of a different species, regardless of the function of the symbiont, which can range from mutualistic to parasitic16. Microbial symbionts can be associated with host organisms by attaching to external surfaces or host cells or occurring intracellularly at least during part of their life cycle. Some microorganisms are obligate intracellular and cannot replicate outside a host cell. To account for these differences in host dependence, we broadly categorized microbial symbionts as ‘host-associated’ or, more specifically, as ‘obligate intracellular’. Host-associated symbionts encompass both facultatively intracellular microorganisms, which can replicate inside a host cell but also independently of a host, and ectosymbionts (or epibionts), which inhabit the exterior of their host, often attached to the host’s surface. In contrast, obligate intracellular symbionts represent a subset of host-associated symbionts that thrive and replicate exclusively inside a host cell. We investigate the prevalence of predicted symbionts across the microbial tree of life, systematically detect the taxonomic extent of symbiont-exclusive clades and analyze the functional and metabolic patterns that characterize host-associated and intracellular lifestyles. The results of our large-scale approach enrich our understanding of microbial symbiont prevalence, diversity and metabolic capacity, providing global insights into microbial symbiosis and establishing a resource for potential biotechnological applications ranging from synthetic consortium design to minimal genome engineering and enzyme discovery.
We constructed a comprehensive symbiont proteome database and extracted genomic features to enable lifestyle classification (Fig. 1). The first part included a reference set of 792 symbiont proteomes from across major clades (Fig. 1a) that was used to capture functional signals of symbiosis. The manually curated origin of the 792 symbiont proteomes and the expert-guided labeling criteria for the 6,751 genomes are described in the Methods. We performed orthogroup inference on these symbiont proteomes, from which we built 20,063 high-confidence profile hidden Markov models (HMMs) for orthogroups that contained ≥5 member proteins. The HMMs serve as potential markers of symbiotic gene content (Fig. 1a). Next, we built a dataset comprising 6,751 labeled microbial genomes (Fig. 1b and Supplementary Table 1) available in the Integrated Microbial Genomes and Metagenomes (IMG/M) database29, spanning three lifestyle classes: free-living (n = 5,959), host-associated (n = 409) and obligately intracellular (n = 383). Detailed definitions of these lifestyle labels are provided in the Methods. With the aim of accounting for the effect of genome incompleteness in the downstream training of our models, all the genomes were artificially reduced to different levels of genome completeness and fragmentation lengths (Fig. 1b). We then applied uniform genome gene calling and annotation across the entire labeled dataset, yielding ~244 million proteins (Fig. 1b). Each protein was scanned against the symbiont HMM library (Fig. 1a) to determine orthogroup memberships. As a result, each genome was represented as a 20,063-dimensional feature vector reflecting the bitscore of symbiont-related orthogroups.
a, Feature engineering workflow for detecting the genomic features specific to symbiont genomes. b, Construction of a genus-level representative database of lifestyle-labeled genomes accounting for artificial genome reductions. c, Schematic of the symbiosis continuum concept applied in the labeling strategy for model training. d, Construction of the training matrix for the symbiont classifier (symcla) and symbiont regressor (symreg) using XGBoost. The final models were optimized for the most relevant 1,000 features. e, Strategy implemented to calibrate the final lifestyle prediction by training a neural network on the outputs of the out-of-fold validations. Right, the neural network was trained on 11,025 rounds of benchmarking and validations of the classifier and regressor models by blinding the models to entire clades from the training data. Left, bar plot showing the total number of unique clades analyzed at each taxonomic rank. f, Architecture of the neural network featuring batch normalization and dropout layers (30% and 20%), with three dense layers using ReLU activation. The output layer also yields confidence scores for each lifestyle class (0, free-living; 1, host-associated; 2, obligate intracellular). g, ROC curves showing classification performance for each lifestyle category (one versus rest). ROC curves are shown for the final neural network model evaluated on the stratified 80:20 split of the integrated out-of-fold benchmarking dataset used for neural network model selection. h, Confusion matrix at 0.725 confidence threshold depicting correct classifications for each category. The data split for the neural network validation was applied at 80% for training (n = 385,800 proteomes) and 20% for testing (n = 96,450 proteomes).
Genome labeling was performed following our conceptual demarcation of the symbiosis continuum (Fig. 1c). Then, the 20,063-dimensional encoding captured each genome’s potential host-dependence traits (Fig. 1c). Accordingly, we used the encodings to train a symbiont classifier model (symcla) on the basis of discrete class labels and to train a regressor model (symreg) derived from the continuous labels (Fig. 1d). To optimize computational efficiency, the top 1,000 most relevant features were selected for each model.
To benchmark the model’s generalizability to novel taxa, we implemented a clade-level leave-one-clade-out cross-validation strategy (Fig. 1e). In total, 11,025 unique clades (across species, genus, family, order, class and phylum levels) were represented in the labeled dataset (Fig. 1e, left). For each validation run, we withheld all genomes from one clade (for example, all members of a given family) from training and then tested the model on those held-out genomes (Fig. 1e, right). This process was repeated such that each clade (at each taxonomic rank) was left out in turn, yielding 11,025 clade-specific validations. By removing entire lineages during training, this rigorous scheme evaluates how well the classifier can predict lifestyles for previously unseen lineages.
We developed a feedforward neural network, termed ‘symclatron’, to provide a final prediction of each genome’s lifestyle on the basis of the clade-level cross-validation data and other interpretable genomic features (Fig. 1f). Rather than using thousands of raw ortholog features directly, the symclatron model integrates seven informative features per genome that summarize its position in symbiotic gene-content space. These features capture both the categorical and the continuous signals of symbiosis, as well as genome quality, and are defined below (Fig. 1f).
Symcla class probabilities (three features): the probabilities of the genome being free-living, host-associated or obligate intracellular, as predicted by an initial gradient-boosted trees classifier trained on the 20,063-dimensional feature matrix.
Symreg score (one feature): a continuous symbiosis score ranging from 0.0 (free-living) to 2.0 (obligate intracellular) produced by a gradient-boosted regression model (symreg). This models the genome’s position along the host-dependence continuum (Fig. 1c).
Distance to training datasets (two features): for both the symcla classifier feature space and the symreg feature space, we computed the genome’s weighted Euclidean distance to the nearest training-set genome. These distance metrics (weighted by feature importances from the symcla/symreg models) quantify how atypical or novel the genome’s gene content is relative to known symbionts and free-living organisms.
Genome completeness (one feature): an estimate of genome completeness (assembly quality) for the proteome, included to ensure that partial genomes are appropriately handled.


