Attention Presenters - please review the Speaker Information Page available here
Schedule subject to change
All times listed are in CET
Monday 31 August
13:30-14:30
Session: Advances in Gut Microbiota Analysis, linking microbiota to health
Proceedings Presentation: A Spectrum of Gut Microbiota Signatures Predicts Long-Term Health Status
Confirmed Presenter: Himmi Lindgren, Department of Computing, University of Turku, Turku, Finland, Finland

Room: Room EF
Moderator(s): Avidan Neumann; Lucas Paoli


Authors List: Show

  • Himmi Lindgren, Department of Computing, University of Turku, Turku, Finland, Finland
  • Chandler Ross, Department of Computing, University of Turku, Turku, Finland, Finland
  • Matti Ruuskanen, Department of Life Technologies, University of Turku, Turku, Finland, Finland
  • Veikko Salomaa, Department of Internal Medicine, University of Turku, Turku, Finland, Finland
  • Rob Knight, Department of Pediatrics, University of California San Diego, LA Jolla, CA, USA, United States
  • Teemu Niiranen, Department of Clinical Medicine, University of Turku, Turku, Finland, Finland
  • Aki Havulinna, Department of Computing, University of Turku, Turku, Finland, Finland
  • Leo Lahti, Department of Computing, University of Turku, Turku, Finland, Finland

Presentation Overview: Show

Motivation: The human gut microbiota is a complex ecosystem composed of hundreds of unique microbial species in a typical adult individual. Its composition has been linked to an individual's current and future health. At the population level, joint analysis of co-abundant species can capture associations with health status and other variables that cannot be robustly detected by analyzing individual taxa alone. Recently, the concept of enterosignatures was proposed to detect such co-abundant groups, decomposing microbial community variation into a mixture of five broad components of co-varying microbial taxa. However, their overall robustness, generalizability, and prospective potential in predicting long-term health status remain to be evaluated.
Results: We demonstrate that robust signatures of co-abundant species can be prospectively associated with mortality, diabetes, liver disease, and sepsis risk among Finnish adults in the FINRISK population cohort based on a 20-year follow-up. We characterized dysbiosis-associated signatures across varying levels of sub-ecosystem granularity, encompassing a previously reported Escherichia-driven signature as well as novel Prevotella subspecies-driven signatures linked to compromised future health status. These results illustrate the benefits of analyzing latent signatures across multiple levels of ecosystem granularity.

Improving metagenome binning by integrating intrinsic features and taxonomy
Confirmed Presenter: Simon Rasmussen, University of Copenhagen, Denmark

Room: Room EF
Moderator(s): Avidan Neumann; Lucas Paoli


Authors List: Show

  • Simon Rasmussen, University of Copenhagen, Denmark
  • Svetlana Kutuzova, University of Copenhagen, Denmark
  • Jakob Nybo Andersen, University of Copenhagen, Denmark
  • Mads Nielsen, University of Copenhagen, Denmark
  • Lars Hestbjerg Hansen, University of Copenhagen, Denmark
  • Knud Nor Nielsen, Technical University of Denmark, Denmark

Presentation Overview: Show

A common procedure for studying the microbiome is binning the sequenced contigs into metagenome-assembled genomes. State-of-the-art binning methods use co-abundance and sequence based motifs such as tetranucleotide frequencies while taxonomic labels derived from alignment based classification have not been widely used. Here, we propose TaxVAMB, a metagenome binning tool based on semi-supervised bi-modal variational autoencoders, combining tetranucleotide frequencies and contig co-abundances with taxonomic information. TaxVAMB outperformed all other binners on CAMI2 human microbiome datasets, returning on average 29% more high quality assemblies than the next best binner, and performed on par with the best binners on short-read datasets. On a human gut long-read dataset TaxVAMB recovered 41% more high quality bins and 29% more species. In a typical single-sample setup, TaxVAMB on average returns 83% more high quality bins compared to VAMB. Furthermore, TaxVAMB enables high quality binning of eukaryotic organisms. Finally, TaxVAMB binned incomplete genomes better than any other tool, returning on average 300% more high quality bins of incomplete genomes than the next best binner.

Exploring functional insights into the human gut microbiome via the structural proteome
Confirmed Presenter: Yu Li, Department of Computer Science and Engineering, The Chinese University of Hong Kong, Hong Kong

Room: Room EF
Moderator(s): Avidan Neumann; Lucas Paoli


Authors List: Show

  • Hongbin Liu, State Key Laboratory of Quantitative Synthetic Biology, Shenzhen Institutes of Advanced Technology, Shenzhen, China, China
  • Jiuming Wang, Department of Computer Science and Engineering, The Chinese University of Hong Kong, Hong Kong
  • Yu Li, Department of Computer Science and Engineering, The Chinese University of Hong Kong, Hong Kong
  • Lei Dai, Shenzhen Institute of Synthetic Biology, SIAT, Chinese Academy of Sciences, Shenzhen 518055, China, China

Presentation Overview: Show

Most gut microbiome proteins remain functionally uncharacterized because sequence-based methods miss remote homologs from deep evolutionary divergence. We leverage AI-driven protein structure prediction and representation learning to systematically decode this functional landscape.

We constructed the Human Gut Microbial Protein Structure (GMPS) database of ~2.7 million AlphaFold2-predicted structures from 968 bacterial and 1,255 phage genomes. Structural alignment and clustering substantially outperformed sequence-based annotation, raising putative functional assignments for phage proteins from 24.8% to 54.5%. This revealed structural diversification of phage endolysins through domain shuffling; we experimentally validated nine endolysins with lytic activity against gut pathobionts, including Enterococcus gallinarum and Clostridium innocuum. Structural searches also identified bacterial enzymes for melatonin biosynthesis (DDC, AANAT, and ASMT), whose physiological roles were confirmed by enzymatic assays and in vivo mouse experiments showing direct impacts on host melatonin levels and colitis severity.

To surpass the sensitivity ceiling of alignment-based methods, we developed Dense Enzyme Retrieval (DEER), an alignment-free model fusing pre-trained protein language models (ESM-2 and SaProt) through contrastive learning to project enzymes into a unified latent space where functionally related proteins cluster regardless of sequence or structural divergence. On a cross-kingdom benchmark, DEER outperformed sequence (BLASTp, MMseqs2) and structure (Foldseek, US-align, DALI) aligners for bacteria–human isozyme detection while being orders of magnitude faster, and recovered validated bacterial melatonin enzymes among its top candidates.

Our work establishes an AI-powered framework for functional discovery in the human gut microbiome, illuminating uncharacterized protein space across microbial ecosystems, with broad implications for microbiome-targeted therapeutics and host–microbe interaction discovery.

14:45-15:45
Session: Metagenomic Sequence Analysis (Cleanup, Inverton Detection, and Alignment Visualisation)
Proceedings Presentation: TheBIGbam: compression and interactive exploration of large-scale sequencing alignments with circular mapping support
Confirmed Presenter: Martin Boutroux, EPFL, Switzerland

Room: Room 8
Moderator(s): Lucas Paoli; Katja Bärenfaller


Authors List: Show

  • Martin Boutroux, EPFL, Switzerland
  • Evan Thomas, EPFL, Switzerland
  • Wai Hoe Chin, EPFL, Switzerland
  • Hannes Peter, EPFL, Switzerland

Presentation Overview: Show

TheBIGbam (https://github.com/bhagavadgitadu22/theBIGbam) is a genome browser and alignment viewer designed for massive metagenomic and metatranscriptomic datasets. The tool accepts genome assemblies in FASTA format or annotated genome sequences in GenBank format, together with BAM files containing read alignments. Alternatively, it can start from raw FASTQ reads and generate alignments using a modified mapper that supports circular genomes, enabling seamless read mapping across genome ends. TheBIGbam can compress hundreds of gigabytes of input files 10- to 100-fold into dedicated databases while retaining key per-position information, including coverage depth and recurrent mismatches, insertions, and deletions between reads and the reference. These databases can be served to a local web browser, enabling interactive exploration of any contig in any sample using DNAFeaturesViewer for genome maps and Bokeh for mapping-derived features. Contig-sample pairs available for visualization can be filtered using a range of summary metrics calculated per contig, per sample, and per contig-sample pair to guide users toward the most relevant signals. Through its interactive visualization, theBIGbam facilitates the exploration of complex datasets, while its integrated database, combining assembly features, annotated features, and mapping-derived features, provides the information needed to investigate biological hypotheses systematically. Designed to complement existing browsing tools like IGV and Anvi'o, theBIGbam is particularly suited for examining misassemblies, subpopulations, microdiversity, and contig topology in large-scale datasets.

Cleanifier: contamination removal from microbial sequences using spaced seeds of a human pangenome index
Confirmed Presenter: Jens Zentgraf, Saarland University, Germany

Room: Room 8
Moderator(s): Lucas Paoli; Katja Bärenfaller


Authors List: Show

  • Johanna Elena Schmitz, Center for Bioinformatics, Saarland University, Saarbrücken, Germany, Germany
  • Jens Zentgraf, Saarland University, Germany
  • Sven Rahmann, Saarland University, Germany

Presentation Overview: Show

Motivation: The first step when working with DNA data of human-derived microbiomes is to remove human contamination for two reasons. First, many countries have strict privacy and data protection guidelines for human sequence data, so microbiome data containing partly human data cannot be easily further processed or published. Second, human contamination may cause problems in downstream analysis, such as metagenomic binning or genome assembly. For large-scale metagenomics projects, fast and accurate removal of human contamination is therefore critical.

Results: We introduce Cleanifier, a fast and memory frugal alignment-free tool for detecting and removing human contamination based on gapped k-mers, or spaced seeds. Cleanifier uses a pangenome index of known human gapped k-mers, and the creation and use of alternative references is also possible. Reads are classified and filtered according to their gapped k-mer content. Cleanifier supports two filtering modes: one that queries all gapped k-mers and one that queries only a sample of them. A comparison of Cleanifier with other state-of-the-art tools shows that the sampling mode makes Cleanifier the fastest method with comparable accuracy. When using a probabilistic Cuckoo filter to store the complete k-mer set, Cleanifier has similar memory requirements to methods that use a sampled minimizer index. At the same time, Cleanifier is more flexible, because it can use different sampling methods on the same index.

Availability and implementation: Cleanifier is available via gitlab (https://gitlab.com/rahmannlab/cleanifier), PyPi (https://pypi.org/project/cleanifier/), and Bioconda (https://anaconda.org/bioconda/cleanifier). The pre-computed human pangenome index is available at Zenodo (https://doi.org/10.5281/zenodo.15639519).

Proceedings Presentation: TPMM: Three-component Posterior Mixture Model Enables Robust Inverton Detection in Low-Depth Metagenomes and Suggests Potential Viral Invertons
Confirmed Presenter: Yi Lu, City University of Hong Kong, Hong Kong

Room: Room 8
Moderator(s): Lucas Paoli; Katja Bärenfaller


Authors List: Show

  • Yi Lu, City University of Hong Kong, Hong Kong
  • Jiaojiao Guan, City University of Hong Kong, Hong Kong
  • Yang Shen, City University of Hong Kong, Hong Kong
  • Jiayu Shang, Chinese University of Hong Kong, Hong Kong
  • Yanni Sun, City University of Hong Kong, Hong Kong

Presentation Overview: Show

Bacterial phase variation enables reversible, locus-specific phenotypic switching, often driven by DNA inversion (invertons). To identify these events, researchers commonly rely on sequencing reads that provide orientation-specific support. Metagenomic sequencing, which captures total genetic material independent of cultivation, offers a powerful platform for the comprehensive study of invertons. However, computational inverton calling from metagenomic data is difficult at low sequencing depth: hard read-support cutoffs can miss true events, while sequence-only predictors lack read-backed interpretability and uncertainty quantification. To address this, we present TPMM, a three-component posterior mixture model for inverton calling in metagenomic data. TPMM explicitly incorporates sequencing depth to formulate inverton detection as a probabilistic mixture problem. Starting from candidates flanked by inverted repeats, the model classifies the candidates into noise, low-probability, or high-probability inversion signals using read evidence. Finally, TPMM assigns posterior probabilities as soft labels and applies cumulative Bayesian False Discovery Rate control to robustly identify true invertons. On two real gut metagenomic datasets, TPMM agrees well with PhaseFinder at high depth but recovers substantially more invertons under systematic downsampling, demonstrating superior performance in sparse-data regimes. We further examine potential reversible inversion elements in viral genomes and provide supporting analyses, suggesting a broader scope for inversion-mediated regulation.

Tuesday 1 September
10:30-11:30
Session: LLM and Agent Based Modelling (in Health and Protein-Protein Interactions)
Proceedings Presentation: Agent-Based Modelling Identifies many paths for Inflammation Resolution
Confirmed Presenter: Hermes Desgrez Dautet, Universite de Toulouse, CNRS UMR 5070, INSERM U1301, Institut RESTORE, Institut of Research in Informatics of Toulouse, France

Room: Room EF
Moderator(s): Katja Bärenfaller; Lucas Paoli


Authors List: Show

  • Hermes Desgrez Dautet, Universite de Toulouse, CNRS UMR 5070, INSERM U1301, Institut RESTORE, Institut of Research in Informatics of Toulouse, France
  • David Bernard, Institut of Research in Informatics of Toulouse, Université Toulouse Capitole, France
  • Cousin Beatrice, Universite de Toulouse, CNRS UMR 5070, INSERM U1301, Institut RESTORE, France
  • Paul Monsarrat, Universite de Toulouse, CNRS UMR 5070, INSERM U1301, Institut RESTORE, Toulouse Dental Faculty, France
  • Sylvain Cussat-Blanc, Institut of Research in Informatics of Toulouse, Université Toulouse Capitole, Institut Universitaire de France, France

Presentation Overview: Show

Sterile inflammation emerges from spatially coordinated interactions between immune recruitment, polarization, signal diffusion, and tissue turnover. Capturing this complexity requires modeling frameworks that integrate stochasticity, spatial heterogeneity, and functional redundancy across immune cell types. We present a spatially explicit agent-based model of sterile inflammation based on a functional mono-agent abstraction, in which immune cells share core capabilities, e.g. polarization, chemotaxis, signal production, and phagocytosis, while differing through parameterization rather than rigid rule sets. Using large-scale global parameter exploration via Latin Hypercube Sampling, we generated 300,000 simulations to characterize the model's dynamical landscape. Dimensionality reduction suggested a continuous manifold of inflammatory trajectories. Applying biologically informed temporal criteria identified a highly constrained subspace consistent with canonical sterile inflammation. Within this region, multiple distinct yet self-consistent kinetic regimes achieved successful resolution. These results indicate that sterile inflammation is not a single dynamical attractor but an emergent coordination constraint in a high-dimensional functional space, providing a systems-level framework for exploring immune regulation and dysregulation.

ProteomeLM: A proteome-scale language model enables accurate and rapid prediction of protein-protein interactions and gene essentiality across taxa
Confirmed Presenter: Cyril Malbranke, EPFL, SIB, Switzerland

Room: Room EF
Moderator(s): Katja Bärenfaller; Lucas Paoli


Authors List: Show

  • Cyril Malbranke, EPFL, SIB, Switzerland
  • Gionata Paolo Zalaffi, EPFL, SIB, Switzerland
  • Cecilia Fruet, EPFL, SIB, Switzerland
  • Anne-Florence Bitbol, EPFL, SIB, Switzerland

Presentation Overview: Show

Language models trained on biological sequences are advancing inference tasks from the scale of single proteins to that of genomic neighborhoods. Here, we introduce ProteomeLM, a transformer-based language model that uniquely operates on entire proteomes from species spanning the tree of life. ProteomeLM is trained to reconstruct masked protein embeddings using the whole proteomic context, yielding contextualized protein representations that reflect proteome-scale functional constraints.
Notably, ProteomeLM's attention coefficients encode protein-protein interactions (PPI), despite being trained without interaction labels. Furthermore, it enables interactome-wide PPI screening that is substantially more accurate, and orders of magnitude faster, than amino-acid coevolution-based methods.
We further develop ProteomeLM-PPI, a supervised model that combines ProteomeLM embeddings and attention coefficients to achieve state-of-the-art PPI prediction across benchmarks and species. Finally, we introduce ProteomeLM-Ess, a supervised gene essentiality predictor that generalizes across diverse taxa. Our results demonstrate the potential of proteome-scale language models for addressing function and interactions at the organism level.

Proceedings Presentation: Relation Extraction for Diet, Non-Communicable Disease and Biomarker Associations (RECoDe): A CoDiet study
Confirmed Presenter: Donghee Choi, Imperial College London; Pusan National University, United Kingdom

Room: Room EF
Moderator(s): Katja Bärenfaller; Lucas Paoli


Authors List: Show

  • Donghee Choi, Imperial College London; Pusan National University, United Kingdom
  • Yajie Gu, University of Nottingham, United Kingdom
  • Kai Qi Zong, Imperial College London, United Kingdom
  • Antoine Lain, Imperial College London, United Kingdom
  • Dimitrios Zaikis, Aristotle University of Thessaloniki, Greece
  • Thomas Rowlands, University of Nottingham, United Kingdom
  • Marek Rei, Imperial College London, United Kingdom
  • Tim Beck, University of Nottingham, United Kingdom
  • Joram Posma, Imperial College London, United Kingdom

Presentation Overview: Show

Diet plays a critical role in human health, with growing evidence linking dietary habits to disease outcomes. However, extracting structured dietary knowledge from biomedical literature remains challenging due to the lack of dedicated relation extraction datasets. To address this gap, we introduce RECoDe, a novel relation extraction (RE) dataset designed specifically for diet, disease, and related biomedical entities.
RECoDe captures a diverse set of relation types, including a broad spectrum of positive association patterns and explicit negative examples, with over 5,000 human-annotated instances validated by up to five independent annotators. Furthermore, we benchmark various natural language processing (NLP) RE models, including BERT-based architectures and enhanced prompting techniques with locally deployed large language models (LLMs) to improve classification performance on underrepresented relation types.
The best performing model was gpt-oss-20B, a locally-deployed open-weight LLM, achieving an F1-score of 61% (macro) for multi-class classification and 89% for binary classification using a hierarchical prompting strategy with a separate reflection step built in. To demonstrate the practical utility of RECoDe, we introduce the Contextual Co-occurrence Summarisation (CoCoS) framework, which aggregates sentence-level relation extractions into document-level summaries and further integrates evidence across multiple documents. CoCoS produces effect estimates consistent with established dietary knowledge, demonstrating its validity as a general framework for systematic evidence synthesis.

Availability: The code, models, and dataset are publicly available at https://github.com/omicsNLP/RECoDe.

Wednesday 2 September
10:30-11:30
Session: Diving into Virus Host Prediction and Coral Microbiome Diversity
Coral microbiomes as reservoirs of unknown genomic and biosynthetic diversity
Confirmed Presenter: Lucas Paoli, School of Life Sciences, EPFL, Switzerland

Room: Room EF
Moderator(s): Avidan Neumann; Katja Bärenfaller


Authors List: Show

  • Fabienne Wiederkehr, Department of Biology, Institute of Microbiology and Swiss Institute of Bioinformatics, ETH Zurich, Switzerland, Switzerland
  • Lucas Paoli, School of Life Sciences, EPFL, Switzerland
  • Daniel Richter, Department of Biology, Institute of Microbiology, ETH Zurich, 8093 Zurich, Switzerland, Switzerland
  • Dora Racunica, Department of Biology, Institute of Microbiology, ETH Zurich, 8093 Zurich, Switzerland, Switzerland
  • Hans-Joachim Ruscheweyh, Department of Biology, Institute of Microbiology and Swiss Institute of Bioinformatics, ETH Zürich, Switzerland
  • Martin Sperfeld, Department of Biology, Institute of Microbiology and Swiss Institute of Bioinformatics, ETH Zurich, Switzerland, Switzerland
  • James O'Brien, Department of Biology, Institute of Microbiology and Swiss Institute of Bioinformatics, ETH Zurich, Switzerland, Switzerland
  • Samuel Miravet-Verde, Department of Biology, Institute of Microbiology and Swiss Institute of Bioinformatics, ETH Zurich, Switzerland, Switzerland
  • Alena Streiff, Department of Biology, Institute of Microbiology, ETH Zurich, 8093 Zurich, Switzerland, Switzerland
  • Jessica Ransome, Department of Biology, Institute of Microbiology, ETH Zurich, 8093 Zurich, Switzerland, Switzerland
  • Clara Chepkirui, Department of Biology, Institute of Microbiology, ETH Zurich, 8093 Zurich, Switzerland, Switzerland
  • Taylor Priest, Department of Biology, Institute of Microbiology and Swiss Institute of Bioinformatics, ETH Zurich, Switzerland, Switzerland
  • Anna Sintsova, Department of Biology, Institute of Microbiology and Swiss Institute of Bioinformatics, ETH Zurich, Switzerland, Switzerland
  • Guillem Salazar, Department of Biology, Institute of Microbiology and Swiss Institute of Bioinformatics, ETH Zürich, Switzerland
  • Kalia Bistolas, Department of Microbiology, Oregon State University, Corvallis, OR, 97330, United States
  • Teresa Sawyer, Electron Microscopy Facility, Oregon State University, Corvallis, OR, 97330, United States
  • Karine Labadie, Genoscope, Institut François Jacob, CEA, CNRS, Université Evry, Université Paris-Saclay, 91057 Evry, France
  • Kim-Isabelle Mayer, Department of Biology, University of Konstanz, 78457 Konstanz, Germany
  • Aude Perdereau, Genoscope, Institut François Jacob, CEA, CNRS, Université Evry, Université Paris-Saclay, 91057 Evry, France
  • Maggie Reddy, Marine Biodiscovery Laboratory, School of Chemistry and Ryan Institute, National University of Ireland, Galway, Ireland
  • Clémentine Moulin, Fondation Tara Océan, Base Tara, 8 rue de Prague, 75012 Paris, France
  • Emilie Boissin, PSL Research University: EPHE-UPVD-CNRS, USR 3278 CRIOBE, Laboratoire d’Excellence CORAIL, Université de Perpignan, 52 Avenue Paul Alduy, 66860, Perpignan, France
  • Guillaume Bourdin, School of Marine Sciences, University of Maine, Orono, ME, USA
  • Juliette Cailliau, Fondation Tara Océan, Base Tara, Paris, France
  • Guillaume Iwankow, PSL Research University: EPHE-UPVD-CNRS, UAR 3278 CRIOBE, Laboratoire d’Excellence CORAIL, Université de Perpignan, Perpignan, France
  • Julie Poulain, Génomique Métabolique, Genoscope, Institut François Jacob, CEA, CNRS, Université Evry, Université Paris-Saclay, Evry; Research Federation for the Study of Global Ocean Systems Ecology and Evolution, FR2022/Tara GOSEE, Paris, France
  • Sarah Romac, Sorbonne Université, CNRS, Station Biologique de Roscoff, AD2M, UMR 7144, ECOMAP, Roscoff, France
  • Serge Planes, PSL Research University: EPHE-UPVD-CNRS, UAR 3278 CRIOBE, Laboratoire d’Excellence CORAIL, Université de Perpignan, Perpignan; Research Federation for the Study of Global Ocean Systems Ecology and Evolution, FR2022/Tara GOSEE, Paris, France
  • Denis Allemand, Laboratoire International Associé Université Côte d’Azur-Centre Scientifique de Monaco (LIA ROPSE), Monaco; Centre Scientifique de Monaco, Monaco
  • Sylvain Agostini, Shimoda Marine Research Center, University of Tsukuba, Shizuoka, Japan
  • Chris Bowler, Institut de Biologie de l’École Normale Supérieure (IBENS), École Normale Supérieure, CNRS, INSERM, Université PSL, Paris, France
  • Eric Douville, Laboratoire des Sciences du Climat et de l’Environnement (LSCE), CEA, CNRS, UVSQ, Université Paris-Saclay, Gif-sur-Yvette, France
  • Didier Forcioli, Laboratoire International Associé Université Côte d’Azur-Centre Scientifique de Monaco Principality of Monaco, France
  • Pierre Galand, Sorbonne Université, CNRS, LECOB, Observatoire Océanologique de Banyuls, France, France
  • Fabien Lombard, Sorbonne Université, Institut de la Mer de Villefranche, Laboratoire d'Océanographie de Villefranche, France, France
  • Pedro Oliveira, Génomique Métabolique, Genoscope, Institut François Jacob, CEA, CNRS, Université Evry, Université Paris-Saclay, France, France
  • Olivier Thomas, School of Biological and Chemical Sciences, Ryan Institute, University of Galway, Ireland, Ireland
  • Rebecca Vega Thurber, Department of Microbiology, Oregon State University, Corvallis, OR 97330, USA, United States
  • Romain Trouble, Fondation Tara Océan, Base Tara, 8 Rue de Prague, 75012 Paris, France, France
  • Christian Voolstra, Department of Biology, University of Konstanz, 78457 Konstanz, Germany, Germany
  • Patrick Wincker, Metabolic Genomics, Genoscope, Institut de Biologie François Jacob, CEA, CNRS, Evry,, France
  • Maren Ziegler, Department of Holobiont Biology, Justus Liebig University Giessen, 35392 Giessen, Germany, Germany
  • Joern Piel, Department of Biology, Institute of Microbiology, ETH Zurich, Switzerland
  • Shinichi Sunagawa, Department of Biology, Institute of Microbiology and Swiss Institute of Bioinformatics, ETH Zurich, Switzerland

Presentation Overview: Show

Coral reefs are marine biodiversity hotspots that provide a wide range of ecosystem services. They are reservoirs of bioactive metabolites, many produced by microorganisms associated with reef invertebrate hosts. However, for the keystone species of coral reefs—the reef-building corals—we still lack a systematic assessment of their microbially encoded biosynthetic potential and the molecular resources at stake due to the alarming decline in reef biodiversity. Here we analysed microbial genomes reconstructed from 820 reef-building coral samples of three representative coral genera collected at 99 reefs across 32 islands throughout the Pacific Ocean (Tara Pacific expedition). By contextualizing our analyses with the microbiomes of other reef species, we found that only 10% of the 4,224 microbial species and less than 1% of the 645 species exclusively identified in Tara Pacific samples had genomic information available. Furthermore, the biosynthetic potential of reef-building coral microbiomes rivalled or surpassed that of traditional natural product sources such as sponges. Among the biosynthetically rich bacteria in the reef microbiome, we identified new groups of Acidobacteriota that encode previously unknown enzymology, in turn opening promising avenues for functional protein engineering. Together, this study underscores the importance of conserving coral reefs as vital reservoirs of molecular diversity.

Proceedings Presentation: GiantHost: A Domain-Adaptive and Uncertainty-Aware Framework for Giant Virus Host Prediction
Confirmed Presenter: Fuchuan Qu, City University of Hong Kong, Hong Kong

Room: Room EF
Moderator(s): Avidan Neumann; Katja Bärenfaller


Authors List: Show

  • Fuchuan Qu, City University of Hong Kong, Hong Kong
  • Guowei Chen, City University of Hong Kong, Hong Kong
  • Yanni Sun, Department of Electrical Engineering, City Univerisity of Hong Kong, Hong Kong

Presentation Overview: Show

Motivation: Nucleocytoplasmic large DNA viruses (NCLDVs) play crucial roles in global ecosystems. Although metagenomics has vastly accelerated the discovery of novel NCLDVs, predicting their hosts from fragmented contigs remains a critical bottleneck, with no dedicated end-to-end computational tools currently available. Addressing this gap requires overcoming three fundamental challenges: the extreme scarcity of labeled reference genomes, the severe domain shift between laboratory isolates and diverse environmental metagenomes, and the inability of traditional deterministic models to quantify prediction uncertainty—a crucial requirement for reliable ecological profiling where novel, divergent viruses are prevalent.
Results: We present GiantHost, the first NCLDV host prediction tool with domain adaptation and uncertainlty awareness. GiantHost employs a dual-tower neural network to integrate dense genome traits and sparse GVOG profiles, allowing better integration of heterogeneous features. To overcome label scarcity and domain shift, we leverage 1,400 environmental viral genomes (GVMAGs) via semi-supervised multi-task learning and Domain Adversarial Neural Networks (DANN), effectively bridging the distributional gap between RefSeq and environmental data. Additionally, GiantHost incorporates Conformal Prediction (CP) to output statistically guaranteed prediction sets rather than overconfident single labels. Evaluated under rigorous genome-level cross-validation, GiantHost demonstrates robust predictive power. Applied to the Tara Ocean dataset, GiantHost successfully captured the vertical stratification of NCLDV hosts—revealing a depth-dependent decline of phytoplankton-infecting viruses and a relative enrichment of Amoebozoa-infecting viruses in the mesopelagic zone.

Proceedings Presentation: ViralQC: A Tool for Assessing Completeness and Contamination of Predicted Viral Contigs
Confirmed Presenter: Cheng Peng, City University of Hong Kong, Hong Kong

Room: Room EF
Moderator(s): Avidan Neumann; Katja Bärenfaller


Authors List: Show

  • Cheng Peng, City University of Hong Kong, Hong Kong
  • Jiayu Shang, The Chinese University of Hong Kong, Hong Kong
  • Jiaojiao Guan, City Univesity of HongKong, China
  • Yanni Sun, City University of Hong Kong, Hong Kong

Presentation Overview: Show

Viruses represent the most abundant biological entities on Earth, playing vital roles in diverse ecosystems. Cataloging viruses across various environments is essential for understanding their properties and functions. Metagenomic sequencing has emerged as the most comprehensive method for virus discovery. However, distinguishing viral sequences from the vast background of microbial organisms in metagenomic data remains a significant challenge. Existing tools experience varying degrees of false positive rates due to noise in sequencing and assembly, and the integration of proviruses into microbial genomes. This highlights the urgent need for an accurate and efficient method to evaluate the quality of viral contigs.
To address these challenges, we introduce ViralQC, a tool designed to assess the quality of viral contigs or bins. ViralQC identifies microbial contamination within putative viral sequences using a multi-modal framework powered by DNA and protein foundation models and estimates completeness by analyzing protein organization. We evaluated ViralQC on multiple datasets and compared its performance against the state-of-the-art tool, CheckV. Leveraging both DNA and protein foundation models, ViralQC achieves higher recall on contamination detection for contigs longer than 10 kbp while maintaining comparable precision. Additionally, ViralQC delivers more accurate estimation on contigs with completeness>50%.