View posters by category

Scroll down to view Results

Session A: Monday 31 August 12:00-13:30
Session B: Tuesday 1 September 16:15-17:45
Session C: Wednesday 2 September 11:30-13:00

Results

C-B.01: Metabolic modeling reveals microbial metabolites associated with infant temperament traits
Track: Biodiversity, sustainability, envirobioinformatic
  • Abhijit Paul, Turku Bioscience Centre, University of Turku and Ã…bo Akademi University, 20520 Turku, Finland, Finland
  • Anna-Katariina Aatsinki, FinnBrain Birth Cohort Study, Department of Clinical Medicine, University of Turku and Turku University Hospital, Turku, Finland
  • Matilda KrÃ¥kström, Turku Bioscience Centre, University of Turku and Ã…bo Akademi University, 20520 Turku, Finland, Finland
  • Minna Lukkarinen, FinnBrain Birth Cohort Study, Department of Clinical Medicine, University of Turku and Turku University Hospital, Turku, Finland
  • Eveliina Munukka, Centre for Population Health Research, University of Turku and Turku University Hospital, Turku, Finland, Finland
  • Venla Huovinen, FinnBrain Birth Cohort Study, Department of Clinical Medicine, University of Turku and Turku University Hospital, Turku, Finland
  • Saara Nolvi, FinnBrain Birth Cohort Study, Department of Clinical Medicine, University of Turku and Turku University Hospital, Turku, Finland
  • Eeva-Leena Kataja, FinnBrain Birth Cohort Study, Department of Clinical Medicine, University of Turku and Turku University Hospital, Turku, Finland
  • Riikka Korja, FinnBrain Birth Cohort Study, Department of Clinical Medicine, University of Turku and Turku University Hospital, Turku, Finland
  • Nitin Bayal, Department of Computing, University of Turku, 20014 Turku, Finland, Finland
  • Hasse Karlsson, FinnBrain Birth Cohort Study, Department of Clinical Medicine, University of Turku and Turku University Hospital, Turku, Finland
  • Leo Lahti, Department of Computing, University of Turku, 20014 Turku, Finland, Finland
  • Santosh Lamichhane, Research Center for Infections and Immunity, Institute of Biomedicine, University of Turku, Turku, Finland, Finland
  • Linnea Karlsson, FinnBrain Birth Cohort Study, Department of Clinical Medicine, University of Turku and Turku University Hospital, Turku, Finland
  • Alex M. Dickens, Turku Bioscience Centre, University of Turku and Ã…bo Akademi University, 20520 Turku, Finland, Finland
  • Matej OreÅ¡ič, Turku Bioscience Centre, University of Turku and Ã…bo Akademi University, 20520 Turku, Finland, Finland


Presentation Overview: Show

Background: Early temperament traits are important predictors of later socioemotional development and mental health outcomes. While infant gut microbiota has been linked to early temperament traits, the underlying microbial functional mechanisms remain largely unexplored. Here we examine whether microbial metabolites at 2.5 months of age are associated with positive and negative reactivity at 6 months, as they may serve as early intermediate phenotypes of later child psychiatric disorders.

Methods: Stool samples from 2.5-month-old infants (n = 207) in the FinnBrain Birth Cohort underwent shotgun metagenomics sequencing. Species-level taxonomic profiles were generated with nf-core/taxprofiler and MetaPhlAn4. Individualized microbial community metabolic models were constructed using Microbiome Modelling Toolbox to estimate metabolite secretion potentials. We examined associations between taxonomic composition, predicted metabolite secretion potentials, and mother-reported Infant Behavior Questionnaire–Revised (IBQ-R) measures of positive and negative reactivity at 6 months of age. Model predictions were compared with fecal metabolomics data.

Results: Community metabolic modeling revealed that bile acids, short-chain fatty acids, amino acids, vitamins, and glycans were associated with temperament traits. Taurine- and glycine-conjugated bile acids positively associated with negative emotionality, while microbially conjugated bile acids with other amino acids positively associated with surgency (positive emotionality). Short-chain fatty acids, particularly butyrate and isobutyrate, were negatively associated with fear reactivity.

Conclusions: By integrating shotgun metagenomics with community metabolic modeling, we identified associations between microbial bile acid and short-chain fatty acid metabolism and infant temperament traits, supported by stool metabolomics. Beyond taxonomic associations, these findings highlight potential microbial functional contributions to early behavioral phenotypes.

C-B.02: Deep Orthogroups: Extending Orthogroups beyond the root of the species tree.
Track: Biodiversity, sustainability, envirobioinformatic
  • Jonathan Holmes, University of Oxford, United Kingdom
  • Steven Kelly, University of Oxford, United Kingdom


Presentation Overview: Show

An orthogroup is the central unit of comparative genomics, which consists of a set of genes descended from a single gene in the last common ancestor of a set of species. Programs such as OrthoFinder, FastOMA and Sonic Paranoid aim to identify these orthogroups and infer orthologs. Orthogroups are often limited to the root of a species tree which ensures that ancient duplications beyond the root are separated. However, further evolutionary history exists between genes of different orthogroups beyond the root of the species tree, which we term deep orthogroups. Here, we present a workflow for mapping these connections to identify deep orthogroups by characterising the evolutionary history of distantly related orthogroups and identifying duplications beyond the root of the species tree. Our workflow first utilises the orthogroups initially derived by OrthoFinder, using the BLAST scores between genes computed by OrthoFinder to linked distantly related orthogroups, we then perform the additional steps of re-clustering existing orthogroups into new clusters then using a modified DendroBLAST distance for phylogenetic inference. We show through the use of simulation studies that we can not only recapture related orthogroups into deep orthogroups, but we also accurately re-create the phylogenetic history of these orthogroups, allowing deep duplication events to be mapped. We then apply this to a range of Eukaryotes across the tree of life in order to infer ancient gene duplications.

C-B.03: Genome-Centric Characterisation of Limnochordia in Biogas Microbiomes: From Phylogenetic Diversity to Clade-Specific Functions
Track: Biodiversity, sustainability, envirobioinformatic
  • Zihan Dai, Institute of Bio- and Geosciences IBG-5: Computational Metagenomics, Forschungszentrum Jülich, Germany
  • Irena Maus, Institute of Bio- and Geosciences IBG-5: Computational Metagenomics, Forschungszentrum Jülich, Germany
  • Benedikt Osterholz, Institute of Bio- and Geosciences IBG-5: Computational Metagenomics, Forschungszentrum Jülich, Germany
  • Tom Tubbesing, Computational Metagenomics Group, Faculty of Technology and Center for Biotechnology (CeBiTec), Bielefeld University, Germany
  • Liren Huang, Computational Metagenomics Group, Faculty of Technology and Center for Biotechnology (CeBiTec), Bielefeld University, Germany
  • Sebastian Jünemann, Institute of Bio- and Geosciences IBG-5: Computational Metagenomics, Forschungszentrum Jülich, Germany
  • Andreas Schlüter, Computational Metagenomics Group, Faculty of Technology and Center for Biotechnology (CeBiTec), Bielefeld University, Germany
  • Alexander Sczyrba, Institute of Bio- and Geosciences IBG-5: Computational Metagenomics, Forschungszentrum Jülich, Germany


Presentation Overview: Show

Anaerobic digestion (AD) of biomass for methane production is a key component of renewable energy systems, with process stability governed by complex, poorly characterised microbial communities. The class Limnochordia has emerged as a consistently abundant and resilient member of biogas microbiomes, yet its genomic and functional diversity remains unresolved. Here, we present a comprehensive, genome-centric characterisation of Limnochordia across 5 biogas plants comprising 10 reactors, combining in-house metagenomic data with publicly available biogas datasets.
Metagenome assembly and binning yielded 77 Limnochordia MAGs, 24 reconstructed from in-house samples and 53 from public data, of which 69 passed quality thresholds (completeness ≥80%, contamination ≤10%). These high-quality MAGs span 5 families, 18 genera, and 36 species. Family DTU010 dominates the dataset (50 MAGs, 9 genera, 23 species), followed by DTU012 (15 MAGs, 5 genera, 9 species); three additional families are represented by 1 to 2 MAGs each. Phylogenetic analysis resolved clade structure across all 69 MAGs. KEGG-based functional profiling indicates that functional repertoires are broadly partitioned along phylogenetic lines, with further within-family heterogeneity revealed at higher pathway completeness thresholds. Detailed differential pathway analysis across clades is ongoing.
Metaproteomic data will be mapped against these MAG references to link genomic potential to active expression under varying reactor conditions. Collectively, this work provides a resolved phylogenetic and functional framework for Limnochordia in engineered anaerobic systems, with direct implications for microbiome management and inocula development in biogas production.

C-B.04: Unmasking the Growing Challenge and Hidden Impact of Sampling Bias in Prokaryotic Genome Collections
Track: Biodiversity, sustainability, envirobioinformatic
  • Hannah Goetsch, JLU Giessen, Germany
  • Franz Baumdicker, JLU Giessen, Germany


Presentation Overview: Show

The rapid growth of fully sequenced prokaryotic genomes improved our ability to study prokaryotic diversity, adaptation, and the spread of traits such as antibiotic resistance. However, many genome collections have been assembled through heterogeneous study designs rather than systematic sampling, often leading to an overrepresentation of closely related individuals of particular clinical, ecological, or epidemiological interest. Such sampling bias reduces the effective information, biases downstream analyses, and limits the ability to capture true prokaryotic diversity.
We investigate the extent and consequences of sampling bias in the NCBI prokaryotic genome database. We show that oversampling has become widespread and increasingly severe, with substantial effects on downstream analyses. To address this problem, we introduce PhyloThin, a method that detects and removes sampling bias from prokaryotic genome collections using phylogenetic relationships within a coalescent framework. Applying PhyloThin to 402 prokaryotic species reveals pervasive oversampling trends, especially in recent years. Across species, the effective sample size is on average nearly half of the total number of genomes. Importantly, this bias distorts gene frequency estimates and leads to underestimation of prokaryotic diversity and adaptive potential.
Our results provide a robust framework for identifying oversampled genomes and improving the accuracy of inferences from large-scale prokaryotic genome datasets. In particular, reducing sampling bias enables more reliable assessments of prokaryotic diversity and evolutionary potential. More broadly, exposing these biases can help prioritize species and lineages for future sequencing efforts to better capture the full extent of prokaryotic biodiversity.

C-B.05: FlaPro: a computational framework for TLR5-phenotype-aware flagellome profiling across human gut metagenomes and metatranscriptomes
Track: Biodiversity, sustainability, envirobioinformatic
  • Anna Bogdanova, Department of Microbiome Science, Max Planck Institute for Biology Tübingen, Germany
  • Andrea Borbón-Garcí­a, Department of Microbiome Science, Max Planck Institute for Biology Tübingen, Germany
  • Ruth Ley, Department of Microbiome Science, Max Planck Institute for Biology Tübingen, Germany
  • Alexander Tyakht, Department of Microbiome Science, Max Planck Institute for Biology Tübingen, Germany


Presentation Overview: Show

Flagellin, the structural protein of bacterial flagella, activates the innate immune receptor Toll-like receptor 5 (TLR5). However, different flagellins vary widely in their ability to stimulate TLR5 - from highly immunostimulatory variants to ""silent"" ones that bind TLR5 without triggering signaling. This suggests that the composition of an individual's flagellin repertoire, or flagellome, is a potential distinct contributor to host-microbiome interactions and inflammation. Despite growing evidence that flagellins contribute to diseases such as inflammatory bowel disease (IBD), computational methods to resolve both flagellome composition and its TLR5-stimulatory potential from microbiome sequencing data have been lacking.
We developed FlaPro, a Snakemake-based pipeline for quantification and functional annotation of human gut flagellomes from metagenomic (MGX) and metatranscriptomic (MTX) data. FlaPro combines ShortBRED marker-based quantification against a curated reference of non-redundant flagellin clusters from human gut microbiome with a Random Forest classifier trained on experimentally characterized flagellins to predict per-flagellin TLR5 phenotype.
We apply FlaPro in a cross-cohort meta-analysis of public gut MGX/MTX datasets spanning diverse diseases, generating harmonized flagellome profiles and identifying both shared and condition-specific TLR5-phenotype signatures. For an IBD dataset, compared to healthy controls we discovered reduced flagellome diversity and a lower silent-to-stimulatory ratio in both Crohn's disease and ulcerative colitis consistently across MGX and MTX – a disease-associated shift toward a more stimulatory flagellome. Similarly, condition-specific alterations have been identified for other diseases. Together, these analyses position the flagellome as a quantifiable, functionally interpretable microbiome feature for microbiome-wide association studies in health and disease.

C-B.06: Fast, flexible gene cluster family delineation with IGUA
Track: Biodiversity, sustainability, envirobioinformatic
  • Martin Larralde, Leiden University Center for Infectious Diseases, Leiden University Medical Center, Leiden, Netherlands
  • Josefin Blom, Department of Clinical Microbiology, SciLifeLab, UmeÃ¥ University, UmeÃ¥, Sweden
  • Hadrien Gourlé, Department of Clinical Microbiology, SciLifeLab, UmeÃ¥ University, UmeÃ¥, Sweden
  • Hale-Seda Radoykova, Global Health Institute, School of Life Science, EPFL, Lausanne, Switzerland
  • Lucas Paoli, Global Health Institute, School of Life Science, EPFL, Lausanne, Switzerland
  • Laura Carroll, Department of Clinical Microbiology, SciLifeLab, UmeÃ¥ University, UmeÃ¥, Sweden
  • Georg Zeller, Leiden University Center for Infectious Diseases, Leiden University Medical Center, Leiden, Netherlands


Presentation Overview: Show

Prokaryotic genomes harbor a variety of functional elements encoded as contiguous multi-gene clusters. To identify homologous gene clusters, computational tools that account for multi-gene architectures are needed to group gene clusters into Gene Cluster Families (GCFs). However, existing GCF delineation methods do not scale well to large datasets and are often limited to certain subclasses of gene clusters.

Here, we present IGUA (Iterative Gene clUster Analysis; https://github.com/zellerlab/IGUA), a scalable, flexible GCF delineation method for genomic segments with multi-gene architectures.

IGUA uses Mmseqs2 clustering of both nucleotide and protein sequences with different parameters tuned for optimal de-replication. The final gene cluster families are obtained with hierarchical clustering over the protein content of the gene clusters, which captures homology despite organizational changes at the nucleotide scale. To that end, we propose an efficient distance metric for measuring pairwise distances between gene clusters without a priori reference features.

On a biosynthetic gene cluster (BGC) clustering task, IGUA is >18x and 5x faster than the state-of-the-art (BiG-SCAPE 2.0 and BiG-SLiCE 2.0, respectively), without sacrificing accuracy. To highlight its scalability, we use IGUA to cluster >2.8 million BGCs from ~1 million prokaryotic genomes in <18 hours (n = 2,829,071 BGCs to 56,960 GCFs). Demonstrating its utility beyond BGC clustering, we use IGUA to cluster five additional prokaryotic gene cluster types (secretion systems, prophages, genomic islands, phage-plasmids, and carbohydrate active enzyme [CAZyme] gene clusters).

Overall, IGUA represents a versatile GCF delineation tool with unmatched computational efficiency and flexibility, enabling (meta)genomic mining applications at unprecedented scales.

C-B.07: MetagenomeWatch: Web-based monitoring and search of public metagenomic data for pathogen surveillance
Track: Biodiversity, sustainability, envirobioinformatic
  • Matthew Huska, Robert Koch Institute, Germany
  • Daniel Desiro, Robert Koch Institute, Germany
  • Martin Hoelzer, Robert Koch Institute, Germany


Presentation Overview: Show

The rapid growth of next-generation sequencing has led to an unprecedented expansion of publicly available metagenomic data. The Sequence Read Archive (SRA) now contains over nine million sequencing datasets spanning diverse environments, including human, animal, and wastewater samples. This resource provides substantial potential for improving our understanding of microbial communities and their relevance to public health. However, effective search and analysis at this scale remain challenging due to the computational and storage demands of petabyte-scale data.

Recent advances in sequence sketching enable highly compressed representations that retain sufficient information to probabilistically detect the presence of sequences within large datasets, making large-scale search of metagenomic repositories computationally feasible.

Here, we present MetagenomeWatch, a web-based system built on sourmash branchwater for large-scale sequence sketching and comparison. The platform enables efficient querying of public metagenomic data for viral, bacterial, and eukaryotic pathogens. Leveraging precomputed sketches and branchwater-based search, typical queries across the full SRA metagenomics collection can be completed in under five minutes while maintaining sensitivity for sequence detection. In addition, the system provides automated monitoring workflows that track newly processed datasets and report matches to user-defined queries.

This approach enables scalable, continuous surveillance of publicly available metagenomic data and facilitates early detection and tracking of pathogens, with potential applications in outbreak monitoring and public health response.

MetagenomeWatch is open-source software (AGPL-3.0) and can be found at: https://github.com/rki-mf1/mgwatch

C-B.08: Comparative methods for RNA Secondary Structure Prediction in RNA Viruses
Track: Biodiversity, sustainability, envirobioinformatic
  • Charlotte Tumescheit, IDIAP, SIB, Switzerland
  • Katherine Brown, Cambridge University, United Kingdom
  • Janna Hastings, IDIAP, SIB, Switzerland
  • Andrew E Firth, Cambridge University, United Kingdom


Presentation Overview: Show

RNA secondary structures play many different roles in the life cycle of RNA viruses, for example in directing translational control and genome replication. However, predicting these structures can be challenging, especially when it comes to pseudoknots, long-range interactions, and mutually exclusive interactions.

Here, we use a multiple sequence alignment based approach to mitigate these challenges by specifically looking for all possibly functionally relevant conserved short-range and long-range interactions, thereby allowing for the detection of pseudoknots and mutually exclusive structures.
To allow functional structures to be identified, we incorporated an offset parameter, allowing imperfect and more divergent alignments; a phylogenetic weighting scheme, to balance diversity with specificity; and a free energy calculation.
For benchmarking purposes, we curated a dataset of verified secondary structures in RNA viruses, showing that our tool successfully predicts known structures (sensitivity of 77.78%) besides suggesting novel functionally relevant structures.

Next, we improve existing transformer models for RNA analysis by incorporating knowledge-injection strategies as well as fine-tuning on our curated database to further improve secondary structure prediction for RNA viruses.

While RNA structure prediction is still struggling to match the performance of protein structure predictions, we present different strategies to improve the predictions of RNA structures, using methods based on comparative genomics, with and without the incorporation of machine learning. Specifically focusing on RNA viruses and using curated databases already shows the usefulness of our approaches, which we aim to expand to other RNA use cases in the future.

C-B.09: ABRomics: a platform for antibiotic resistance research and public health using an integrated One Health approach
Track: Biodiversity, sustainability, envirobioinformatic
  • Brieuc Quemeneur, CNRS IFB-core & Nantes Université, CNRS, INSERM, l’institut du thorax, France
  • Alban Gaignard, CNRS, IFB-core & Nantes Université, CNRS, INSERM, l’institut du thorax, France
  • Samuel Chaffron, Nantes Université, École Centrale Nantes, CNRS, LS2N, UMR 6004, F-44000 Nantes, France
  • Audrey Bihouée, Nantes Université, CNRS, INSERM, l’institut du thorax, F-44000 Nantes, France
  • Abromics Consortium, CNRS IFB-core, UAR 3601 - Villejuif, France
  • Gildas Le Corguillé, CNRS, IFB-core & ABiMS, Station Biologique-Roscoff, France
  • Etienne Ruppé, Univ. Paris Cité and Univ.Sorbonne Paris Nord, Inserm, IAME - Paris, France
  • Nadia Goué, CNRS IFB-core & AuBi platform, Université Clermont-Auvergne, France
  • Bérénice Batut, CNRS IFB-core & AuBi platform, Université Clermont-Auvergne, France
  • Pierre Marin, CNRS IFB-core & AuBi platform, Université Clermont-Auvergne, France
  • Claudine Médigue, CNRS IFB-core & CEA, Genoscope, LABGeM, France
  • Hugo Lefeuvre, Nantes Université, CNRS, INSERM, l’institut du thorax, France
  • Cléa Siguret, CNRS IFB-core & AuBi platform Université Clermont-Auvergne, France
  • Thomas Mignon, CNRS FB-core & ABiMS, Station Biologique - Roscoff, France
  • Amanda Dieuaide, CNRS IFB-core & Institut Pasteur, Bioinformatics and Biostatistics Hub, France
  • Raphaël Tackx, CNRS IFB-core & Institut Pasteur, Bioinformatics and Biostatistics Hub, France
  • Julie Lao, CNRS, Institut Français de Bioinformatique, IFB-core, UAR 3601 - Villejuif, France
  • Philippe Glaser, Institut Pasteur, Unité EERA, CNRS UMR604, Paris, France
  • Fabien Mareuil, Institut Pasteur, Université Paris Cité, Bioinformatics and Biostatistics Hub, Paris, France


Presentation Overview: Show

Antibiotic resistance is a major public health concern. Whole Genome Sequencing (WGS) is essential for characterizing the diffusion and the monitoring of Multidrug resistant bacteria (MDRB) within and between the human, animal and environmental sectors. In this context, the French Priority Plan on Antimicrobial Resistance has funded the development of an online platform dedicated to the surveillance and research of antibiotic resistance within a One Health framework (www.abromics.fr/). The ABRomics platform is hosted at the French Institute of Bioinformatics (www.ifb-elixir.fr/), which has extensive data analysis and storage capacity.

The web service ABRomics-analysis (analysis.abromics.fr/) offers a user-friendly interface for uploading and managing genomic samples as well as for running, through usegalaxy.fr, a standardized workflow comprising: quality control, species identification, MLST (MultiLocus Sequence Typing), genome assembly and annotation, cgMLST (core genome MLST) using standardized nomenclatures, plasmid typing, and the detection of antibiotic resistance and virulence genes. The ABRomics database aims to facilitate data exploration and foster collaboration across sectors and communities. To this end, it provides tools for exploring freely accessible data on all submitted isolates and for connecting scientists who share related strains.

Ongoing developments of the ABRomics platform include data brokering of genomic samples to the European Nucleic Archive (ENA), standardized workflows for metagenomics analysis, and tools for analyzing (meta)pangenomes. Importantly, data management procedures comply with the FAIR principle (Findable, Accessible, Interoperable and Reusable) thuswill enablinge retrospective epidemiological studies to be conducted.

C-B.10: Phylogenetic ordering and batching for better compression of million-genome bacterial collections
Track: Biodiversity, sustainability, envirobioinformatic
  • Tam Truong, INRIA Rennes, France, France
  • Dominique Lavenier, CNRS, Rennes, France,, France
  • Pierre Peterlongo, INRIA Rennes, France, France
  • Karel Břinda, INRIA Rennes, France, France


Presentation Overview: Show

Modern bacterial genome collections such as AllTheBacteria (ATB) and GTDB contain millions of isolate genomes and metagenome-assembled genomes (MAGs). Efficient compression of such collections requires batching and ordering of genomes that co-localizes shared redundancy before low-level compression. While phylogenetic compression enables such an ordering efficiently using evolutionary history, its state-of-the-art implementation MiniPhy, used as the core method for ATB, is limited by fixed-size batching based on metadata proxies like species and accessions. Consequently, compression degrades on collections with oversampled species, high taxonomic diversity, and mixed isolate and MAG content.

We present PhyloPack, a scalable method for global phylogenetic ordering and batching via skeleton phylogenies. We formalize both as an optimization problem aimed at maximizing compression efficiency under fixed computational constraints. PhyloPack employs a two-step heuristic: first, a skeleton phylogeny is inferred from a subsampled set of genomes; second, the remaining sequences are placed within the skeleton tree using sketch-based nearest neighbor identification. On ATBv0.3, PhyloPack with MBGC2 reduces the major-species subset (2.2 millions genomes) from 57 GB (XZ) to 15 GB, and the whole collection (2.4 millions genomes) from 103GB (XZ) to 41 GB (60% reduction to the current ATB distribution). Combined with XZ and AGC, we observe a consistent 40% improvement over MiniPhy. On GTDB (r226), using MBGC2, PhyloPack compresses the isolate to 35GB and MAG subsets to 70GB, respectively (totaled to 8% of the size of the official gzip files). Overall, PhyloPack makes phylogenetic compression practical for heterogeneous genome collections, scalable beyond millions of genomes.

C-B.11: A research-oriented framework for viral discovery and interpretation in metagenomic data
Track: Biodiversity, sustainability, envirobioinformatic
  • Amanj Bajalan, Department of Microbiology, Tumor and Cell Biology, Karolinska Institutet, Stockholm, Sweden, Sweden
  • Björn Andersson, Department of Cell and Molecular Biology (CMB), Karolinska Institutet, Stockholm, Sweden, Sweden
  • Tobias Allander, Department of Microbiology, Tumor and Cell Biology, Karolinska Institutet, Stockholm, Sweden, Sweden


Presentation Overview: Show

Viruses from many families are important human pathogens, and advances in metagenomic sequencing have enabled broad and unbiased virus detection. At the same time, there is a need for reproducible workflows that support both detection and interpretation of ambiguous and potentially novel viral findings. We developed a scalable and modular viral discovery framework for reproducible metagenomic analysis. The workflow is implemented in Nextflow and integrates read-based analysis, assembly creation, assembly-based classification, assembly statistics, machine learning-based post-classification, and interactive HTML reporting. In addition to automated processing, it includes inspection tools for interpretation of ambiguous and unclassified sequences, including full lineage reporting, metadata, filterable human-virus tagging, built-in support for querying NCBI resources, a built-in ORF viewer for unknown viral hits, and a viral likelihood scoring system for prioritization of uncertain findings. This makes the workflow useful for both virus detection and research-oriented interpretation of unclear results. The workflow was applied to blood plasma, CSF, and fecal samples from Swedish clinical data, together with publicly available oral microbiome data generated for bacterial profiling rather than viral discovery. The analysis identified both RNA and DNA viruses from multiple viral families across diverse sample types, and supported prioritization of uncertain or potentially novel viral sequences. In the oral microbiome dataset, the workflow recovered and supported inspection of a contig matching HIV despite the dataset's original bacterial focus, highlighting utility for exploratory cross-domain metagenomic analysis. The workflow generates portable HTML reports for rapid overview and downstream analysis, and is being applied in ongoing research studies.

C-B.12: Characterization of Microbial Succession During Spontaneous Fermentation Using Long-Read Amplicon Sequencing
Track: Biodiversity, sustainability, envirobioinformatic
  • Oliver Scharinger, Hochschule Campus Wien, Austria
  • Lukas Fürnwein, Hochschule Campus Wien, Austria
  • Alexandra Graf, Hochschule Campus Wien, Austria


Presentation Overview: Show

Regional distinctiveness of wines is strongly shaped by the composition and activity of site-specific microorganisms, which is known as the microbial terroir concept. For this reason, spontaneously fermented wines are gaining popularity, whereas wines relying on inoculated fermentations with commercial yeast strains tend to standardized aroma profiles and reduced sensorical complexity. Therefore the microbial dynamcis during a spontaneous fermentation of an Austrian Grüner Veltliner from a winery in Gols (Burgenland), was analyzed using long-read amplicon sequencing. Samples of unripe grapes, mash and fermenting must (every three days) were collected, DNA was extracted and then sequenced on an Oxford Nanopore platform as 16S- and ITS-amplicons. The results show that the washed grape samples and the mash sample are
strongly affected by co-amplification of chloroplast and mitochondrial Vitis vinifera rRNA genes.Over the course of fermentation, Tatumella/Pantoea (29.78-59.02%) were the most abundant genera. At fermentation start, lactic acid bacteria of the genera Lactococcus (32.63%) and Lactobacillus/Apilactobacillus (27.43%) showed high relative abundances and decreased over time,whereas Lactiplantibacillus increased with the course of fermentation. In one sample, Gluconobacter was detected at a relative abundance of 9%. The fungal community showed a transition from Hanseniaspora at the beginning of fermentation towards a Saccharomyces-dominated community from mid to end of fermentation. The study demonstrates the potential of long-read amplicon sequencing for determining the microbial composition of spontaneously fermented wine samples, but also identifies plant-associated DNA contaminations as a technical limitation that should be improved in future studies by optimized host-depletion strategies.

C-B.13: Plant species pool size predicts arbuscular mycorrhizal fungal diversity and biomass better than local richness
Track: Biodiversity, sustainability, envirobioinformatic
  • Bo Stevens, University of Tartu, Estonia
  • Darkdivnet Consortium, University of Tartu, Estonia
  • Martti Vasar, University of Tartu, Estonia
  • Koit Herodes, University of Tartu, Estonia
  • Ayesh Wipulasena, University of Tartu, Estonia
  • Tanel Vahter, University of Tartu, Estonia
  • Martin Zobel, University of Tartu, Estonia
  • Meelis Pärtel, University of Tartu, Estonia


Presentation Overview: Show

Arbuscular mycorrhizal (AM) fungi are globally important plant symbionts, yet the relative contributions of environmental conditions and plant diversity metrics to AM fungal diversity and biomass remain incompletely resolved at broad spatial scales. We evaluated factors associated with AM fungal diversity (effective number of species) and two fungal biomass proxies, NLFA 16:1ω5 and PLFA 16:1ω5, across 99 globally distributed DarkDivNet sites using Random Forest models integrating soil, climate, fungal guild, and plant diversity metrics. We specifically tested whether plant diversity metrics derived from dark-diversity theory -- species pool size (regional biodiversity potential) and community completeness (realization of that potential) -- improved prediction beyond abiotic covariates and local plant richness alone. Models incorporating plant diversity metrics consistently outperformed abiotic-only models. The best-performing models explained 18% of held-out variance for AM fungal diversity, 20% for NLFA, and 28% for PLFA. Plant species pool size was retained in the best-performing model for all three responses, either alone or together with plant community completeness. AM fungal diversity was positively associated with plant species pool size and soil pH, whereas PLFA was more strongly associated with soil carbon and precipitation. Ectomycorrhizal fungal diversity showed negative associations with AM fungal diversity and NLFA. Overall, our results indicate that broad-scale patterns of AM fungal diversity and biomass are linked more strongly to regional plant diversity potential than to local plant richness and environmental conditions.

C-B.14: ViHostFinder: A Hierarchical Multilabel Host Predictor for Viral Sequences Using DNA Language Models
Track: Biodiversity, sustainability, envirobioinformatic
  • Alfred Ferrer Florensa, Technical University of Denmark, Denmark
  • Frank Møller Aarestrup, Technical University of Denmark, Denmark
  • Henrik Nielsen, Technical University of Denmark, Denmark
  • Philip Thomas Conradsen Clausen, Technical University of Denmark, Denmark


Presentation Overview: Show

Metagenomic sequencing has dramatically expanded our view of the global virome, yet our ability to assess the biological relevance of novel viral sequences remains severely limited. Current approaches rely almost exclusively on sequence similarity to known viruses — a strategy that fails systematically for the vast, uncharacterized fraction of environmental viral diversity. Predicting which hosts a virus can infect is a critical step toward evaluating zoonotic risk, understanding transmission dynamics, and enabling actionable surveillance.
We present ViHostFinder, a universal viral host predictor built on a hierarchical multilabel classification framework using DNA language models. By operating directly on raw genomic sequences, ViHostFinder bypasses protein annotation — a step that remains unreliable for highly divergent viruses — making it broadly applicable across the virome. The model natively supports multilabel prediction, capturing the capacity of viruses to infect multiple host species and accounting for vector-mediated transmission. Hierarchical classification enables confident prediction at higher taxonomic levels when fine-grained labels are data-limited.
To handle the unequal distribution of sequence information across viral genomes — a key challenge for prediction on partial inputs — ViHostFinder incorporates mixture-of-experts and contrastive learning strategies, enabling robust inference even from fragmented metagenomic assemblies. Finally, the framework includes a metric to estimate host jump capacity, directly relevant to the emergence of novel viral threats. ViHostFinder will be freely available as a web server, designed with applicability in mind for the broader virology community.

C-B.15: Inter-individual variation in gut microbiome composition impacts identification of tissue type and cancer status in colorectal cancer
Track: Biodiversity, sustainability, envirobioinformatic
  • Alexander Bartholomew, Claremont McKenna College - Kravis Department of Integrated Sciences, United States
  • Shibu Yooseph, Claremont McKenna College - Kravis Department of Integrated Sciences, United States


Presentation Overview: Show

Colorectal cancer (CRC) is a leading cause of cancer-related deaths worldwide. There is growing evidence linking CRC to the gut microbiome. We present here a novel analysis of a previously published CRC gut microbiome study (NCBI Bioproject PRJNA743150). The 16S rRNA sequence dataset in this study was generated from biopsy samples collected from a cohort of 51 patients. Each patient in this cohort was labelled as recurrent or nonrecurrent for CRC and contributed two tissue types (normal and tumor). We assessed microbiome composition variation in individuals and across tissue types. Analysis of the 16S based taxonomic profiles (after centered log-ratio transformation) revealed that paired normal and tumor profiles from the same individual had significantly higher similarity (vector dot product) compared to randomized pairings of normal and tumor profiles (0.72 vs. 0.42, p < 0.0001). This implies that samples group more closely based on source individual than tissue type, which has implications for distinguishing between normal vs. tumor samples. We also trained random forest classifiers (RFC) to use taxonomic profiles to predict CRC recurrence. One RFC was trained using unpaired sample data, and the second model trained using taxonomic profile distance between matched normal and tumor samples. The first RFC model performed poorly (57% accuracy), while the second had a higher performance (80% accuracy). Our analysis emphasizes the importance of considering the longitudinal sampling of patients when predicting CRC recurrence and carefully designing studies to counteract the high variation in taxonomic profiles across individuals and tissue types.

C-B.16: ANETO: A Host-Agnostic Platform for Microbiome and Multi-Omics Discovery
Track: Biodiversity, sustainability, envirobioinformatic
  • Vincent Darbot, Aviwell, France
  • Erwann Chinal, Aviwell, France
  • Alexandre Fourment, Aviwell, France
  • Reda Mekdad, Aviwell, France
  • Arnaud Di Franco, Aviwell, France


Presentation Overview: Show

Host-associated microbiomes shape health, productivity and resilience, yet their functional integration with host biology remains challenging, particularly in under-represented agricultural, environmental and aquaculture species. We present ANETO, Aviwell's Discovery Platform, a host-agnostic framework designed to transform metagenomic and multi-omics data into interpretable biological hypotheses and intervention strategies.

ANETO combines standardized bioinformatics workflows, curated microbial genome resources and AI-driven discovery of host–microbiome interactions. Aviwell has generated sequencing data supporting a collection of more than 580,000 microbial genomes, expanding taxonomic resolution for microbiomes from uncommon hosts, including avian and aquatic species. To evaluate metagenomic profiling in poultry-relevant contexts while addressing the limitations of human-centred benchmarks, we built synthetic chicken caecal microbiome datasets. Across multiple community compositions, ANETO outperformed MetaPhlAn and METEOR, achieving F1 median scores at least 10% higher while improving both taxonomic accuracy and quantitative recovery of microbial community structure.

Beyond microbial composition, ANETO integrates molecular and host gene-expression measurements using machine learning-based network inference to model multimodal host–microbiome interactions, including gut–brain-axis-relevant associations. Interactive exploration of these networks enables prioritization of links between phenotypes, microbial communities, molecules and host genes. Resulting associations showed a two-fold enrichment for genes belonging to the same metabolic pathways or functional ontologies compared with random associations, supporting biological coherence.

ANETO provides a scalable framework for converting microbiome and multi-omics data into actionable insights. It has supported applications including feed-conversion improvement, welfare-related traits, pathogen resilience and sustainable aquaculture, while enabling biomarker discovery, postbiotic identification and microbiome-informed intervention design.

C-B.17: Consistent, scalable prokaryotic genome annotation and clustering to support biodiversity representation across EMBL-EBI resources
Track: Biodiversity, sustainability, envirobioinformatic
  • Christina Vasilopoulou, EMBL-EBI, United Kingdom
  • Johanna von Wachsmann, EMBL-EBI, United Kingdom
  • Tatiana A. Gurbich, EMBL-EBI, United Kingdom
  • Robert D. Finn, EMBL-EBI, United Kingdom


Presentation Overview: Show

We are entering an era of rapidly advancing biodiversity and surveillance genomics, with millions of prokaryotic genomes now available. This introduces challenges in downstream analysis, including high redundancy (the 20 more frequent species account for 90% of the genomes, with a bias toward human pathogens). However, different resources at EMBL-EBI represent this genomic information at different granularities: the AMR portal aims to capture all genomes with an AMR phenotype, while resources such as Ensembl aim to provide a subset that represents the biodiversity. Furthermore, MGnify Genomes focuses on generating genome catalogues that represent microbial diversity within individual biomes.

To address this challenge, we have developed a reproducible and scalable Nextflow pipeline that integrates the following: (i) comprehensive genome quality controls and codon usage estimation; (ii) gemsparcl, a tool for ultra-fast bacterial genome clustering into genomically coherent units, enabling the grouping of similar genomes for pangenome analysis, while enabling the removal of redundancy; (iii) flexible rule-based system for representative genome selection; and (iv) mettannotator, a Nextflow pipeline for comprehensive prokaryotic genome annotation. We benchmarked our approach using tens of thousands of prokaryotic genomes from the European Nucleotide Archive. Coupled to this, we intend to provide a lightweight tracking database to store and record metadata, such as genome quality, biome and clustering status, which we aim to make publicly available.

Our pipeline demonstrates a use case for scalable representation of prokaryotic genomes across species, which can be integrated across open-access resources such as MGnify and Ensembl.