View posters by category
Scroll down to view Results
Session A: Monday 31 August 12:00-13:30
|
Session B: Tuesday 1 September 16:15-17:45
|
|
|
|
Session C: Wednesday 2 September 11:30-13:00
|
|
|
|
Results
C-B.01: Metabolic modeling reveals microbial metabolites associated with infant temperament traits
Track: Biodiversity, sustainability, envirobioinformatic
-
Abhijit Paul, Turku Bioscience Centre, University of Turku and Ã…bo Akademi University, 20520 Turku,
Finland, Finland
-
Anna-Katariina Aatsinki, FinnBrain Birth Cohort Study, Department of Clinical Medicine, University of
Turku and Turku University Hospital, Turku, Finland
-
Matilda KrÃ¥kström, Turku Bioscience Centre, University of Turku and Ã…bo Akademi University, 20520 Turku,
Finland, Finland
-
Minna Lukkarinen, FinnBrain Birth Cohort Study, Department of Clinical Medicine, University of Turku and
Turku University Hospital, Turku, Finland
-
Eveliina Munukka, Centre for Population Health Research, University of Turku and Turku University
Hospital, Turku, Finland, Finland
-
Venla Huovinen, FinnBrain Birth Cohort Study, Department of Clinical Medicine, University of Turku and
Turku University Hospital, Turku, Finland
-
Saara Nolvi, FinnBrain Birth Cohort Study, Department of Clinical Medicine, University of Turku and Turku
University Hospital, Turku, Finland
-
Eeva-Leena Kataja, FinnBrain Birth Cohort Study, Department of Clinical Medicine, University of Turku and
Turku University Hospital, Turku, Finland
-
Riikka Korja, FinnBrain Birth Cohort Study, Department of Clinical Medicine, University of Turku and Turku
University Hospital, Turku, Finland
-
Nitin Bayal, Department of Computing, University of Turku, 20014 Turku, Finland, Finland
-
Hasse Karlsson, FinnBrain Birth Cohort Study, Department of Clinical Medicine, University of Turku and
Turku University Hospital, Turku, Finland
-
Leo Lahti, Department of Computing, University of Turku, 20014 Turku, Finland, Finland
-
Santosh Lamichhane, Research Center for Infections and Immunity, Institute of Biomedicine, University of
Turku, Turku, Finland, Finland
-
Linnea Karlsson, FinnBrain Birth Cohort Study, Department of Clinical Medicine, University of Turku and
Turku University Hospital, Turku, Finland
-
Alex M. Dickens, Turku Bioscience Centre, University of Turku and Ã…bo Akademi University, 20520 Turku,
Finland, Finland
-
Matej OreÅ¡iÄ, Turku Bioscience Centre, University of Turku and Ã…bo Akademi University, 20520 Turku,
Finland, Finland
Presentation Overview: Show
Background: Early temperament traits are important predictors of later socioemotional development and mental
health outcomes. While infant gut microbiota has been linked to early temperament traits, the underlying
microbial functional mechanisms remain largely unexplored. Here we examine whether microbial metabolites at
2.5 months of age are associated with positive and negative reactivity at 6 months, as they may serve as early
intermediate phenotypes of later child psychiatric disorders.
Methods: Stool samples from 2.5-month-old infants (n = 207) in the FinnBrain Birth Cohort underwent shotgun
metagenomics sequencing. Species-level taxonomic profiles were generated with nf-core/taxprofiler and
MetaPhlAn4. Individualized microbial community metabolic models were constructed using Microbiome Modelling
Toolbox to estimate metabolite secretion potentials. We examined associations between taxonomic composition,
predicted metabolite secretion potentials, and mother-reported Infant Behavior Questionnaire–Revised (IBQ-R)
measures of positive and negative reactivity at 6 months of age. Model predictions were compared with fecal
metabolomics data.
Results: Community metabolic modeling revealed that bile acids, short-chain fatty acids, amino acids,
vitamins, and glycans were associated with temperament traits. Taurine- and glycine-conjugated bile acids
positively associated with negative emotionality, while microbially conjugated bile acids with other amino
acids positively associated with surgency (positive emotionality). Short-chain fatty acids, particularly
butyrate and isobutyrate, were negatively associated with fear reactivity.
Conclusions: By integrating shotgun metagenomics with community metabolic modeling, we identified associations
between microbial bile acid and short-chain fatty acid metabolism and infant temperament traits, supported by
stool metabolomics. Beyond taxonomic associations, these findings highlight potential microbial functional
contributions to early behavioral phenotypes.
C-B.02: Deep Orthogroups: Extending Orthogroups beyond the root of the species tree.
Track: Biodiversity, sustainability, envirobioinformatic
-
Jonathan Holmes, University of Oxford, United Kingdom
- Steven Kelly, University of Oxford, United Kingdom
Presentation Overview: Show
An orthogroup is the central unit of comparative genomics, which consists of a set of genes descended from a
single gene in the last common ancestor of a set of species. Programs such as OrthoFinder, FastOMA and Sonic
Paranoid aim to identify these orthogroups and infer orthologs. Orthogroups are often limited to the root of a
species tree which ensures that ancient duplications beyond the root are separated. However, further
evolutionary history exists between genes of different orthogroups beyond the root of the species tree, which
we term deep orthogroups. Here, we present a workflow for mapping these connections to identify deep
orthogroups by characterising the evolutionary history of distantly related orthogroups and identifying
duplications beyond the root of the species tree. Our workflow first utilises the orthogroups initially
derived by OrthoFinder, using the BLAST scores between genes computed by OrthoFinder to linked distantly
related orthogroups, we then perform the additional steps of re-clustering existing orthogroups into new
clusters then using a modified DendroBLAST distance for phylogenetic inference. We show through the use of
simulation studies that we can not only recapture related orthogroups into deep orthogroups, but we also
accurately re-create the phylogenetic history of these orthogroups, allowing deep duplication events to be
mapped. We then apply this to a range of Eukaryotes across the tree of life in order to infer ancient gene
duplications.
C-B.03: Genome-Centric Characterisation of Limnochordia in Biogas Microbiomes: From Phylogenetic Diversity to
Clade-Specific Functions
Track: Biodiversity, sustainability, envirobioinformatic
-
Zihan Dai, Institute of Bio- and Geosciences IBG-5: Computational Metagenomics, Forschungszentrum
Jülich, Germany
-
Irena Maus, Institute of Bio- and Geosciences IBG-5: Computational Metagenomics, Forschungszentrum Jülich,
Germany
-
Benedikt Osterholz, Institute of Bio- and Geosciences IBG-5: Computational Metagenomics, Forschungszentrum
Jülich, Germany
-
Tom Tubbesing, Computational Metagenomics Group, Faculty of Technology and Center for Biotechnology
(CeBiTec), Bielefeld University, Germany
-
Liren Huang, Computational Metagenomics Group, Faculty of Technology and Center for Biotechnology
(CeBiTec), Bielefeld University, Germany
-
Sebastian Jünemann, Institute of Bio- and Geosciences IBG-5: Computational Metagenomics, Forschungszentrum
Jülich, Germany
-
Andreas Schlüter, Computational Metagenomics Group, Faculty of Technology and Center for Biotechnology
(CeBiTec), Bielefeld University, Germany
-
Alexander Sczyrba, Institute of Bio- and Geosciences IBG-5: Computational Metagenomics, Forschungszentrum
Jülich, Germany
Presentation Overview: Show
Anaerobic digestion (AD) of biomass for methane production is a key component of renewable energy systems,
with process stability governed by complex, poorly characterised microbial communities. The class Limnochordia
has emerged as a consistently abundant and resilient member of biogas microbiomes, yet its genomic and
functional diversity remains unresolved. Here, we present a comprehensive, genome-centric characterisation of
Limnochordia across 5 biogas plants comprising 10 reactors, combining in-house metagenomic data with publicly
available biogas datasets.
Metagenome assembly and binning yielded 77 Limnochordia MAGs, 24 reconstructed from in-house samples and 53
from public data, of which 69 passed quality thresholds (completeness ≥80%, contamination ≤10%). These
high-quality MAGs span 5 families, 18 genera, and 36 species. Family DTU010 dominates the dataset (50 MAGs, 9
genera, 23 species), followed by DTU012 (15 MAGs, 5 genera, 9 species); three additional families are
represented by 1 to 2 MAGs each. Phylogenetic analysis resolved clade structure across all 69 MAGs. KEGG-based
functional profiling indicates that functional repertoires are broadly partitioned along phylogenetic lines,
with further within-family heterogeneity revealed at higher pathway completeness thresholds. Detailed
differential pathway analysis across clades is ongoing.
Metaproteomic data will be mapped against these MAG references to link genomic potential to active expression
under varying reactor conditions. Collectively, this work provides a resolved phylogenetic and functional
framework for Limnochordia in engineered anaerobic systems, with direct implications for microbiome management
and inocula development in biogas production.
C-B.04: Unmasking the Growing Challenge and Hidden Impact of Sampling Bias in Prokaryotic Genome
Collections
Track: Biodiversity, sustainability, envirobioinformatic
-
Hannah Goetsch, JLU Giessen, Germany
- Franz Baumdicker, JLU Giessen, Germany
Presentation Overview: Show
The rapid growth of fully sequenced prokaryotic genomes improved our ability to study prokaryotic diversity,
adaptation, and the spread of traits such as antibiotic resistance. However, many genome collections have been
assembled through heterogeneous study designs rather than systematic sampling, often leading to an
overrepresentation of closely related individuals of particular clinical, ecological, or epidemiological
interest. Such sampling bias reduces the effective information, biases downstream analyses, and limits the
ability to capture true prokaryotic diversity.
We investigate the extent and consequences of sampling bias in the NCBI prokaryotic genome database. We show
that oversampling has become widespread and increasingly severe, with substantial effects on downstream
analyses. To address this problem, we introduce PhyloThin, a method that detects and removes sampling bias
from prokaryotic genome collections using phylogenetic relationships within a coalescent framework. Applying
PhyloThin to 402 prokaryotic species reveals pervasive oversampling trends, especially in recent years. Across
species, the effective sample size is on average nearly half of the total number of genomes. Importantly, this
bias distorts gene frequency estimates and leads to underestimation of prokaryotic diversity and adaptive
potential.
Our results provide a robust framework for identifying oversampled genomes and improving the accuracy of
inferences from large-scale prokaryotic genome datasets. In particular, reducing sampling bias enables more
reliable assessments of prokaryotic diversity and evolutionary potential. More broadly, exposing these biases
can help prioritize species and lineages for future sequencing efforts to better capture the full extent of
prokaryotic biodiversity.
C-B.05: FlaPro: a computational framework for TLR5-phenotype-aware flagellome profiling across human gut
metagenomes and metatranscriptomes
Track: Biodiversity, sustainability, envirobioinformatic
-
Anna Bogdanova, Department of Microbiome Science, Max Planck Institute for Biology Tübingen,
Germany
-
Andrea Borbón-García, Department of Microbiome Science, Max Planck Institute for Biology Tübingen,
Germany
-
Ruth Ley, Department of Microbiome Science, Max Planck Institute for Biology Tübingen, Germany
-
Alexander Tyakht, Department of Microbiome Science, Max Planck Institute for Biology Tübingen, Germany
Presentation Overview: Show
Flagellin, the structural protein of bacterial flagella, activates the innate immune receptor Toll-like
receptor 5 (TLR5). However, different flagellins vary widely in their ability to stimulate TLR5 - from highly
immunostimulatory variants to ""silent"" ones that bind TLR5 without triggering signaling. This suggests that
the composition of an individual's flagellin repertoire, or flagellome, is a potential distinct contributor to
host-microbiome interactions and inflammation. Despite growing evidence that flagellins contribute to diseases
such as inflammatory bowel disease (IBD), computational methods to resolve both flagellome composition and its
TLR5-stimulatory potential from microbiome sequencing data have been lacking.
We developed FlaPro, a Snakemake-based pipeline for quantification and functional annotation of human gut
flagellomes from metagenomic (MGX) and metatranscriptomic (MTX) data. FlaPro combines ShortBRED marker-based
quantification against a curated reference of non-redundant flagellin clusters from human gut microbiome with
a Random Forest classifier trained on experimentally characterized flagellins to predict per-flagellin TLR5
phenotype.
We apply FlaPro in a cross-cohort meta-analysis of public gut MGX/MTX datasets spanning diverse diseases,
generating harmonized flagellome profiles and identifying both shared and condition-specific TLR5-phenotype
signatures. For an IBD dataset, compared to healthy controls we discovered reduced flagellome diversity and a
lower silent-to-stimulatory ratio in both Crohn's disease and ulcerative colitis consistently across MGX and
MTX – a disease-associated shift toward a more stimulatory flagellome. Similarly, condition-specific
alterations have been identified for other diseases. Together, these analyses position the flagellome as a
quantifiable, functionally interpretable microbiome feature for microbiome-wide association studies in health
and disease.
C-B.06: Fast, flexible gene cluster family delineation with IGUA
Track: Biodiversity, sustainability, envirobioinformatic
-
Martin Larralde, Leiden University Center for Infectious Diseases, Leiden University Medical Center,
Leiden, Netherlands
-
Josefin Blom, Department of Clinical Microbiology, SciLifeLab, Umeå University, Umeå, Sweden
-
Hadrien Gourlé, Department of Clinical Microbiology, SciLifeLab, UmeÃ¥ University, UmeÃ¥, Sweden
-
Hale-Seda Radoykova, Global Health Institute, School of Life Science, EPFL, Lausanne, Switzerland
-
Lucas Paoli, Global Health Institute, School of Life Science, EPFL, Lausanne, Switzerland
-
Laura Carroll, Department of Clinical Microbiology, SciLifeLab, Umeå University, Umeå, Sweden
-
Georg Zeller, Leiden University Center for Infectious Diseases, Leiden University Medical Center, Leiden,
Netherlands
Presentation Overview: Show
Prokaryotic genomes harbor a variety of functional elements encoded as contiguous multi-gene clusters. To
identify homologous gene clusters, computational tools that account for multi-gene architectures are needed to
group gene clusters into Gene Cluster Families (GCFs). However, existing GCF delineation methods do not scale
well to large datasets and are often limited to certain subclasses of gene clusters.
Here, we present IGUA (Iterative Gene clUster Analysis; https://github.com/zellerlab/IGUA), a scalable,
flexible GCF delineation method for genomic segments with multi-gene architectures.
IGUA uses Mmseqs2 clustering of both nucleotide and protein sequences with different parameters tuned for
optimal de-replication. The final gene cluster families are obtained with hierarchical clustering over the
protein content of the gene clusters, which captures homology despite organizational changes at the nucleotide
scale. To that end, we propose an efficient distance metric for measuring pairwise distances between gene
clusters without a priori reference features.
On a biosynthetic gene cluster (BGC) clustering task, IGUA is >18x and 5x faster than the state-of-the-art
(BiG-SCAPE 2.0 and BiG-SLiCE 2.0, respectively), without sacrificing accuracy. To highlight its scalability,
we use IGUA to cluster >2.8 million BGCs from ~1 million prokaryotic genomes in <18 hours (n = 2,829,071
BGCs to 56,960 GCFs). Demonstrating its utility beyond BGC clustering, we use IGUA to cluster five additional
prokaryotic gene cluster types (secretion systems, prophages, genomic islands, phage-plasmids, and
carbohydrate active enzyme [CAZyme] gene clusters).
Overall, IGUA represents a versatile GCF delineation tool with unmatched computational efficiency and
flexibility, enabling (meta)genomic mining applications at unprecedented scales.
C-B.07: MetagenomeWatch: Web-based monitoring and search of public metagenomic data for pathogen
surveillance
Track: Biodiversity, sustainability, envirobioinformatic
-
Matthew Huska, Robert Koch Institute, Germany
- Daniel Desiro, Robert Koch Institute, Germany
- Martin Hoelzer, Robert Koch Institute, Germany
Presentation Overview: Show
The rapid growth of next-generation sequencing has led to an unprecedented expansion of publicly available
metagenomic data. The Sequence Read Archive (SRA) now contains over nine million sequencing datasets spanning
diverse environments, including human, animal, and wastewater samples. This resource provides substantial
potential for improving our understanding of microbial communities and their relevance to public health.
However, effective search and analysis at this scale remain challenging due to the computational and storage
demands of petabyte-scale data.
Recent advances in sequence sketching enable highly compressed representations that retain sufficient
information to probabilistically detect the presence of sequences within large datasets, making large-scale
search of metagenomic repositories computationally feasible.
Here, we present MetagenomeWatch, a web-based system built on sourmash branchwater for large-scale sequence
sketching and comparison. The platform enables efficient querying of public metagenomic data for viral,
bacterial, and eukaryotic pathogens. Leveraging precomputed sketches and branchwater-based search, typical
queries across the full SRA metagenomics collection can be completed in under five minutes while maintaining
sensitivity for sequence detection. In addition, the system provides automated monitoring workflows that track
newly processed datasets and report matches to user-defined queries.
This approach enables scalable, continuous surveillance of publicly available metagenomic data and facilitates
early detection and tracking of pathogens, with potential applications in outbreak monitoring and public
health response.
MetagenomeWatch is open-source software (AGPL-3.0) and can be found at: https://github.com/rki-mf1/mgwatch
C-B.08: Comparative methods for RNA Secondary Structure Prediction in RNA Viruses
Track: Biodiversity, sustainability, envirobioinformatic
-
Charlotte Tumescheit, IDIAP, SIB, Switzerland
- Katherine Brown, Cambridge University, United Kingdom
- Janna Hastings, IDIAP, SIB, Switzerland
- Andrew E Firth, Cambridge University, United Kingdom
Presentation Overview: Show
RNA secondary structures play many different roles in the life cycle of RNA viruses, for example in directing
translational control and genome replication. However, predicting these structures can be challenging,
especially when it comes to pseudoknots, long-range interactions, and mutually exclusive interactions.
Here, we use a multiple sequence alignment based approach to mitigate these challenges by specifically looking
for all possibly functionally relevant conserved short-range and long-range interactions, thereby allowing for
the detection of pseudoknots and mutually exclusive structures.
To allow functional structures to be identified, we incorporated an offset parameter, allowing imperfect and
more divergent alignments; a phylogenetic weighting scheme, to balance diversity with specificity; and a free
energy calculation.
For benchmarking purposes, we curated a dataset of verified secondary structures in RNA viruses, showing that
our tool successfully predicts known structures (sensitivity of 77.78%) besides suggesting novel functionally
relevant structures.
Next, we improve existing transformer models for RNA analysis by incorporating knowledge-injection strategies
as well as fine-tuning on our curated database to further improve secondary structure prediction for RNA
viruses.
While RNA structure prediction is still struggling to match the performance of protein structure predictions,
we present different strategies to improve the predictions of RNA structures, using methods based on
comparative genomics, with and without the incorporation of machine learning. Specifically focusing on RNA
viruses and using curated databases already shows the usefulness of our approaches, which we aim to expand to
other RNA use cases in the future.
C-B.09: ABRomics: a platform for antibiotic resistance research and public health using an integrated One
Health approach
Track: Biodiversity, sustainability, envirobioinformatic
-
Brieuc Quemeneur, CNRS IFB-core & Nantes Université, CNRS, INSERM, l’institut du thorax, France
-
Alban Gaignard, CNRS, IFB-core & Nantes Université, CNRS, INSERM, l’institut du thorax, France
-
Samuel Chaffron, Nantes UniversiteÌ, EÌcole Centrale Nantes, CNRS, LS2N, UMR 6004, F-44000 Nantes, France
-
Audrey Bihouée, Nantes Université, CNRS, INSERM, l’institut du thorax, F-44000 Nantes, France
- Abromics Consortium, CNRS IFB-core, UAR 3601 - Villejuif, France
- Gildas Le Corguillé, CNRS, IFB-core & ABiMS, Station Biologique-Roscoff, France
-
Etienne Ruppé, Univ. Paris Cité and Univ.Sorbonne Paris Nord, Inserm, IAME - Paris, France
- Nadia Goué, CNRS IFB-core & AuBi platform, Université Clermont-Auvergne, France
-
Bérénice Batut, CNRS IFB-core & AuBi platform, Université Clermont-Auvergne, France
-
Pierre Marin, CNRS IFB-core & AuBi platform, Université Clermont-Auvergne, France
-
Claudine Médigue, CNRS IFB-core & CEA, Genoscope, LABGeM, France
- Hugo Lefeuvre, Nantes Université, CNRS, INSERM, l’institut du thorax, France
- Cléa Siguret, CNRS IFB-core & AuBi platform Université Clermont-Auvergne, France
- Thomas Mignon, CNRS FB-core & ABiMS, Station Biologique - Roscoff, France
-
Amanda Dieuaide, CNRS IFB-core & Institut Pasteur, Bioinformatics and Biostatistics Hub, France
-
Raphaël Tackx, CNRS IFB-core & Institut Pasteur, Bioinformatics and Biostatistics Hub, France
-
Julie Lao, CNRS, Institut Français de Bioinformatique, IFB-core, UAR 3601 - Villejuif, France
- Philippe Glaser, Institut Pasteur, Unité EERA, CNRS UMR604, Paris, France
-
Fabien Mareuil, Institut Pasteur, Université Paris Cité, Bioinformatics and Biostatistics Hub, Paris,
France
Presentation Overview: Show
Antibiotic resistance is a major public health concern. Whole Genome Sequencing (WGS) is essential for
characterizing the diffusion and the monitoring of Multidrug resistant bacteria (MDRB) within and between the
human, animal and environmental sectors. In this context, the French Priority Plan on Antimicrobial Resistance
has funded the development of an online platform dedicated to the surveillance and research of antibiotic
resistance within a One Health framework (www.abromics.fr/). The ABRomics platform is hosted at the French
Institute of Bioinformatics (www.ifb-elixir.fr/), which has extensive data analysis and storage capacity.
The web service ABRomics-analysis (analysis.abromics.fr/) offers a user-friendly interface for uploading and
managing genomic samples as well as for running, through usegalaxy.fr, a standardized workflow comprising:
quality control, species identification, MLST (MultiLocus Sequence Typing), genome assembly and annotation,
cgMLST (core genome MLST) using standardized nomenclatures, plasmid typing, and the detection of antibiotic
resistance and virulence genes. The ABRomics database aims to facilitate data exploration and foster
collaboration across sectors and communities. To this end, it provides tools for exploring freely accessible
data on all submitted isolates and for connecting scientists who share related strains.
Ongoing developments of the ABRomics platform include data brokering of genomic samples to the European
Nucleic Archive (ENA), standardized workflows for metagenomics analysis, and tools for analyzing
(meta)pangenomes. Importantly, data management procedures comply with the FAIR principle (Findable,
Accessible, Interoperable and Reusable) thuswill enablinge retrospective epidemiological studies to be
conducted.
C-B.10: Phylogenetic ordering and batching for better compression of million-genome bacterial
collections
Track: Biodiversity, sustainability, envirobioinformatic
-
Tam Truong, INRIA Rennes, France, France
- Dominique Lavenier, CNRS, Rennes, France,, France
- Pierre Peterlongo, INRIA Rennes, France, France
- Karel Břinda, INRIA Rennes, France, France
Presentation Overview: Show
Modern bacterial genome collections such as AllTheBacteria (ATB) and GTDB contain millions of isolate genomes
and metagenome-assembled genomes (MAGs). Efficient compression of such collections requires batching and
ordering of genomes that co-localizes shared redundancy before low-level compression. While phylogenetic
compression enables such an ordering efficiently using evolutionary history, its state-of-the-art
implementation MiniPhy, used as the core method for ATB, is limited by fixed-size batching based on metadata
proxies like species and accessions. Consequently, compression degrades on collections with oversampled
species, high taxonomic diversity, and mixed isolate and MAG content.
We present PhyloPack, a scalable method for global phylogenetic ordering and batching via skeleton
phylogenies. We formalize both as an optimization problem aimed at maximizing compression efficiency under
fixed computational constraints. PhyloPack employs a two-step heuristic: first, a skeleton phylogeny is
inferred from a subsampled set of genomes; second, the remaining sequences are placed within the skeleton tree
using sketch-based nearest neighbor identification. On ATBv0.3, PhyloPack with MBGC2 reduces the major-species
subset (2.2 millions genomes) from 57 GB (XZ) to 15 GB, and the whole collection (2.4 millions genomes) from
103GB (XZ) to 41 GB (60% reduction to the current ATB distribution). Combined with XZ and AGC, we observe a
consistent 40% improvement over MiniPhy. On GTDB (r226), using MBGC2, PhyloPack compresses the isolate to 35GB
and MAG subsets to 70GB, respectively (totaled to 8% of the size of the official gzip files). Overall,
PhyloPack makes phylogenetic compression practical for heterogeneous genome collections, scalable beyond
millions of genomes.
C-B.11: A research-oriented framework for viral discovery and interpretation in metagenomic data
Track: Biodiversity, sustainability, envirobioinformatic
-
Amanj Bajalan, Department of Microbiology, Tumor and Cell Biology, Karolinska Institutet, Stockholm,
Sweden, Sweden
-
Björn Andersson, Department of Cell and Molecular Biology (CMB), Karolinska Institutet, Stockholm, Sweden,
Sweden
-
Tobias Allander, Department of Microbiology, Tumor and Cell Biology, Karolinska Institutet, Stockholm,
Sweden, Sweden
Presentation Overview: Show
Viruses from many families are important human pathogens, and advances in metagenomic sequencing have enabled
broad and unbiased virus detection. At the same time, there is a need for reproducible workflows that support
both detection and interpretation of ambiguous and potentially novel viral findings. We developed a scalable
and modular viral discovery framework for reproducible metagenomic analysis. The workflow is implemented in
Nextflow and integrates read-based analysis, assembly creation, assembly-based classification, assembly
statistics, machine learning-based post-classification, and interactive HTML reporting. In addition to
automated processing, it includes inspection tools for interpretation of ambiguous and unclassified sequences,
including full lineage reporting, metadata, filterable human-virus tagging, built-in support for querying NCBI
resources, a built-in ORF viewer for unknown viral hits, and a viral likelihood scoring system for
prioritization of uncertain findings. This makes the workflow useful for both virus detection and
research-oriented interpretation of unclear results. The workflow was applied to blood plasma, CSF, and fecal
samples from Swedish clinical data, together with publicly available oral microbiome data generated for
bacterial profiling rather than viral discovery. The analysis identified both RNA and DNA viruses from
multiple viral families across diverse sample types, and supported prioritization of uncertain or potentially
novel viral sequences. In the oral microbiome dataset, the workflow recovered and supported inspection of a
contig matching HIV despite the dataset's original bacterial focus, highlighting utility for exploratory
cross-domain metagenomic analysis. The workflow generates portable HTML reports for rapid overview and
downstream analysis, and is being applied in ongoing research studies.
C-B.12: Characterization of Microbial Succession During Spontaneous Fermentation Using Long-Read Amplicon
Sequencing
Track: Biodiversity, sustainability, envirobioinformatic
-
Oliver Scharinger, Hochschule Campus Wien, Austria
- Lukas Fürnwein, Hochschule Campus Wien, Austria
- Alexandra Graf, Hochschule Campus Wien, Austria
Presentation Overview: Show
Regional distinctiveness of wines is strongly shaped by the composition and activity of site-specific
microorganisms, which is known as the microbial terroir concept. For this reason, spontaneously fermented
wines are gaining popularity, whereas wines relying on inoculated fermentations with commercial yeast strains
tend to standardized aroma profiles and reduced sensorical complexity. Therefore the microbial dynamcis during
a spontaneous fermentation of an Austrian Grüner Veltliner from a winery in Gols (Burgenland), was analyzed
using long-read amplicon sequencing. Samples of unripe grapes, mash and fermenting must (every three days)
were collected, DNA was extracted and then sequenced on an Oxford Nanopore platform as 16S- and ITS-amplicons.
The results show that the washed grape samples and the mash sample are
strongly affected by co-amplification of chloroplast and mitochondrial Vitis vinifera rRNA genes.Over the
course of fermentation, Tatumella/Pantoea (29.78-59.02%) were the most abundant genera. At fermentation start,
lactic acid bacteria of the genera Lactococcus (32.63%) and Lactobacillus/Apilactobacillus (27.43%) showed
high relative abundances and decreased over time,whereas Lactiplantibacillus increased with the course of
fermentation. In one sample, Gluconobacter was detected at a relative abundance of 9%. The fungal community
showed a transition from Hanseniaspora at the beginning of fermentation towards a Saccharomyces-dominated
community from mid to end of fermentation. The study demonstrates the potential of long-read amplicon
sequencing for determining the microbial composition of spontaneously fermented wine samples, but also
identifies plant-associated DNA contaminations as a technical limitation that should be improved in future
studies by optimized host-depletion strategies.
C-B.13: Plant species pool size predicts arbuscular mycorrhizal fungal diversity and biomass better than
local richness
Track: Biodiversity, sustainability, envirobioinformatic
-
Bo Stevens, University of Tartu, Estonia
- Darkdivnet Consortium, University of Tartu, Estonia
- Martti Vasar, University of Tartu, Estonia
- Koit Herodes, University of Tartu, Estonia
- Ayesh Wipulasena, University of Tartu, Estonia
- Tanel Vahter, University of Tartu, Estonia
- Martin Zobel, University of Tartu, Estonia
- Meelis Pärtel, University of Tartu, Estonia
Presentation Overview: Show
Arbuscular mycorrhizal (AM) fungi are globally important plant symbionts, yet the relative contributions of
environmental conditions and plant diversity metrics to AM fungal diversity and biomass remain incompletely
resolved at broad spatial scales. We evaluated factors associated with AM fungal diversity (effective number
of species) and two fungal biomass proxies, NLFA 16:1ω5 and PLFA 16:1ω5, across 99 globally distributed
DarkDivNet sites using Random Forest models integrating soil, climate, fungal guild, and plant diversity
metrics. We specifically tested whether plant diversity metrics derived from dark-diversity theory -- species
pool size (regional biodiversity potential) and community completeness (realization of that potential) --
improved prediction beyond abiotic covariates and local plant richness alone. Models incorporating plant
diversity metrics consistently outperformed abiotic-only models. The best-performing models explained 18% of
held-out variance for AM fungal diversity, 20% for NLFA, and 28% for PLFA. Plant species pool size was
retained in the best-performing model for all three responses, either alone or together with plant community
completeness. AM fungal diversity was positively associated with plant species pool size and soil pH, whereas
PLFA was more strongly associated with soil carbon and precipitation. Ectomycorrhizal fungal diversity showed
negative associations with AM fungal diversity and NLFA. Overall, our results indicate that broad-scale
patterns of AM fungal diversity and biomass are linked more strongly to regional plant diversity potential
than to local plant richness and environmental conditions.
C-B.14: ViHostFinder: A Hierarchical Multilabel Host Predictor for Viral Sequences Using DNA Language
Models
Track: Biodiversity, sustainability, envirobioinformatic
-
Alfred Ferrer Florensa, Technical University of Denmark, Denmark
- Frank Møller Aarestrup, Technical University of Denmark, Denmark
- Henrik Nielsen, Technical University of Denmark, Denmark
- Philip Thomas Conradsen Clausen, Technical University of Denmark, Denmark
Presentation Overview: Show
Metagenomic sequencing has dramatically expanded our view of the global virome, yet our ability to assess the
biological relevance of novel viral sequences remains severely limited. Current approaches rely almost
exclusively on sequence similarity to known viruses — a strategy that fails systematically for the vast,
uncharacterized fraction of environmental viral diversity. Predicting which hosts a virus can infect is a
critical step toward evaluating zoonotic risk, understanding transmission dynamics, and enabling actionable
surveillance.
We present ViHostFinder, a universal viral host predictor built on a hierarchical multilabel classification
framework using DNA language models. By operating directly on raw genomic sequences, ViHostFinder bypasses
protein annotation — a step that remains unreliable for highly divergent viruses — making it broadly
applicable across the virome. The model natively supports multilabel prediction, capturing the capacity of
viruses to infect multiple host species and accounting for vector-mediated transmission. Hierarchical
classification enables confident prediction at higher taxonomic levels when fine-grained labels are
data-limited.
To handle the unequal distribution of sequence information across viral genomes — a key challenge for
prediction on partial inputs — ViHostFinder incorporates mixture-of-experts and contrastive learning
strategies, enabling robust inference even from fragmented metagenomic assemblies. Finally, the framework
includes a metric to estimate host jump capacity, directly relevant to the emergence of novel viral threats.
ViHostFinder will be freely available as a web server, designed with applicability in mind for the broader
virology community.
C-B.15: Inter-individual variation in gut microbiome composition impacts identification of tissue type and
cancer status in colorectal cancer
Track: Biodiversity, sustainability, envirobioinformatic
-
Alexander Bartholomew, Claremont McKenna College - Kravis Department of Integrated Sciences, United
States
-
Shibu Yooseph, Claremont McKenna College - Kravis Department of Integrated Sciences, United States
Presentation Overview: Show
Colorectal cancer (CRC) is a leading cause of cancer-related deaths worldwide. There is growing evidence
linking CRC to the gut microbiome. We present here a novel analysis of a previously published CRC gut
microbiome study (NCBI Bioproject PRJNA743150). The 16S rRNA sequence dataset in this study was generated from
biopsy samples collected from a cohort of 51 patients. Each patient in this cohort was labelled as recurrent
or nonrecurrent for CRC and contributed two tissue types (normal and tumor). We assessed microbiome
composition variation in individuals and across tissue types. Analysis of the 16S based taxonomic profiles
(after centered log-ratio transformation) revealed that paired normal and tumor profiles from the same
individual had significantly higher similarity (vector dot product) compared to randomized pairings of normal
and tumor profiles (0.72 vs. 0.42, p < 0.0001). This implies that samples group more closely based on
source individual than tissue type, which has implications for distinguishing between normal vs. tumor
samples. We also trained random forest classifiers (RFC) to use taxonomic profiles to predict CRC recurrence.
One RFC was trained using unpaired sample data, and the second model trained using taxonomic profile distance
between matched normal and tumor samples. The first RFC model performed poorly (57% accuracy), while the
second had a higher performance (80% accuracy). Our analysis emphasizes the importance of considering the
longitudinal sampling of patients when predicting CRC recurrence and carefully designing studies to counteract
the high variation in taxonomic profiles across individuals and tissue types.
C-B.16: ANETO: A Host-Agnostic Platform for Microbiome and Multi-Omics Discovery
Track: Biodiversity, sustainability, envirobioinformatic
-
Vincent Darbot, Aviwell, France
- Erwann Chinal, Aviwell, France
- Alexandre Fourment, Aviwell, France
- Reda Mekdad, Aviwell, France
- Arnaud Di Franco, Aviwell, France
Presentation Overview: Show
Host-associated microbiomes shape health, productivity and resilience, yet their functional integration with
host biology remains challenging, particularly in under-represented agricultural, environmental and
aquaculture species. We present ANETO, Aviwell's Discovery Platform, a host-agnostic framework designed to
transform metagenomic and multi-omics data into interpretable biological hypotheses and intervention
strategies.
ANETO combines standardized bioinformatics workflows, curated microbial genome resources and AI-driven
discovery of host–microbiome interactions. Aviwell has generated sequencing data supporting a collection of
more than 580,000 microbial genomes, expanding taxonomic resolution for microbiomes from uncommon hosts,
including avian and aquatic species. To evaluate metagenomic profiling in poultry-relevant contexts while
addressing the limitations of human-centred benchmarks, we built synthetic chicken caecal microbiome datasets.
Across multiple community compositions, ANETO outperformed MetaPhlAn and METEOR, achieving F1 median scores at
least 10% higher while improving both taxonomic accuracy and quantitative recovery of microbial community
structure.
Beyond microbial composition, ANETO integrates molecular and host gene-expression measurements using machine
learning-based network inference to model multimodal host–microbiome interactions, including
gut–brain-axis-relevant associations. Interactive exploration of these networks enables prioritization of
links between phenotypes, microbial communities, molecules and host genes. Resulting associations showed a
two-fold enrichment for genes belonging to the same metabolic pathways or functional ontologies compared with
random associations, supporting biological coherence.
ANETO provides a scalable framework for converting microbiome and multi-omics data into actionable insights.
It has supported applications including feed-conversion improvement, welfare-related traits, pathogen
resilience and sustainable aquaculture, while enabling biomarker discovery, postbiotic identification and
microbiome-informed intervention design.
C-B.17: Consistent, scalable prokaryotic genome annotation and clustering to support biodiversity
representation across EMBL-EBI resources
Track: Biodiversity, sustainability, envirobioinformatic
-
Christina Vasilopoulou, EMBL-EBI, United Kingdom
- Johanna von Wachsmann, EMBL-EBI, United Kingdom
- Tatiana A. Gurbich, EMBL-EBI, United Kingdom
- Robert D. Finn, EMBL-EBI, United Kingdom
Presentation Overview: Show
We are entering an era of rapidly advancing biodiversity and surveillance genomics, with millions of
prokaryotic genomes now available. This introduces challenges in downstream analysis, including high
redundancy (the 20 more frequent species account for 90% of the genomes, with a bias toward human pathogens).
However, different resources at EMBL-EBI represent this genomic information at different granularities: the
AMR portal aims to capture all genomes with an AMR phenotype, while resources such as Ensembl aim to provide a
subset that represents the biodiversity. Furthermore, MGnify Genomes focuses on generating genome catalogues
that represent microbial diversity within individual biomes.
To address this challenge, we have developed a reproducible and scalable Nextflow pipeline that integrates the
following: (i) comprehensive genome quality controls and codon usage estimation; (ii) gemsparcl, a tool for
ultra-fast bacterial genome clustering into genomically coherent units, enabling the grouping of similar
genomes for pangenome analysis, while enabling the removal of redundancy; (iii) flexible rule-based system for
representative genome selection; and (iv) mettannotator, a Nextflow pipeline for comprehensive prokaryotic
genome annotation. We benchmarked our approach using tens of thousands of prokaryotic genomes from the
European Nucleotide Archive. Coupled to this, we intend to provide a lightweight tracking database to store
and record metadata, such as genome quality, biome and clustering status, which we aim to make publicly
available.
Our pipeline demonstrates a use case for scalable representation of prokaryotic genomes across species, which
can be integrated across open-access resources such as MGnify and Ensembl.