View posters by category
Scroll down to view Results
Session A: Monday 31 August 12:00-13:30
|
Session B: Tuesday 1 September 16:15-17:45
|
|
|
|
Session C: Wednesday 2 September 11:30-13:00
|
|
|
|
Results
C-G.C.01: Invariom-Derived Geometric Priors Improve Conformer Generation Beyond Crystallographic Chemical
Space
Track: General computational biology
-
Rok Breznikar, Biozentrum, University of Basel; SIB Swiss Institute of Bioinformatics,
Switzerland
-
Janani Durairaj, Biozentrum, University of Basel; SIB Swiss Institute of Bioinformatics, Switzerland
-
Torsten Schwede, Biozentrum, University of Basel; SIB Swiss Institute of Bioinformatics, Switzerland
- Birger Dittrich, Novartis Pharma AG; University of Zurich, Switzerland
Presentation Overview: Show
Predicting the three-dimensional conformations of drug-like molecules is central to computational drug
discovery. Molecular docking, pharmacophore modeling, shape-based virtual screening, and structure–activity
analysis all depend on conformer ensembles, whose quality directly shapes downstream predictions. Current
distance geometry methods such as ETKDGv3, implemented in the widely used RDKit cheminformatics library, rely
on generic geometric priors and torsional preferences derived from crystallographic databases, limiting
coverage and requiring ongoing manual curation. Here, we replace these empirical priors with geometric
parameters derived from invariom model compounds computed at a semiempirical quantum-mechanical level of
theory. These parameters are precomputed once and reused across target molecules, so no per-molecule
quantum-mechanical calculation is required at conformer generation time. Coupling distance geometry with Monte
Carlo sampling guided by invariom-based energy terms yields conformational ensembles that are systematically
lower in energy than those generated by ETKDGv3, without the need for manually curated torsion rules. Because
invariom parameters are computed rather than extracted from experimental data, this approach extends naturally
to chemistries underrepresented in crystallographic repositories, providing a scalable route to improved
conformer generation for drug discovery.
C-G.C.02: IDEAL-GENOM: A Python toolkit for automation of genomic analysis
Track: General computational biology
-
Luis Giraldo Gonzalez Ricardo, University of Tubingen, Germany
- Amabel Tenghe, Independent, Germany
- Ashwin Ashok Kumar Sreelatha, Independent, Germany
- Manu Sharma, Universitätsklinikum Tübingen, Germany
Presentation Overview: Show
There has been an exponential increase in the development of technologies that have offered unprecedented
opportunities to unravel the genetic underpinnings of rare and complex diseases. However, despite the
progress, and availability of tools, there are still several challenges to overcome, such as the development
of a unified framework that facilitates downstream analysis. To address this challenge, we present IDEAL-GENOM
(Integrated Downstream Analytical Toolkit for Genomic Analysis), a Python-based wrapper designed to streamline
the analytical workflow commonly implemented in genome-wide association studies (GWAS) settings. IDEAL-GENOM
integrates widely used tools such as PLINK, GCTA and bcftools along with custom-developed functionalities,
enabling reproducible results through parameter sharing. IDEAL-GENOM provides a simplified framework for
quality control (QC) and GWAS and downstream analysis that will enable beginners and advanced users to
leverage the in-built functionalities to perform GWAS analysis for complex diseases. IDEAL-GENOM is a
customizable tool that enables users to efficiently analyze genotyping data with reproducible QC,
post-imputation, and GWAS pipelines.
C-G.C.03: Global antimicrobial resistance patterns in human gut metagenomes are structured along
socio-economic gradients
Track: General computational biology
-
Mahkameh Salehi, Dept. of Computing, University of Turku, Dept of Microbiology University of
Helsinki, Finland
-
Shivang Bhanushali, Dept. of Computing, University of Turku, Dept of Microbiology University of Helsinki,
Finland
- Eetu Tammi, Dept. of Computing, University of Turku, Finland
-
Peter Collingon, ACT Pathology, Canberra Hospital, Australia, Medical School, Australian National Univ.,
Australia, Australia
- John J. Beggs, Independent researcher, Melbourne, Australia, Australia
-
Johan Bengtsson-Palme, Chalmers Univ. of Technology, Centre for Antibiotic Resistance Research (CARe),
Sahlgrenska Academy, Univ. of Gothenburg, Sweden
- Leo Lahti, Dept. of Computing, University of Turku, Finland
-
Katariina Pärnänen, Dept of Microbiology University of Helsinki, Dept. of Computing, University of Turku,
Finland
Presentation Overview: Show
Antimicrobial resistance (AMR) surveillance is primarily based on clinical isolates, providing limited insight
into population-level resistance dynamics. Metagenomic sequencing enables large-scale profiling of antibiotic
resistance genes (ARGs), but integrating such data with heterogeneous socio-economic and ecological variables
requires robust statistical and computational frameworks.
We analysed over 58,000 publicly available human gut metagenomes from NCBI and ENA, spanning more than 60
countries. ARG profiles were quantified by mapping reads to the ResFinder database, and microbiome composition
was derived from marker-gene-based taxonomic annotation. These data were integrated with country-level
indicators, including antibiotic consumption, socioeconomic indicators, and global connectivity.
We combined generalized additive models and Bayesian multilevel models to capture non-linear associations and
region-specific variability while accounting for hierarchical structure. To explicitly model relationships
between countries, we constructed dissimilarity matrices and apply dyadic regression to quantify how resistome
differences correlate with geographic, socioeconomic, and microbiome distances and similarity.
ARG load and diversity are higher in low- and middle-income countries, with several predictors showing
context-dependent or reversed associations across income groups. At the country level, resistome dissimilarity
is strongly associated with microbiome composition and socio-economic similarity. Adjustment for microbiome
structure reduces the strength of associations between socioeconomic distance and resistome variation, which
suggests that microbiome composition partly accounts for the association between socio-economic factors and
resistance patterns
These results show that global AMR patterns are structured by interacting socioeconomic and ecological
factors, and demonstrate the value of combining hierarchical and dyadic modelling for analysing
population-scale metagenomic data.
C-G.C.04: Active Learning-Guided Docking for Efficient Discovery of LasR-Targeting Anti-Virulence Compounds
for Diabetic Foot Infections
Track: General computational biology
-
Jonasz Czajkowski, Sano Centre for Computational Medicine, Poland
-
Anna Górska-Ratusznik, Lukasiewicz Research Network—Kraków Institute of Technology, Poland
-
Agata Barzowska-Gogola, Lukasiewicz Research Network—Kraków Institute of Technology, Poland
-
Joanna Budziaszek, Lukasiewicz Research Network—Kraków Institute of Technology, Poland
- Adam Sułek, Sano Centre for Computational Medicine, Poland
- Tomasz Kościółek, Sano Centre for Computational Medicine, Poland
-
Barbara Pucelik, Lukasiewicz Research Network—Kraków Institute of Technology, Poland
Presentation Overview: Show
Active learning-guided docking enables efficient exploration of ultra-large chemical libraries otherwise
inaccessible due to the cost of exhaustive screening. By iteratively selecting the most informative compounds,
it prioritizes candidates with high predicted docking scores and model uncertainty, ensuring balanced
exploration of chemical space. Only a small fraction of molecules is sampled, reducing computational demands
while maintaining strong enrichment of promising candidates.
In this work, the method was applied to LasR, a key regulator of bacterial virulence. Targeting virulence
factors such as LasR is particularly attractive because it may reduce the selective pressure associated with
traditional antibiotics and limit the emergence of resistance. Moreover, the scarcity of known ligands for
LasR makes structure-based approaches especially relevant.
Overall, active learning-guided docking enables scalable and flexible navigation of chemically diverse,
billion-scale libraries, supporting the identification of novel drug candidates and providing a practical
framework for integration with experimental validation workflows. Importantly, experimental validation of
selected candidates confirmed the models predictive capability, highlighting its promise as an effective tool
for accelerating drug discovery.
Importantly, the iterative nature of active learning allows continuous refinement of the model as new data is
incorporated. This closed-loop workflow efficiently allocates computational resources to the most relevant
regions of chemical space and bridges the gap between in silico screening and real-world drug discovery.
Acknowledgements: This work was supported by the National Science Centre (NCN), Poland, under the SONATA-19
grant no. 2023/51/D/NZ7/0259.
C-G.C.05: Grounding AI Agents in Viral Genomics: A Deterministic Retrieval Layer for Reliable Autonomous
Surveillance
Track: General computational biology
-
Ferdous Nasri, Hasso Plattner Institute; Broad Institute of MIT and Harvard, Germany
-
Sarah Gurev, Broad Institute of MIT and Harvard; Department of Biology at MIT; FutureHouse, United States
- Patrick Varilly, Broad Institute of MIT and Harvard, United States
- Krithik Ramesh, Lyra Labs, United States
- Nuala A. O'Leary, National Center for Biotechnology Information, United States
- Jonah Cool, Anthropic, United States
- Bernhard Y. Renard, Hasso Plattner Institute, Germany
-
Pardis C. Sabeti, Broad Institute of MIT and Harvard; Harvard University; Howard Hughes Medical Institute,
United States
-
Laura Luebbert, Broad Institute of MIT and Harvard; FutureHouse; Harvard University, United States
Presentation Overview: Show
As Large Language Models (LLMs) transition from conversational assistants to autonomous research agents, their
ability to interface with primary genomic databases becomes critical. However, the stochastic nature of LLMs
is fundamentally at odds with the deterministic precision required for viral sequence retrieval. When
navigating the multi-dimensional metadata of repositories like NCBI Virus, state-of-the-art agents (including
GPT-5.5, Claude Sonnet 4, Edison and Biomni) frequently "hallucinate" query parameters or fail to navigate
complex web-based filtering logic, leading to corrupted datasets and invalid downstream biological
conclusions.
We present gget virus, a programmatic framework that formalizes the full filtering logic of the NCBI Virus
portal into a deterministic system. To evaluate its impact, we developed VirBench, a curated benchmark of 120
complex virological queries. We demonstrate that unconstrained AI agents achieve a retrieval accuracy of only
16.9-91.3% with high inter-run variability. By equipping these agents with gget virus as a tool-use layer,
accuracy increases to >90% across all systems, while ensuring exact-match semantics. Furthermore, the
framework's "metadata-first" retrieval strategy reduces data transfer by >98% compared to bulk download
methods. By bridging the gap between natural language intent and precise data acquisition, gget virus provides
the necessary infrastructure for reproducible, AI-driven viral discovery and genomic surveillance. The manual
and source code are available at https://scverse.org/gget/en/virus.html and
https://github.com/scverse/gget/blob/main/gget/gget_virus.py, respectively.
C-G.C.06: Exploration of Epitranscriptomic's Landscape through Self-Supervised Language Models
Track: General computational biology
-
Michael Jopiti, University of Bern, Switzerland
- Vincent Jung, EPFL, Switzerland
- Raphaëlle Luisier, University of Bern, Switzerland
Presentation Overview: Show
RNA is increasingly recognized as a central regulator of diverse cellular processes, with non-coding mRNA
regions encoding information that shapes subcellular localization, stability, and translation. This regulation
cannot be explained by individual factors in isolation; it emerges from combinatorial interactions in which
multiple RNA-binding proteins (RBPs) and RNA modifications co-occupy the same transcripts, assembling local
interactomes that vary across transcript regions and cellular states. Mapping this landscape experimentally is
constrained by the scale of the combinatorial space, the cost of specialized assays, and the context-specific
nature of each measurement, leaving existing atlases confined to few contexts. Pretrained RNA Language Models
(RNA-LMs) offer a complementary route: as shown by preliminary work in our group, even without modification
labels their frozen embeddings recover modified nucleotides along with their canonical motifs, indicating they
capture regulatory signal extractable from their representations.
We extend this paradigm to RBPs, treating the latent geometry of RNA-LM representations as a context-agnostic
map of the regulatory landscape that encodes the implicit rules of post-transcriptional regulation. Probing
this signal across model depth and sequence context quantifies how well the representation space captures
known RBP binding preferences. Applied to a broad panel of RBPs profiled by CLIP-based assays across cellular
contexts, the framework reveals a spectrum of recoverability. RNA-LM-derived predictions then feed
transcriptome-wide non-negative matrix factorization (NMF), uncovering latent regulatory programs defined by
co-occurring modifications and RBPs.
C-G.C.07: Hecate: A Modular Genomic Compressor
Track: General computational biology
- Kamila Szewczyk, Algorithmic Bioinformatics, Saarland University, Germany
-
Sven Rahmann, Algorithmic Bioinformatics, Saarland University, Germany
Presentation Overview: Show
We present hecate, a modular lossless genomic compression framework. It is designed around uncommon but
practical source-coding choices.
Unlike many single-method compressors, hecate treats compression as a conditional coding problem over coupled
FASTA/FASTQ streams
(control, headers, nucleotides, case, quality, extras). It uses per-stream codecs under a shared indexed block
container. Codecs include alphabet-
aware packing with an explicit side channel for out-of-alphabet residues, an auxiliary-index Burrows-Wheeler
pipeline with custom arithmetic
coding, and a blockwise Markov mixture coder with explicit model-competition signaling. This architecture
yields high throughput, exact
random-access slicing, and referential mode through streamwise binary differencing. In a comprehensive
benchmark suite, hecate provides
the best compression vs. speed trade-offs against state-of-the-art established tools (MFCompress, NAF, bzip3,
AGC), with notably stronger
behaviour on large genomes and high-similarity referential settings. For the same compression ratio, hecate is
2 to 10 times faster. When given
the same time budget as other algorithms, hecate achieves up to 5% to 10% better compression.
C-G.C.08: Automated Machine Learning to Identify Drivers of Microbial Community Composition
Track: General computational biology
-
Ianis Vilela, Interfaculty Bioinformatics Unit, University of Bern, Switzerland
-
Adamandia Kapopoulou, Theoretical Ecology and Evolution, University of Bern, Switzerland
- Claudia Bank, Theoretical Ecology and Evolution, University of Bern, Switzerland
- Stephan Peischl, Interfaculty Bioinformatics Unit, University of Bern, Switzerland
Presentation Overview: Show
Microbial communities play critical roles across diverse biological systems, yet identifying the environmental
factors that shape their composition remains challenging. Advances in sequencing technologies have increased
data availability, intensifying the need for analytical approaches capable of handling high-dimensional,
compositional microbiome data. While there is strong interest in uncovering multi-factor environmental
drivers, conventional methods such as PERMANOVA, redundancy analysis (RDA), and Mantel tests rely on linear
assumptions that often fail to capture non-linear, threshold-like patterns and complex interactions common in
ecological systems.
Here, we propose an automated, machine learning–based framework to identify environmental drivers of
microbial community structure in a user-friendly and generalizable manner. By relaxing strict model
assumptions, this approach enables detection of complex, non-linear relationships between environmental
variables and community composition.
Simulation studies using synthetic microbial communities demonstrate that the proposed machine learning
framework outperforms traditional methods in models of with cluster like threshold influence patterns from
environmental factors. These results highlight the potential of the proposed machine learning framework, with
an accompanying R package currently under development to facilitate its application to complex datasets.
C-G.C.09: Rotavirus A genotype diversity and antigenic profile in Central Ethiopia: implications for Rotarix®
vaccine efficacy
Track: General computational biology
-
Yisehak Tsegaye Redda, Mekelle University, Ethiopia
- Haileeyesus Adamu, Addis Ababa University, Ethiopia
- Julia Bergholm, Swedish University of Agricultural Sciences, Sweden
- Johanna Lindahl, Uppsala University, Sweden
- Anne-Lie Blomstrom, Swedish University of Agricultural Sciences, Sweden
- Mikael Berg, Swedish University of Agricultural Sciences, Sweden
- Tesfaye Sisay Tessema, Addis Ababa, Ethiopia
Presentation Overview: Show
Introduction:Rotavirus remains a major cause of severe gastroenteritis among children under five years of age
worldwide, including in Ethiopia. Despite the introduction of vaccines, the virus continues to evolve through
frequent mutation and genetic reassortment, resulting in substantial genetic diversity and raising concerns
about potential vaccine escape. This study aimed to assess the prevalence, genotype distribution, and genetic
characteristics of rotavirus A (RVA) strains circulating among children with diarrhea in central Ethiopia, and
to compare these strains with the Rotarix® vaccine strain.
Methods: A cross-sectional study was conducted between April 2022 and December 2023 in health centers located
in Debre Birhan and Addis Ababa. Stool samples were collected from children under five years presenting with
diarrhea. RVA detection was performed using quantitative real-time PCR (qPCR), and genotyping was carried out
through Sanger sequencing of the VP7 and VP4 genes. Amino acid sequences were analyzed to identify
substitutions within key antigenic epitopes in comparison with the Rotarix® vaccine strain.
Results: RVA was detected in 30 out of 247 samples (12.14%), with 28 successfully genotyped. G9 was the
predominant G genotype (50%), followed by G12, G2, G1, and G3, while a proportion remained untyped. Among P
genotypes, P[4] was most common, followed by P[6] and P[8]. The dominant G/P combination was G9P[4]. Notably,
multiple amino acid substitutions were observed in both VP7 and VP4 antigenic regions compared to the Rotarix®
strain. Conclusion: These findings suggest ongoing viral evolution and highlight the need for continuous
surveillance and evaluation of vaccine effectiveness.
C-G.C.10: Population structure of wild and cultivated grapevines in Armenia
Track: General computational biology
-
Maria Nikoghosyan, The Institute of Molecular Biology of the NAS RA, Armenia
- Emma Hovhannisyan, Armenian Bioinformatics Institute, Armenia
- Nate Zadirako, The Institute of Molecular Biology of the NAS RA, Armenia
-
Shengchang Duan, Yunnan Research Institute for Local Plateau Agriculture and Industry, China
- Armine Asatryan, Armenian Bioinformatics Institute, Yerevan, Armenia, Armenia
- Arsen Arakelyan, Institute of Molecular Biology, NAS RA, Armenia
- Kristine Margaryan, Institute of Molecular Biology, NAS RA, Armenia
- Anush Baloyan, Armenian Bioinformatics Institute, Armenia
- Tomas Konecny, Armenian Bioinformatics Institute, Armenia
- Hans Binder, Armenian Bioinformatics Institute, Armenia
Presentation Overview: Show
The South Caucasus region, including Armenia, is recognized as a center of early viticulture, the oldest known
winery, and a tradition of winemaking. Armenia's topography has contributed to the preservation of genetically
diverse grapevine populations. Cultivated grapevines (Vitis vinifera ssp. vinifera) and their wild ancestor
(V. vinifera ssp. sylvestris) exhibit high genetic diversity, making them valuable resources for understanding
domestication, adaptation, and breeding. Despite Armenia's historical and economic importance, the genomic
diversity of its wild and cultivated grapevines remains underexplored. We re-analyzed whole-genome sequencing
data of 164 grapevine accessions from Armenia to characterize genomic diversity, population structure, and
domestication history.
Our analysis uncovered genetic patterns partly unique to Armenia. Population structure analysis revealed a
genetic separation between wild and cultivated groups and three distinct ancestral components within the
cultivated gene pool, reflecting a west-to-east geographical gradient in Armenia. This genetic cline
correlates with a shift in usage, from table to wine grapes, and a transition in berry skin color from white
to black. Additionally, we identified four distinct subgroups within wild populations in Syunik, suggesting
notable diversity. Evolutionary history analysis indicates that wild and cultivated lineages began to separate
~18.5k years ago, with divergence intensifying ~4k years ago under human cultivation. Comparative genomic
scans for divergent selection identified genomic regions associated with domestication traits. GWAS uncovered
candidate markers linked to agronomic traits, such as berry skin color and bunch density.
The results underscore the importance of conserving local grapevine diversity in Armenia, a historically
significant and genetically rich viticultural region.
C-G.C.11: FuFisOr Finder: A Bioinformatics Pipeline for Detecting Orthology Based Fusion and Fission Events
in Protein Sequences
Track: General computational biology
-
Fareha Masood, Center for Synthetic Microbiology (SYNMIKRO), University of Marburg, Germany
-
Paul Klemm, Center for Synthetic Microbiology (SYNMIKRO), University of Marburg, Germany
-
Marcus Lechner, Center for Synthetic Microbiology (SYNMIKRO), University of Marburg, Germany
Presentation Overview: Show
The prediction of homologous genes is fundamental for functional annotation transfer, evolutionary analysis,
and comparative genomics. Orthology inference tools such as Proteinortho, OrthoFinder, SonicParanoid, and
OrthoMCL distinguish orthologs and paralogs but typically overlook gene fusion and fission events, which also
play a key role in protein evolution by reshaping domain architectures and enabling functional
diversification.
We demonstrate that orthology data can be leveraged to detect fusion and fission events, extending its utility
beyond conventional applications. To this end, we developed FuFisOr, a lightweight post-processing pipeline
designed to identify such events from orthology inference outputs, using Proteinortho as an example. The
pipeline detects candidate fission sets aligning to potential fusion proteins by exploiting reciprocal best
hits, avoiding additional similarity searches. FuFisOr integrates seamlessly with Proteinortho, requires
minimal pre-processing, and is adaptable to other tools.
FuFisOr was evaluated on three bacterial proteomes: Pseudomonas aeruginosa, Bacillus subtilis, and Escherichia
coli, identifying 15 high-confidence fusion/fission events validated using InterPro domain annotations and
AlphaFold2 structural models superimposed in UCSF Chimera. One representative case involves the riboflavin
biosynthesis protein RibBA in B. subtilis, which contains both DHBP_synthase and GTP_cyclohydro2 domains; in
E. coli, these exist separately as RibB and RibA, indicating a fission event.
These results highlight FuFisOr ability to detect biologically meaningful events and extend orthology-based
analyses with a scalable, reproducible solution for studying protein evolution.
C-G.C.12: OncoSpat: A Harmonized Pan-Cancer Spatial Transcriptomics Atlas for Quality-Aware Exploration of
Tumor Architecture
Track: General computational biology
-
Tonmoy Das, Institute of Medical Bioinformatics and Systems Medicine, University of Freiburg,
Germany
-
Andreas Tsouris, Institute of Medical Bioinformatics and Systems Medicine, University of Freiburg, Germany
-
Melanie Boerries, Institute of Medical Bioinformatics and Systems Medicine, University of Freiburg,
Germany
-
Sajib Chakraborty, Institute of Medical Bioinformatics and Systems Medicine, University of Freiburg,
Germany
-
Geoffroy Andrieux, Institute of Medical Bioinformatics and Systems Medicine, University of Freiburg,
Germany
Presentation Overview: Show
Spatial transcriptomics is becoming an important approach for studying cancer tissue architecture,
heterogeneity, and microenvironment organization. However, publicly available cancer datasets remain highly
heterogeneous in terms of preprocessing, metadata quality, format variability, which limits their usability
and complicates large-scale, cross-study analyses. To address this, we are developing OncoSpat, a pan-cancer
spatial transcriptomics atlas and interactive portal that currently integrates 1,229 spatial samples from
manually curated public Visium and related in situ-capturing transcriptomics datasets across human and mouse
tumors. OncoSpat is designed as a harmonized and continuously expandable atlas framework rather than a static
collection. At this stage, the project focuses on building a standardized sample registry with unified
metadata, source information, spatial asset availability, and sample readiness information. In parallel, we
are establishing a harmonization workflow for consistent sample formatting, gene identifier standardization,
and quality-aware spatial preprocessing. Planned analysis layers include spot-level quality control, spatial
graph diagnostics, tissue fragmentation, hotspot-based region definition, and co-localization analysis, with
the broader aim of identifying conserved spatial phenotypes and recurring niches across cancers. We are also
developing a scalable Python-based benchmarking module to evaluate data transformation strategies for Visium
spatial transcriptomics at cohort scale. The framework is further being organized as a modular workflow system
with checkpointed analysis stages, explicit quality control review steps, and structured tracking of datasets
and outputs. The first public release is planned for next year, with continued expansion through integration
of newly published public cancer spatial datasets.
C-G.C.13: Influence of Peptide-HLA Binding Modes, Different HLA Alleles, and Epitope Mutations on TCR
Recognition
Track: General computational biology
-
Yan Liu, UNIL, Switzerland
- David Gfeller, UNIL, Switzerland
Presentation Overview: Show
T cells play a crucial role in eliminating cancer cells and pathogen-infected cells, underscoring their
significance in cancer immunotherapy. The functionality of T cells depends on the recognition of peptides
presented by HLA (human leukocyte antigen) molecules via the T-cell receptor (TCR) on the cell surface. This
study aims to elucidate the interaction dynamics between peptide-HLA complexes and TCRs, focusing on the
effects of different binding modes, HLA alleles, and epitope mutations on TCR recognition.
We stimulated human primary T cells with various epitopes to characterize these interactions. Our findings
reveal that the bulge binding mode of peptides induces significant changes in the TCR library. Additionally,
changes in the HLA molecule lead to substantial shifts in TCR recognition patterns. The extent of TCR library
alteration varies depending on the mutation's location within the epitope and the degree of amino acid
difference from the wild-type sequence.
These results not only uncover fundamental principles governing peptide-HLA and TCR interactions but also
provide insights into predicting cross-reactivity between different peptide-HLA molecules. Understanding these
mechanisms enhances our ability to forecast TCR responses to various epitopes, which is crucial for designing
effective cancer immunotherapies and vaccines. This study lays the groundwork for leveraging public TCR
datasets to predict cross-reactivity, potentially reducing the reliance on repetitive functional assays.
C-G.C.14: Copy number aberrations. The challenge of WES, FFPE, tumor only samples.
Track: General computational biology
- Jose Juan Almagro Armenteros, Bristol Myers Squibb, Spain
- Dharanya Sampath, Bristol Myers Squibb, United States
- Elena Buscaroli, Bristol Myers Squibb, Spain
-
Konstantinos Mavrommatis, Bristol Myers Squibb, Switzerland
Presentation Overview: Show
Copy number aberrations (CNAs) are a major class of somatic alterations in cancer and provide clinically
actionable signals through oncogene amplifications and tumor suppressor deletions. Whole-exome sequencing
(WES) is widely adopted in translational and clinical genomics, yet CNA calling from WES remains
difficult—especially for formalin-fixed paraffin-embedded (FFPE) material and tumor-only workflows—because
capture and GC biases, variable mappability, degraded DNA, and unknown purity/ploidy can obscure true copy
number states and inflate false positives, particularly for focal events. In this work, we review the
principles and limitations of WES-based CNA inference and outline practical strategies for normalization,
segmentation, and quality control in challenging sample contexts. We further explore how additional genomic
features (e.g., depth and allelic imbalance summaries, segment-level noise metrics, and concordance across
callers) can be combined with machine learning (ML) approaches to improve confidence estimation and
prioritization of CNA calls. As initial results, we describe a framework for leveraging multi-caller ensembles
and feature-driven scoring to better separate likely biological CNAs from technical artefacts, with an
emphasis on enhancing detection and interpretation of focal deletions and other clinically relevant
aberrations in tumor-only and FFPE WES data.
C-G.C.15: Analysis of current needs, barriers, and formats for computational life sciences training - Swiss
and global views.
Track: General computational biology
- Valeria Di Cola, SIB Swiss Institute of Bioinformatics, Switzerland
- Geert van Geest, SIB Swiss Institute of Bioinformatics, Switzerland
- Monique Zahn, SIB Swiss Institute of Bioinformatics, Switzerland
- Gregoire Rossier, SIB Swiss Institute of Bioinformatics, Switzerland
- Patricia M. Palagi, SIB Swiss Institute of Bioinformatics, Switzerland
-
Diana Marek, SIB Swiss Institute of Bioinformatics, Switzerland
Presentation Overview: Show
Understanding the evolving training needs of researchers in computational life sciences is essential to ensure
that training remains relevant, impactful, and accessible. The SIB Training Group delivers over 60 training
events annually to a diverse audience spanning academia, industry, and healthcare sectors. To guide the future
development of our training portfolio, we conducted a survey to identify emerging scientific topics, preferred
learning formats, and barriers to participation.
Responses were collected across career stages, and scientific domains. A core section of the survey focused on
identifying training topics not currently covered by the SIB portfolio. Using Likert-scale ratings,
respondents evaluated the relevance of prospective topics across multiple areas of computational life
sciences.
A second section examined barriers to training participation, including limited funding, time constraints
during the work week, prerequisite knowledge requirements, and geographical distance. Participants rated the
extent to which each factor affected their ability to attend SIB courses, enabling quantification of
obstacles.
Finally, the survey assessed preferences for learning formats, including multi-day intensive courses,
project-based learning, self-paced e-learning, webinars, blended approaches, AI-assisted learning, multi-week
programmes, flipped classrooms, and bring-your-own-data models. These insights support the development of
flexible and scalable training offerings aligned with researchers' needs and constraints.
The results provide an overview of the Swiss life science training landscape and will directly inform the
design of the 2027 SIB Training Programme. This poster presents the survey design, preliminary analyses, and
implications for future training development, offering a model applicable to other bioinformatics training
initiatives.
C-G.C.16: Depictio: an open-source platform for building interactive dashboards from bioinformatics workflow
outputs
Track: General computational biology
-
Thomas Weber, European Molecular Biology Laboratory, Germany
- Jan Korbel, European Molecular Biology Laboratory, Germany
Presentation Overview: Show
Bioinformatics relies on standardized workflows executed through engines such as Nextflow, Snakemake, or
Galaxy, producing large, heterogeneous outputs across multiple runs. Yet researchers still lack a unified way
to aggregate, explore, and share these results interactively, without writing custom code for each project.
We present Depictio, an open-source web platform that transforms bioinformatics workflow outputs into
interactive, shareable dashboards. Depictio's command-line interface scans workflow output directories,
automatically detects file structures through configurable patterns, and ingests tabular data into a scalable
backend. Data from multiple workflow executions are aggregated into unified, versioned collections ready for
exploration.
A central recent feature is Depictio's template system. Community-driven and peer-reviewed dashboard
templates, currently developed for nf-core pipelines, define both the data ingestion logic and the dashboard
layout for a given workflow. A single CLI command points Depictio at a pipeline's output directory and
automatically generates a fully populated, interactive dashboard without manual configuration required.
Dashboards are organized into multi-tab interactive views, each composed from a rich component library
including interactive components, figures, tables, metric cards, MultiQC plots, image galleries, and
geospatial maps. Cross-component filtering, through lasso, box-select and table selection propagates in real
time across linked components, enabling fluid and live exploratory analysis.
Depictio is cloud-ready, deployable via Docker Compose or Kubernetes, and available as a hosted application on
SciLifeLab Serve (https://serve.scilifelab.se) through a collaboration with SciLifeLab Data Centre. Depictio
is released under the MIT license at https://github.com/depictio/depictio.
C-G.C.17: Fostering Sustainability and Environmental Awareness at Work
Track: General computational biology
- Qingyao Huang, Swiss Institute of Bioinformatics, Switzerland
- Chiara Bortoluzzi, Swiss Institute of Bioinformatics, Switzerland
-
Natasha Glover, Swiss Institute of Bioinformatics, Switzerland
- Christian de Guttry, Swiss Institute of Bioinformatics, Switzerland
- Samuel Neuenschwander, Swiss Institute of Bioinformatics, Switzerland
- Marco Pagni, Swiss Institute of Bioinformatics, Switzerland
- Clément Parisato, Swiss Institute of Bioinformatics, Switzerland
- Charlotte Tumescheit, IDIAP, Switzerland
- Silvano Alda, Swiss Institute of Bioinformatics, Switzerland
Presentation Overview: Show
Do you know that 90% of data is never accessed three months after storage? Do you know that a third of the
8,000 plant types present in Switzerland in the 1990s can no longer be traced? Do you know that the
manufacture of devices often has a higher carbon impact than their usage?
We, the EcoImpact Group, tackle these challenges at SIB. We promote eco-friendly commuting through Bike2Work,
measure our collective carbon footprint to identify areas for improvement, and raise awareness of biodiversity
with BioBlitz events using the iNaturalist app. We participate in Digital Cleanup Day to reduce digital
clutter and will host a Learn@Lunch seminar to share knowledge on planetary boundaries and global
environmental limits. Join us, propose your ideas, and make an impact at SIB and beyond.
C-G.C.18: AD-CLIP: An Adenylation Domain-Substrate Contrastive Representation Learning for Virtual
Screening
Track: General computational biology
-
Andreas Rouvalis, Spoken Language Systems, Saarland Informatics Campus, Saarland University,
Saarbrücken, Saarland, Germany, Germany
-
Kenan Bozhüyük, Department of Drug Bioinformatics, HIPS/HZI, Saarland, Germany, Germany
-
Olga Kalinina, Department of Drug Bioinformatics, HIPS/HZI, Saarland, Germany, Germany
-
Zhengjian Li, Spoken Language Systems, Saarland Informatics Campus, Saarland University, Saarbrücken,
Saarland, Germany, Germany
-
Dietrich Klakw, Spoken Language Systems, Saarland Informatics Campus, Saarland University,
PharmaScienceHub (PSH), Germany, Germany
Presentation Overview: Show
Non-Ribosomal Peptide Synthetases (NRPSs) are modular enzymatic assembly lines that synthesise Non-Ribosomal
Peptides (NRPs), secondary metabolites often used as therapeutics. Within NRPSs, Adenylation (A) domains
dictate substrate specificity, governing the chemical composition of the NRP. Computational efforts have
traditionally framed A-domain specificity prediction as a supervised classification problem, producing
discrete predictions over a closed substrate vocabulary. This is at odds with NRPS research, where novel
specificities are routinely characterised, requiring output-layer extension. We propose AD-CLIP, the first
retrieval-based approach to A-domain specificity, embedding A-domains and substrates in a shared
chemical-protein space where compatibility is a continuous similarity score and substrate vocabulary remains
open. We curate 4,250 A-domain sequences spanning 43 substrates. We use 16 specificity-determining residues,
partially overlapping with canonical Stachelhaus residues, identified in our prior classification work. The
dual encoder pairs an MLP for substrates with an A-domain module of per-residue MLPs, an E(3)-equivariant
graph neural network (EGNN) over Boltz2-predicted Cα coordinates, and mean pooling. We evaluate AD-CLIP using
recall and Normalised Discounted Cumulative Gain (NDCG) under two scenarios: zero-shot retrieval, and n-shot
continual pretraining (n ∈ {1, …, 5}), where pretraining continues on incrementally added examples per
held-out substrate. Zero-shot retrieval achieves mean recall@200 = 0.61 and NDCG@200 = 0.45. AD-CLIP adapts
rapidly under minimal supervision: a single example raises NDCG@200 / recall@200 to 0.74 / 0.85, averaging to
0.97 / 0.91 with five while avoiding catastrophic forgetting. AD-CLIP is thus effective and extensible,
accommodating new substrates at minimal cost.
C-G.C.19: LD-Attention: A Generalized Linkage Disequilibrium Aware Attention Layer for Transformer-Based
Genetic Analysis
Track: General computational biology
-
Omar Abdelwahab, Laval University, Canada
- Davoud Torkamaneh, Laval University, Canada
Presentation Overview: Show
Transformer models have become central to biological data analysis, yet most genomic workflows still
incorporate linkage disequilibrium (LD) through explicit preprocessing, typically relying on pairwise
correlation estimates that are computationally demanding and task-specific. Here, we introduce LD-Attention, a
generalized LD-aware attention layer as a modular and reusable transformer component. By embedding LD directly
within attention, the model learns locus-to-locus dependency structure implicitly during end-to-end training,
removing the need for precomputed LD matrices. In this formulation, LD is not treated as an external feature
but as an inductive bias encoded within the architecture, enabling attention to align with the correlation
structure of genetic variation. The resulting framework is task-agnostic and readily extensible across
LD-dependent analyses, including genotype imputation, genome-wide association studies, fine-mapping, haplotype
phasing, and polygenic risk prediction. By abstracting LD awareness into a portable attention layer,
LD-Attention provides a scalable foundation for integrating LD structure into transformer-based models for
genetic data.
C-G.C.20: Robust Estimation of SARS-CoV-2 Lineage-Specific, Confirmed Case Burden Under Limited Sequencing
Using a Bayesian Framework
Track: General computational biology
-
Mohammad-Hadi Foroughmand-Araabi, Helmholtz Center for Infection Research (HZI), Braunschweig,
Germany., Germany
-
Sama Goliaei, Helmholtz Center for Infection Research (HZI), Braunschweig, Germany., Germany
-
Alice Carolyn McHardy, Helmholtz Center for Infection Research (HZI), Braunschweig, Germany., Germany
Presentation Overview: Show
Continuous genomic surveillance has been central to monitoring SARS-CoV-2 evolution and supporting public
health decisions. Public health authorities, including the World Health Organization (WHO), monitor lineage
dynamics to inform variant designation, including Variants of Interest (VOI). In our previous work, we
developed CoVerage (sarscoverage.org), a large-scale analytics platform that tracks lineage dynamics and
enables early detection of rapidly increasing lineages as potential variants of concern through statistical
analysis of prevalence changes and phylogeny-informed lineage structure. While CoVerage provides
high-resolution insights into lineage frequency trends and antigenic evolution, it relies on sequencing data,
which has recently declined in many regions.
This reduction in sequencing effort introduces uncertainty in estimating population-level prevalence of
circulating lineages, limiting robustness of downstream interpretation for variant monitoring and designation.
To address this gap, we present a Bayesian probabilistic framework that integrates sequencing data with
epidemiological case counts to estimate population-level lineage-attributed, confirmed case numbers under
limited sampling.
We model observed lineage counts as multinomial realizations of latent lineage proportions and place a
symmetric Dirichlet prior over these proportions. Posterior uncertainty is propagated via Monte Carlo
sampling, and sampled proportions are scaled by weekly reported case counts to derive mean lineage-specific
case estimates and 95% highest posterior density intervals.
We implemented this method within CoVerage, providing lineage-specific, confirmed case estimates at global and
regional levels, for example per country. This extension enables uncertainty-aware inference of lineage
dynamics in low-coverage settings, supporting more robust and timely variant monitoring and designation.
C-G.C.21: Oligo Designer Toolsuite – lightweight development of custom oligo design pipelines
Track: General computational biology
- Lisa Barros de Andrade E Sousa, Helmholtz AI, Helmholtz Zentrum München, Germany
- Isra Mekki, Helmholtz AI, Helmholtz Zentrum München, Germany
- Francesco Campi, Helmholtz AI, Helmholtz Zentrum München, Germany
- Louis Kümmerle, Computational Health Center, Helmholtz Zentrum München, Germany
-
Jonas Hagenberg, Helmholtz AI, Helmholtz Zentrum München, Germany
- Chelsea Bright, Computational Health Center, Helmholtz Zentrum München, Germany
- Yarkin Eren, Helmholtz AI, Helmholtz Zentrum München, Germany
- Simon Ament, Hasso-Plattner-Institut, Potsdam, Germany
- Felix Dille, Hasso-Plattner-Institut, Potsdam, Germany
- Louisa Mölm, Hasso-Plattner-Institut, Potsdam, Germany
- Erin Sommer, Hasso-Plattner-Institut, Potsdam, Germany
- Max Timmermann, Hasso-Plattner-Institut, Potsdam, Germany
- Malte Lücken, Computational Health Center, Helmholtz Zentrum München, Germany
- Fabian Theis, Computational Health Center, Helmholtz Zentrum München, Germany
- Marie Piraud, Helmholtz AI, Helmholtz Zentrum München, Germany
Presentation Overview: Show
Oligonucleotides are short, synthetic strands of DNA or RNA that have many application areas. Based on the
intended application and experimental design, researchers have to customize the length, sequence composition,
and thermodynamic properties of the designed oligos. While various tools exist that provide customized oligo
sequences, they each use their own implementation of the basic pipeline steps. This makes it difficult to
create new pipelines. We tackle this issue with our open-source Oligo Designer Toolsuite (ODT). ODT is a
collection of modules that provide all basic oligo design functionalities with a common underlying data
structure within a flexible Python framework. This framework allows a lightweight development of custom oligo
design pipelines. We have successfully developed custom probe design pipelines for cycleHCR, Oligo-seq,
SCRINSHOT, SeqFISH+ and MERFISH protocols with ODT. ODT provides a command line interface, which can be hard
to use by scientists unfamiliar with computational tools. Therefore, we currently develop a web service that
provides a graphical user interface (GUI). ODT-Cloud runs on Helmholtz infrastructure and allows users to
easily set parameters, select target regions and download a table with ready-to-order oligos via a GUI.
ODT-Cloud also handles the automatic preprocessing of reference genomes and stores the results of the pipeline
runs. ODT is available at https://github.com/HelmholtzAI-Consultants-Munich/oligo-designer-toolsuite and
ODT-Cloud will be available at https://odt.helmholtz-munich.de/
C-G.C.22: Wastewater surveillance of endemic human coronaviruses in Switzerland
Track: General computational biology
-
Barbara Mühlemann, D-BSSE, ETH Zürich, Basel, Switzerland; Swiss Institute of Bioinformatics (SIB),
Lausanne, Switzerland, Switzerland
-
Jolinda de Korne-Elenbaas, Eawag, Swiss Federal Institute of Aquatic Science and Technology, Dübendorf,
Switzerland, Switzerland
-
Nadine Hürlimann, Eawag, Swiss Federal Institute of Aquatic Science and Technology, Dübendorf,
Switzerland, Switzerland
-
Daniela Yordanova, Eawag, Swiss Federal Institute of Aquatic Science and Technology, Dübendorf,
Switzerland, Switzerland
-
Adrian Lison, D-BSSE, ETH Zürich, Basel, Switzerland; Swiss Institute of Bioinformatics (SIB), Lausanne,
Switzerland, Switzerland
-
Lara Jeworowski, Institute of Virology, Charité - Universitätsmedizin Berlin, Berlin, Germany, Germany
-
Victor Corman, Institute of Virology, Charité - Universitätsmedizin Berlin; Labor Berlin–Charité
Vivantes GmbH, Berlin, Germany, Germany
-
Timothy R. Julian, Eawag, Swiss Federal Institute of Aquatic Science and Technology, Dübendorf,
Switzerland, Switzerland
-
Tanja Stadler, D-BSSE, ETH Zürich, Basel, Switzerland; Swiss Institute of Bioinformatics (SIB), Lausanne,
Switzerland, Switzerland
Presentation Overview: Show
Endemic human coronaviruses (HCoV) cause seasonal peaks of respiratory disease. As infections with endemic
coronaviruses are usually mild, symptoms-based surveillance may provide an incomplete characterisation of the
circulation of these viruses. This study evaluates the efficacy of wastewater-based epidemiology as a tool to
monitor the circulation of endemic human coronaviruses HCoV-229E, HCoV-NL63, and HCoV-HKU1 in Zurich,
Switzerland.
A multiplex digital PCR assay was used to quantify viral RNA in wastewater samples collected between December
2021 and December 2024 in Zurich, Switzerland. Using EpiSewer, a Bayesian semi-mechanistic model, we estimated
the effective reproduction number and characterized epidemic waves. The findings were compared against
clinical data from the Swiss Sentinel reporting system (Sentinella).
Aggregated weekly flow-normalised viral concentrations correlated with the weekly number of clinical cases
reported through the Sentinella surveillance system, although the overall reported number of weekly cases was
low. Circulation generally peaked in the first quarter of the year, though peaks occurred later in the season
compared to pre-pandemic trends observed in the literature. HCoV-NL63 showed increased activity in early 2023
and late 2024, while HCoV-229E and HCoV-HKU1 showed peaks in winter 2022, 2023, and 2024. HCoV-229E circulated
even during periods of non-pharmaceutical interventions due to the SARS-CoV-2 pandemic.
In summary, we demonstrate that wastewater surveillance can offer a robust, population-level approach to
tracking endemic coronaviruses independent of healthcare-seeking behavior and can provide critical insights
into the evolving seasonality of respiratory viruses in the post-pandemic era.
C-G.C.23: Predicting Molecular Hotspots for Efficient Drug Discovery using Deep Learning
Track: General computational biology
-
Inken Fender, University of Bern, Switzerland
- Thomas Lemmin, University of Bern, Switzerland
Presentation Overview: Show
Protein-ligand interactions play a central role in many biological processes ranging from cell signaling to
disease treatment. However, traditional drug discovery struggles with the immense number of potential drug
molecules and the limited availability of detailed protein-ligand complex data. This bottleneck hinders the
development of highly selective drugs with strong binding affinity to their target proteins. Recent research
has opened an exciting new avenue by demonstrating the potential to predict the most probable orientations and
positions of chemical groups solely from the local backbone conformation and the identity of the interacting
amino acid. This eliminates the need for scarce protein-ligand complex data, offering a significant advantage
in drug design.
Our research builds upon this promising avenue and takes a significant step forward by extending the concept
to the broader protein micro-environment. We propose formulating the identification of functional hotspots
within these micro-environments as a semantic segmentation task. Our method harnesses a convolutional neural
network (CNN) trained on a meticulously curated database, encompassing a diverse array of protein
micro-environments surrounding various functional groups. To assess the efficacy of our approach, we evaluated
the predicted hotspots based on their pharmacological relevance and compared them to existing protein-ligand
complexes. This evaluation demonstrates the significant potential of our deep learning approach to expedite
the drug discovery process by prioritizing promising target sites for further investigation.
C-G.C.24: Translating omics data into discovery: the Biomedical Data Science Facility
Track: General computational biology
-
Tane Kafle, Biomedical Data Science Facility, University of Geneva & SIB Swiss Institute of
Bioinformatics, Switzerland
-
Tania Wyss, Biomedical Data Science Facility, University of Geneva & SIB Swiss Institute of
Bioinformatics, Switzerland
-
Nadine Fournier, Biomedical Data Science Facility, University of Geneva & SIB Swiss Institute of
Bioinformatics, Switzerland
Presentation Overview: Show
The Biomedical Data Science Facility (BDSF) provides bioinformatics services to biomedical research groups in
the Lemanic region (largely UNIGE, HUG, UNIL, CHUV and EPFL). Our platform handles a comprehensive range of
omics data, including single-cell RNA-sequencing, ATAC-sequencing, spatial transcriptomics, and metabolomics.
Beyond analysis, we guide researchers in technology selection, statistical method implementation and data
interpretability and sharing.
In this poster, we showcase projects that illustrate how we help transform complex biological data into
meaningful scientific discoveries. These projects highlight the computational methods and analyses that our
facility provides to the local scientific community. For example, we develop custom visualization tools, such
as the B2CViz R package for spatial gene expression, and build interactive Shiny web applications (e.g.,
BrainTIME), which allow researchers to explore their own datasets without requiring advanced coding skills.
Additionally, in collaboration with researchers and clinicians, the facility implements machine learning
models for biomarker discovery. Finally, we also leverage the richness of publicly available data, which
allows researchers to translate large-scale raw data into interpretable biological discovery.
Overall, the BDSF acts as a collaborative partner in navigating the full lifecycle of bioinformatics research,
from technology selection to result interpretation. In addition to data analysis, we also offer training,
including open-source courses in R, statistics, and omics data analysis, alongside tailored support for Ph.D.
and postdoctoral researchers. We invite you to visit our poster to explore our projects in depth, learn more
about our services, or discuss specific bioinformatics challenges in your own research.
C-G.C.25: The MicrobeAtlas database in 2026: an expanded view of the global microbial ecosphere
Track: General computational biology
-
Janko Tackmann, University Zurich, Switzerland
- David Patsch, University Zurich, Switzerland
- Daniela Gaio, University Zurich, Switzerland
- Lukas Malfertheiner, University Zurich, Switzerland
- Eugenio Perez-Molphe-Montoya, University Zurich, Switzerland
- Matteo Eustachio Peluso, University Zurich, Switzerland
- Nicolas Näpflin, University Zurich, Switzerland
- Marija Dmitrijeva, ETH Zurich, Switzerland
- Shinichi Sunagawa, ETH Zurich, Switzerland
- Christian von Mering, University Zurich, Switzerland
Presentation Overview: Show
Earth is teeming with microbial life, pervading all layers of the planet – from the deep subsurface to the
upper atmosphere. Although characterizing this diversity remains a formidable challenge, microbial ecosystems
are being explored at an unprecedented pace, continually revealing new taxa and communities. MicrobeAtlas
integrates these studies into a unified, global resource for comparisons across environments, hosts, and
technologies.
Here, we report on the latest updates and feature additions to MicrobeAtlas. The database now connects more
than 5.3 million microbial sequencing samples into a unified dataset, substantially expanding environmental
coverage and doubling community-level diversity. Using an improved, scalable workflow, these communities were
clustered into tens of thousands of community types at adjustable resolution, enabling a bird's-eye view on
global microbial ecosystems. Large language models now support the metadata extraction pipeline, improving
environmental annotations, generating detailed keyword summaries, and enabling quantitative comparisons
through semantic embeddings. Integration of the mOTUs4 and proGenomes4 databases, alongside existing links to
BacDive, provides deeper functional and genomic context for lineages in MicrobeAtlas.
Ongoing efforts on reference database extension, strain-level biogeographic analyses, and microbial eukaryote
integration will provide an increasingly comprehensive understanding of our microbial planet.
C-G.C.26: Tumour evolution as ground truth for cancer whole-genome sequencing
Track: General computational biology
-
Lucrezia Valeriani, University of Trieste, AREA Science Park, Italy
-
Giorgia Gandolfi, University of Trieste, IRCCS San Raffaele Scientific Institute, Italy
- Elena Buscaroli, University of Trieste, Italy
-
Katsiaryna Davydzenka, International School of Advanced Studies, University Campus Bio-Medico of Rome,
Italy
- Giovanni Santacatterina, University of Trieste, Italy
- Alice Antonello, University of Trieste, Italy
- Azad Sadr Haghighi, University of Trieste, Italy
-
Virginia Anna Gazziero, University of Trieste, Fondazione Toscana Life Sciences, Italy
- Salvatore Milite, Centre for Computational Biology, Human Technopole, Italy
- Elena Rivaroli, University of Trieste, Italy
- Anna Kabanova, Fondazione Toscana Life Sciences, Italy
- Guido Sanguinetti, International School of Advanced studies, Italy
- Alessio Ansuini, AREA Science Park, Italy
- Leonardo Egidi, University of Trieste, Italy
- Stefano Cozzini, AREA Science Park, Italy
- Alberto Cazzaniga, AREA Science Park, Italy
-
Giovanni Tonon, IRCCS San Raffaele Scientific Institute-Vita-Salute San Raffaele University, Italy
- Trevor Graham, Institute of Cancer Research, UK
- Andrea Sottoriva, Centre for Computational Biology, Human Technopole, Italy
- Riccardo Bergamin, University of Trieste, Italy
- Nicola Calonaci, University of Trieste, Italy
- Alberto Casagrande, University of Udine, Italy
- Giulio Caravagna, University of Trieste, Italy
Presentation Overview: Show
Cancer genomes are shaped by evolutionary processes that couple mutagenesis, clonal selection, chromosomal
instability, spatial growth and treatment response into structured genomic patterns, yet current benchmarking
strategies largely ignore this evolutionary dependency. Here, we present SCOUT, a large-scale synthetic
whole-genome sequencing resource of over 200 samples, designed for systematic benchmarking of tumour genomic
analysis and evolutionary inference under controlled evolutionary ground truth. Unlike conventional
task-specific simulations, SCOUT models tumour evolution as a latent generative process that simultaneously
shapes mutations, copy-number alterations, variant allele frequencies, mutational signatures and clonal
architectures. SCOUT recapitulates key features of solid and haematological malignancies, including driver
mutations, chromosomal instability, intratumour heterogeneity, spatial sampling and treatment-associated
evolutionary dynamics in tumour and matched-normal longitudinal and multi-region sequencing designs. Using
SCOUT, we benchmarked widely used methods for somatic variant detection, copy-number analysis, mutational
signature inference and tumour evolutionary reconstruction. Across analytical tasks, performance deteriorated
in low-purity, highly subclonal and structurally complex tumours, while spatial sampling bias and
hypermutation generated spurious evolutionary signals that confounded tumour interpretation across multiple
inference layers. Evolutionary simulations further distinguished lineage-restricted genetic bottlenecks from
multi-lineage resistance dynamics associated with tumour plasticity. Tumour purity consistently exerted a
stronger effect on inference accuracy than sequencing depth. Together, our results establish evolutionary
ground truth as a prerequisite for reproducible benchmarking and biologically interpretable analysis of cancer
whole-genome sequencing data.
C-G.C.27: CLUES A Comprehensive Workflow for Integrating Geospatial Data in Biomedical Research
Track: General computational biology
-
Marcel Jentsch, Charite BIH, Germany
- Sven Twardziok, BIH@Charite, Germany
- Elli Polemiti, BIH@Charite, Germany
Presentation Overview: Show
Environmental exposures play a critical role in shaping physical and mental health, yet integrating such data
into biomedical research remains technically complex and fragmented. The EnvironMENTAL Climate, Urbanicity,
Environment and Society (CLUES) framework is an open-source, end-to-end workflow for generating
individual-level environmental exposure data. CLUES automates the selection and download of open-access
geospatial datasets, standardises spatial and temporal formats, and map projections, and links resulting
environmental variables to individual-level biomedical data, requiring no prior expertise in geospatial data.
CLUES covers key environmental domains, including urban and natural space, climate and weather extremes, air
pollution, and regional socioeconomic conditions. Designed for extensibility and cross-cohort applicability,
it enables multidimensional exposure mapping across global settings and adheres to FAIR (Findability,
Accessibility, Interoperability and Reusability) and privacy-compliant data protection principles. In this
work, we present the CLUES framework and evaluate its scalability, computational performance, and
reproducibility for large-scale biomedical research.
C-G.C.28: Geometry of Antigenic Space Enables Optimization of HCV Vaccine Antigen Design
Track: General computational biology
-
Sama Goliaei, Helmholtz Center for Infection Research (HZI), Braunschweig, Germany., Germany
-
Mohammad-Hadi Foroughmand-Araabi, Helmholtz Center for Infection Research (HZI), Braunschweig, Germany.,
Germany
- Dorothea Bankwitz, Experimental Virology, TWINCORE, Hannover, Germany., Germany
- Julie Sheldon, Experimental Virology, TWINCORE, Hannover, Germany., Germany
- Mandy Döpke, Experimental Virology, TWINCORE, Hannover, Germany., Germany
- Corinne Ginkel, Experimental Virology, TWINCORE, Hannover, Germany., Germany
-
Birthe Reinecke, Institute for Experimental Virology, TWINCORE, Hannover, Germany., Germany
- Ye Ke, Institute for Biochemistry, Luebeck, Germany., Germany
- Thomas Pietschmann, Experimental Virology, TWINCORE, Hannover, Germany., Germany
-
Alice Carolyn McHardy, Helmholtz Center for Infection Research (HZI), Braunschweig, Germany., Germany
Presentation Overview: Show
Hepatitis C virus (HCV) vaccine development is fundamentally constrained by extreme genetic and antigenic
diversity, which complicates rational antigen selection and the design of antigens inducing potent and broad
neutralizing responses. Here, we apply a quantitative geometric framework to analyze and predict
cross-neutralizing antibody responses based on antigen positioning in antigenic space.
Using a previously established antigenic cartography of virus-antibody interactions, we consider a
low-dimensional embedding in which distances reflect functional neutralization similarity. Within this space,
we introduce the concept of an antigenic center and derive distance-based metrics to characterize candidate
immunogens and their combinations. We further integrate neutralization breadth and potency into a composite
functional score, the cross-neutralization index (CNI), enabling consistent ranking across heterogeneous
vaccine platforms.
Applying this framework to diverse recombinant E2 immunogens, we show that some, but not all, antigens closer
to the antigenic center yield stronger cross-neutralizing responses, possibly linking spatial position to
functional outcome. Moreover, viral antigen features independent of antigen cartography may influence antigen
immunogenicity. Notably, rationally selected minimal antigen combinations that effectively cover this space
outperform larger multi-cluster antigen pools, demonstrating that coverage--not antigen count--governs
performance.
Anchored to human benchmarks of vaccine efficacy and HCV humoral immune protection, the framework provides a
bridge computational modeling and translational vaccine evaluation. Together, this work demonstrates the use
of a geometry-based approach for HCV vaccine antigen design, which may facilitate optimization of immunogens.