B-B.09: Genome context interpretation of dark protein functions with metaGCsnap
Over the past two decades, metagenomic sequencing has enabled the exploration of vast and novel regions of the protein universe, revealing an immense diversity of sequences from a wide range of environments. However, most proteins collected from metagenomic experiments have low to no similarity to proteins with experimentally characterized function;
Even the most extensively characterized biome, the human gut microbiome (HGM), resists full-scale functional annotation. HGM-derived sequence catalogs have identified, with homology methods, over 170 million putative proteins, but nearly 40% of these remain functionally uncharacterized, representing a large reservoir of biological ""dark matter"". Current functional annotation methods are even less effective for other biomes.
Compared to structural alignment and deep learning, genome-context-based methods offer a complementary approach to curating the functional annotation of metagenomic data. Nevertheless, there is a lack of tools to gather and organise these data on a large scale in a customisable way.
Here, we introduce metaGCsnap, a tool that leverages comparative genomics by analyzing large repositories of isolated and metagenomic assemblies, allowing to elucidate the functions of uncharacterized proteins. Unlike current methods, metaGCsnap's native integration with MGnify enables it to capture protein occurrence by linking sequences with environmental samples, paving the way for the investigation of sequence diversity and environmental coupling.
By enabling the systematic exploration of uncharted protein space, this study aims to accelerate functional discovery by amplifying scientist capabilities to interrogate their genomic data.
Co-authors: Joana Soares Pereira, Torsten Schwede
Contact Attendee
Warning: Attempt to read property "user_email" on string in /home/1276969.cloudwaysapps.com/ydbgzhdjeq/public_html/wp-content/plugins/my-conference-now/functions.php on line 2826