View posters by category

Scroll down to view Results

Session A: Monday 31 August 12:00-13:30
Session B: Tuesday 1 September 16:15-17:45
Session C: Wednesday 2 September 11:30-13:00

Results

C-G.C.29: Reference proteomes in UniProt
Track: Posters
  • Pedro Raposo
  • Daniel Rice
  • Minjoon Kim
  • UniProt Consortium


Presentation Overview: Show

The ongoing revolution in genome sequencing, driven by initiatives such as the Earth BioGenome Project and the Darwin Tree of Life, is delivering an unprecedented number of high-quality genome assemblies to global repositories such as the International Nucleotide Sequence Database Collaboration (INSDC). Each genome is imported to the UniProt database as a proteome, i.e. the set of all translated proteins for that genome. To meet the challenge of integrating this wealth of biodiverse data, whilst maintaining data quality and usability for the scientific community, UniProt undertook major improvements to its data content and curation pipelines. A critical step in this effort is the selection of reference proteomes, the proteome that best represents the protein space for that species. A new pipeline was developed, using a clustering system based on MMseqs2, to select reference proteomes which maximizes the protein space for each species, whilst minimizing the number of chosen proteomes. Additionally, we aligned our viral reference proteomes with the genome set defined by the International Committee on Taxonomy of Viruses (ICTV). As a result of these improvements, the UniProt Knowledgebase (UniProtKB) saw a 36% increase in reference proteomes, covering 34% more species in the Tree of Life, whilst simultaneously reducing the overall number of proteins by 41%, leading to a more representative knowledgebase. Proteins deleted from UniProtKB still have its protein sequences present in UniParc, now being annotated by UniFIRE, an automatic annotation software based on UniRules and ARBA. This work marked a major step forward in our mission to provide the scientific community with accurate, comprehensive, and accessible protein data, tailored to the rapidly evolving landscape of biodiversity genomics.