The ongoing revolution in genome sequencing, driven by initiatives such as the Earth BioGenome Project and the
Darwin Tree of Life, is delivering an unprecedented number of high-quality genome assemblies to global
repositories such as the International Nucleotide Sequence Database Collaboration (INSDC). Each genome is
imported to the UniProt database as a proteome, i.e. the set of all translated proteins for that genome. To
meet the challenge of integrating this wealth of biodiverse data, whilst maintaining data quality and
usability for the scientific community, UniProt undertook major improvements to its data content and curation
pipelines. A critical step in this effort is the selection of reference proteomes, the proteome that best
represents the protein space for that species. A new pipeline was developed, using a clustering system based
on MMseqs2, to select reference proteomes which maximizes the protein space for each species, whilst
minimizing the number of chosen proteomes. Additionally, we aligned our viral reference proteomes with the
genome set defined by the International Committee on Taxonomy of Viruses (ICTV). As a result of these
improvements, the UniProt Knowledgebase (UniProtKB) saw a 36% increase in reference proteomes, covering 34%
more species in the Tree of Life, whilst simultaneously reducing the overall number of proteins by 41%,
leading to a more representative knowledgebase. Proteins deleted from UniProtKB still have its protein
sequences present in UniParc, now being annotated by UniFIRE, an automatic annotation software based on
UniRules and ARBA. This work marked a major step forward in our mission to provide the scientific community
with accurate, comprehensive, and accessible protein data, tailored to the rapidly evolving landscape of
biodiversity genomics.