C-P.19: Generating tailored high-quality datasets for benchmarking structure-based computational drug design tools: a docking assessment
Authors
Ute F. Röhrig
Swiss Institute of Bioinformatics
Vincent Zoete
Swiss Institute of Bioinformatics, University of Lausanne
Keywords
Molecular docking; benchmarking; structure-based drug design
[sponser-meet-now-chat][/sponser-meet-now-chat]
Structural bioinformatics is critical for drug discovery, as it provides methods and tools to predict, analyze, and validate 3D data of macromolecules. The ELIXIR 3D-BioInfo community is organized around five main activities, one of which aims to create the tools necessary for developing large-scale, high-quality, and sustainable datasets of ligand-protein complexes. These datasets will be used to assess and benchmark structure-based computer-aided drug design (SB-CADD) algorithms, such as docking, virtual screening and binding site detection tools.
We developed three interconnected Nextflow pipelines capable of constructing datasets from complex, protein, or ligand PDB identifiers. Data sources include multiple APIs and dedicated software developed by ELIXIR participating groups for data retrieval and processing.
The pipelines generate multifaceted datasets with comprehensive annotations. Complex analysis involves retrieving structural details including proteins, bound ligands and their interactions, as well as quality metrics such as resolution, missing atoms, crystal contacts, and electronic density support from the Protein Data Bank. Coordinate files are retrieved for complexes and ligands, as well as binding affinity, refined structures, tautomers, and protonation states. The protein characterization pipeline extracts UniProt KB descriptions, associated 3D structures, bound ligands, binding sites, active and inactive molecules, and protein flexibility. For ligands, molecular properties, 3D conformers, partial charges, and associated complex structures are provided.
We present benchmark sets generated from complex, protein, or ligand identifiers, along with an initial docking assessment. Docking with AutoDock Vina, GNINA, and SMINA provides a first evaluation of SB-CADD tool performance on the curated, high-quality data.
Co-authors: Ute F. Röhrig, Vincent Zoete
Contact Attendee
Warning: Attempt to read property "user_email" on string in /home/1276969.cloudwaysapps.com/ydbgzhdjeq/public_html/wp-content/plugins/my-conference-now/functions.php on line 2826