Reproducible cleaning and standardization of SMILES-based chemical datasets
-
Updated
Apr 15, 2026 - Python
Reproducible cleaning and standardization of SMILES-based chemical datasets
R is nowadays probably the most powerful tool for calculations of all kinds. There are plenty of modules available for work with molecular data. Those will be introduced during the course.
Multimodal Integration with Modality-agnostic Autoencoders - Developed by LMIB @ KU Leuven
A Python package for diagnostic assessment of data consistency in molecular datasets
A small utility for parsing PDB files into useable JSON
Supplementary material for the paper The Visual Story of Data Storage: From Storage Properties to User Interfaces, by Aleksandar Anžel, Dominik Heider, and Georges Hattab
CSV templates for preparing molecular data submission via GFBio
Fortran code to investigate the Thermodynamics of PAH stepwise-hydrogenation reaction
Explainability study of Graph Neural Networks for fraud detection (DGraphFin) and molecular property prediction (B-XAIC), using multiple GNN explainers.
Curated chemistry and medicinal chemistry datasets for machine learning drug discovery.
A collection of the Chemical JSON files ive made
Open Babel workflow for molecular file conversion, optimization, and cheminformatics analysis.
⌬ Predictive Modeling of Pharmacological Activites
To associate your repository with the molecular-data topic, visit your repo's landing page and select "manage topics."