Key Findings
A preprint submitted to ChemRxiv on September 9, 2026, details a comparative study on the performance of various machine learning models—tree models, ChemBERTa, graph networks—and a Late Fusion approach for plasma protein binding (PPB) prediction under challenging “scaffold split” conditions. This research aims to enhance the accuracy of drug candidate selection during early drug discovery, thereby improving development efficiency.
Technical Details
- Importance of Plasma Protein Binding (PPB): PPB is a critical pharmacokinetic (PK) property that determines how a drug distributes and exerts its effects within the body. Highly bound drugs may not reach effective concentrations at their target sites, leading to reduced efficacy. Accurate PPB prediction is therefore essential for early-stage screening and optimization of drug candidates in the drug discovery pipeline.
- Challenges of Scaffold Split: While traditional PPB prediction models perform well with “random split” data (where chemically similar molecules are in both training and test sets), “scaffold split” presents a more rigorous challenge. It involves predicting for molecules with different chemical scaffolds, preventing models from overfitting to known structures and assessing their true generalization performance on novel molecules. This is a crucial benchmark for discovering genuinely new drug candidates.
- Applied Machine Learning Models:
- Tree Models: Powerful models like XGBoost and Random Forest, effective for tabular data and molecular descriptors.
- ChemBERTa: A transformer-based language model trained on large chemical text datasets (e.g., SMILES strings), excelling at capturing semantic information about molecules.
- Graph Networks: Models (e.g., GNNs) that represent molecules as graph structures, learning directly from the relationships between atoms (nodes) and bonds (edges). They effectively capture structural molecular features.
- Late Fusion: An approach that integrates the prediction outputs from different models to combine their individual strengths, thereby further improving overall prediction performance.
Background & Context
Drug discovery success rates are notoriously low, and development costs continue to soar. A significant factor contributing to this is inadequate prediction of pharmacokinetic (ADME) and toxicity (Tox) properties during preclinical stages. Improving the accuracy of predicting critical PK properties like PPB is crucial for reducing late-stage failures and identifying more promising drug candidates earlier. Rigorous evaluation using scaffold split, as demonstrated in this research, is vital for assessing the “true” generalization capability of AI drug discovery models and contributes to overall industry efficiency.
Strategic Significance & Outlook
The insights gained from this research will directly translate into more efficient lead compound selection and optimization within drug discovery pipelines, reducing development time and costs. In the future, such robust prediction models could be leveraged not only for pharmaceuticals but also as tools to accelerate safety and efficacy assessment processes in the development of new functional materials, agrochemicals, cosmetic ingredients, and more. Further optimization of late fusion approaches and their application to a wider range of ADME/Tox property predictions are anticipated.
Source: https://chemrxiv.org/browse?pageSize=20&startPage=1
Get our weekly technology intelligence — free
Receive an infographic that lets you judge at a glance whether each field’s analysis report is worth reading.
Subscribe Free — Weekly Tech Intelligence
By subscribing, you’ll receive Troy-Technical’s weekly technology intelligence newsletter.
- Your email and selected fields are used only to deliver the newsletter.
- We never share your information with third parties.
- You can unsubscribe anytime via the link in each email.
See our Privacy Policy for details.
Takes about a minute · Unsubscribe anytime

Comments