Skip to main content

Write a PREreview

Benchmark Label Definition, Not Model Capability, Explains a Machine-Learning Advantage over TD-DFT for λmax Prediction

Posted
Server
ChemRxiv
DOI
10.26434/chemrxiv.15008972/v1

Machine-learning models trained on literature-mined UV-Vis data are frequently reported to outperform time-dependent density functional theory (TD-DFT) for predicting absorption maxima (λ max ). Such comparisons usually evaluate the model on held-out molecules from the same corpus that supplied its training labels, and they set experimental labels of heterogeneous transition type against computed vertical excitations of a single electronic transition. Using a Gaussian process (GP) with 256-bit Morgan fingerprints, trained on 6,878 curated compounds from the Beard literature-mined database, the study shows that two features of the benchmark, rather than model capability, govern the reported margin. On leakage-free in-corpus holdouts the GP outperforms the quantum-chemical references (RMSE 87.3 vs 108.7 nm against TD-DFT, n = 36; 81.6 vs 117.1 nm against sTDA, n = 1,011), and the advantage survives Bemis–Murcko scaffold splitting. Applied unchanged to genuinely external corpora, the same model keeps working where the label convention is compatible with training (R² = +0.209 on 5,811 compounds) and fails where it is not (R² = −0.269 with a +53.5 nm offset on 318 photoswitches labelled specifically to the E -isomer π→π* transition); chemical novelty alone does not break it, reassigning the label does. Against published ωB97X-D3 vertical excitations on 5,501 training-disjoint compounds, the calculations are the more accurate predictor once each method’s removable offset is discounted (≈58 vs 95 nm residual scatter) and correlate with experiment nearly twice as strongly (r = 0.85 vs 0.46). On the transition-specific corpus, CAM-B3LYP/6-31G** reaches 25.4 nm against the GP’s 80.9 nm. The 57–65 nm quantum-chemical underestimation seen in-corpus reappears out of sample at −61.9 nm against independently computed values, supporting a definitional rather than a functional origin. GP predictive uncertainty is poorly calibrated in all external regimes and does not reliably identify high-error predictions; its apparent correlation with error reflects the label variance of the evaluation set rather than structural distance from the training distribution. For this model class, reported margins over quantum chemistry should therefore be regarded as properties of the benchmark—its label definition and its same-corpus provenance,not of the method, unless the benchmark is external and its label convention is stated.

You can write a PREreview of Benchmark Label Definition, Not Model Capability, Explains a Machine-Learning Advantage over TD-DFT for λmax Prediction. A PREreview is a review of a preprint and can vary from a few sentences to a lengthy report, similar to a journal-organized peer-review report.

Before you start

We will ask you to log in with your ORCID iD. If you don’t have an iD, you can create one.

What is an ORCID iD?

An ORCID iD is a unique identifier that distinguishes you from everyone with the same or similar name.

Start now