Comentarios
Escribir un comentarioNo se han publicado comentarios aún.
This manuscript addresses an important and underappreciated problem in computational chemistry and molecular machine learning: apparent differences in predictive performance can arise from benchmark construction and target definition rather than from intrinsic superiority of one modelling approach over another. The distinction between literature-reported absorption maxima and calculated vertical excitation energies is scientifically important, and the emphasis on genuinely external validation is a clear strength of the work.
However, I find the central conclusion stronger than the evidence presently described appears to support. The results convincingly show that model performance changes substantially across datasets and label conventions. They do not yet demonstrate that label definition, rather than model capability, is the principal cause of the apparent machine-learning advantage over TD-DFT.
Several factors change simultaneously across the comparisons presented in the manuscript, including chemical domain, dataset provenance, transition assignment, computational protocol, sample size, and the nature of the machine-learning representation. These factors make it difficult to isolate label definition as the causal explanation.
The manuscript raises a valuable methodological concern, but the title and several conclusions should either be moderated or supported by additional controlled analyses.
The title states that benchmark label definition, “not model capability,” explains the machine-learning advantage over TD-DFT. This is a strong exclusionary claim.
The results show that the apparent ranking between a Gaussian-process model and quantum-chemical calculations changes between evaluation datasets. However, label definition is not the only property that changes between these datasets. Their chemical composition, source literature, distribution of chromophores, spectral range, experimental conditions, and possibly computational protocols also differ.
Therefore, demonstrating that a GP performs differently on a dataset with transition-specific labels does not by itself establish that label definition is the cause of the difference.
A substantially stronger test would use the same set of molecules with two alternative target definitions—for example, a literature-reported overall λmax and an independently assigned transition-specific λmax—and evaluate both methods against both labels. A paired analysis of this type would isolate label definition while holding molecular distribution constant.
Unless such a controlled analysis is provided, I suggest moderating the central conclusion to state that label definition and dataset provenance are important contributors to the reported ML advantage rather than demonstrating that they fully explain it.
The machine-learning side of the comparison is represented by a Gaussian process using 256-bit Morgan fingerprints. This is a relatively specific and comparatively simple model representation.
The conclusions therefore apply directly to this GP/fingerprint model, but not necessarily to molecular machine learning more broadly.
Modern λmax prediction methods may use larger fingerprints, graph neural networks, message-passing architectures, pretrained molecular representations, three-dimensional information, explicit solvent descriptors, or quantum-chemical features. Such models may respond differently to both chemical and label distribution shifts.
This issue is particularly important because the title contrasts “model capability” with benchmark effects in general terms.
I recommend evaluating several substantially different ML approaches or restricting the conclusions explicitly to the model class studied here.
The use of only 256 fingerprint bits also deserves justification. Such a short fingerprint may produce substantial bit collisions for chemically diverse datasets. Performance should ideally be compared with conventional 1024- or 2048-bit Morgan fingerprints to establish that the observed transfer behaviour is not partly a consequence of the selected molecular representation.
The quantum-chemical reference data appear to originate from multiple published sources and include TD-DFT, sTDA, ωB97X-D3, and CAM-B3LYP/6-31G** calculations.
These are not interchangeable computational protocols.
Differences in functional, basis set, molecular geometry, conformer selection, solvent model, treatment of protonation or tautomerism, and transition assignment can all substantially affect calculated excitation energies.
Consequently, describing the comparison simply as “machine learning versus TD-DFT” risks combining methodological heterogeneity with model-class differences.
A stronger comparison would apply a single, clearly defined quantum-chemical workflow to the same molecules used for ML evaluation. At minimum, performance should be stratified by computational protocol rather than interpreted as representing a single quantum-chemical method.
The reported in-corpus TD-DFT comparison contains only 36 compounds.
An RMSE difference of 87.3 versus 108.7 nm may appear substantial, but with n = 36 the uncertainty around both values could also be substantial, particularly if a small number of molecules have very large errors.
Confidence intervals or bootstrap distributions for the difference in RMSE should therefore be reported.
More generally, model comparisons should be paired whenever predictions are available for the same molecules. Reporting distributions of paired absolute-error differences would be more informative than comparing headline RMSE values alone.
The larger sTDA comparison is more convincing statistically, but it answers a somewhat different computational question and should not simply be treated as confirming the TD-DFT result.
A particularly important issue is the use of RMSE in nanometres.
Excitation energy and wavelength are related non-linearly:
E ∝ 1/λ.
Therefore, a 50 nm error at 250 nm does not represent the same energetic error as a 50 nm error at 650 nm. RMSE measured in nm gives disproportionate influence to errors occurring at longer wavelengths and complicates comparisons across datasets with different spectral ranges.
This is especially relevant here because different external datasets may contain substantially different wavelength distributions.
The main comparisons should therefore also be reported in energy units, preferably eV or cm⁻¹. Conclusions that depend strongly on whether errors are expressed in nm or eV would require careful interpretation.
The same consideration applies to the reported 57–65 nm systematic offsets. A constant wavelength offset does not correspond to a constant error in excitation energy across the spectrum.
One of the key claims is that quantum-chemical calculations become substantially more accurate than the GP after each method's “removable offset” is discounted.
The procedure used to estimate and remove this offset is therefore crucial.
If the offset is estimated from the same test dataset on which residual error is subsequently reported, the corrected error is optimistically biased. This effectively gives the method access to information from the evaluation set.
A fair calibration procedure would estimate the correction on an independent calibration or training dataset and then apply that fixed correction to an untouched external test set.
The manuscript should explicitly state how offsets were estimated and whether test labels influenced the correction.
Both uncorrected and prospectively calibrated results should be reported. Otherwise, comparing a calibrated quantum-chemical method against an uncalibrated GP may not constitute a balanced comparison.
The observation that an approximately 57–65 nm underestimation in the in-corpus data reappears as approximately −61.9 nm in an external dataset is interesting.
However, the conclusion that this supports a “definitional rather than functional origin” seems too strong.
A reproducible systematic bias could also arise from features shared across quantum-chemical calculations, including the use of vertical excitations, similar density functionals, incomplete treatment of solvent effects, neglect of vibronic structure, conformational differences, or geometry approximations.
The recurrence of the offset establishes reproducibility of the bias, but not necessarily its mechanistic origin.
To distinguish label-definition effects from computational bias more convincingly, the manuscript would need controlled comparisons involving multiple electronic-structure methods and experimentally assigned transitions on the same molecules.
The external photoswitch dataset is an important test because it uses a specifically defined E-isomer π→π* transition.
However, photoswitches are also a specialised chemical domain.
Therefore, deterioration in GP performance may simultaneously reflect changes in molecular structure, photochemical behaviour, chromophore families, wavelength distribution, and target definition.
The statement that “chemical novelty alone does not break it, reassigning the label does” is potentially one of the strongest contributions of the manuscript, but it requires particularly rigorous evidence.
Chemical novelty should be quantified rather than inferred. I suggest reporting nearest-neighbour Tanimoto similarities to the training corpus, scaffold overlap, descriptor-space distances, and performance as a function of molecular similarity.
Most importantly, label effects should ideally be demonstrated using alternative labels on the same molecular population.
The manuscript describes the in-corpus holdouts as leakage-free and additionally uses Bemis–Murcko scaffold splitting. These are appropriate safeguards, but literature-mined chemical datasets can contain less obvious forms of dependence.
For example, the same molecule may appear under different solvents or experimental conditions; closely related analogue series may originate from the same article; stereoisomers, tautomers, salts, or protonation states may be represented separately; and extraction procedures may generate correlated labels from a single source publication.
A scaffold split does not necessarily eliminate these dependencies.
A particularly informative analysis would perform a publication-level or source-level split, ensuring that all compounds extracted from a particular paper are confined to either training or test data. This would provide a much stronger test of whether the GP is learning transferable chemical relationships rather than corpus-specific regularities.
On the 5,811-compound external dataset described as label-compatible, the reported R² is only +0.209.
This does indicate some transfer, but it represents limited predictive performance.
Describing the model as continuing to “work” may therefore be somewhat generous unless performance is contextualised against simple baselines.
The manuscript should compare the GP against at least a training-set mean predictor, nearest-neighbour regression, and potentially simple descriptor-based models.
MAE, RMSE, correlation, bias, and calibration plots should accompany R². A positive R² alone does not demonstrate practically useful external prediction.
The observation that GP predictive uncertainty is poorly calibrated under external distribution shift is useful.
However, this result concerns uncertainty generated by a particular GP kernel operating on a particular fingerprint representation. It should not be interpreted as a general result about machine-learning uncertainty.
The analysis would benefit from standard quantitative calibration measures, including predictive interval coverage, negative log-likelihood, calibration curves, and error-versus-uncertainty ranking metrics.
It would also be valuable to compare GP uncertainty against simple structural-distance measures. If uncertainty primarily reflects distance in fingerprint space, this comparison could clarify what information the GP uncertainty actually contains.
Post-hoc or conformal calibration using a validation set would also help distinguish a fundamental failure of uncertainty estimation from correctable miscalibration.
Many of the manuscript's conclusions are based on differences in RMSE, correlation coefficients, or residual scatter.
Confidence intervals are therefore important.
This is particularly necessary when comparing datasets with sample sizes ranging from only 36 molecules to more than 5,000.
Bootstrap confidence intervals for RMSE, MAE, correlation coefficients, systematic offsets, and differences between methods would make it possible to distinguish robust effects from potentially unstable numerical differences.
The manuscript correctly highlights that an algorithm can perform well against a benchmark without necessarily describing the same physical quantity as a quantum-chemical calculation.
However, this does not necessarily make the ML result misleading.
If the intended task is to predict the wavelength reported in experimental literature, learning the empirical conventions of those measurements may itself be useful.
Conversely, if the scientific objective is to predict a defined electronic transition, a physics-based method may be the more appropriate comparator.
I therefore suggest framing the result less as determining which method is genuinely “more capable” and more as showing that method rankings are conditional on the definition of the prediction task.
This interpretation is both more precise and more broadly useful.
The manuscript would benefit from a dataset summary table reporting the number of molecules, λmax range, mean and variance of λmax, chemical domain, experimental label definition, computational method where applicable, degree of training-set overlap, and structural-similarity statistics for every benchmark.
Performance distributions rather than only aggregate metrics would also be informative. Scatter plots and residual plots stratified by wavelength region could reveal whether apparent method differences are driven primarily by particular spectral regimes.
Because the central argument concerns transition identity, representative examples in which the literature-defined λmax and calculated transition clearly refer to different electronic states would make the conceptual problem much more tangible.
The author should also clarify how ambiguous or multi-peak spectra were handled during curation and whether solvent, pH, protonation state, aggregation, and isomeric composition were consistently available. These factors can alter experimental λmax independently of either model capability or label convention.
Finally, full release of molecular identifiers, dataset splits, source-publication mappings, model hyperparameters, prediction files, and analysis scripts would be particularly valuable for a paper whose central argument concerns benchmarking practice.
This manuscript raises an important and credible concern about how machine-learning models and quantum-chemical calculations are compared for UV–Vis prediction. The external-validation strategy and explicit attention to target definition are significant strengths, and the work could make a useful contribution to benchmark methodology in computational chemistry.
However, the current central claim appears more definitive than the evidence warrants. The study demonstrates that benchmark choice and label convention strongly influence the apparent relative performance of the evaluated GP and quantum-chemical predictions. It does not yet convincingly establish that label definition, rather than model capability, is the dominant or exclusive explanation for the reported ML advantage.
In particular, the conclusions would be substantially strengthened by matched same-molecule comparisons under alternative label definitions, evaluation in energy units as well as wavelength, prospectively calibrated rather than test-set-derived offset correction, stronger controls for literature-source leakage, statistical uncertainty estimates, and comparison with a broader set of machine-learning representations.
With these revisions, the manuscript could provide a strong and useful caution against interpreting benchmark leaderboards as direct measurements of intrinsic methodological superiority.
The authors declare that they have no competing interests.
The authors declare that they did not use generative AI to come up with new ideas for their review.
No se han publicado comentarios aún.