Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Preprint . 2026
License: CC BY
Data sources: ZENODO
ZENODO
Preprint . 2026
License: CC BY
Data sources: Datacite
ZENODO
Preprint . 2026
License: CC BY
Data sources: Datacite
ZENODO
Preprint . 2026
License: CC BY
Data sources: Datacite
versions View all 2 versions
addClaim

The OMol25 Paradox: Training-Data Domain Specificity Outweighs Architecture in Neural Network Potential Performance on Zn²⁺ Metalloenzyme Conformer Energies

Authors: Han, Cheongwoo;

The OMol25 Paradox: Training-Data Domain Specificity Outweighs Architecture in Neural Network Potential Performance on Zn²⁺ Metalloenzyme Conformer Energies

Abstract

Correction (new version, 2026-07-18) — integrity correction. Corrected version (2026-07-18), fabricated ligand-panel annotations removed. A primary-source audit (2026-07-16) established that the 15-ligand calibration panel used here carried fabricated compound names, potencies and literature attributions (all seven named entries are a different molecule than named; 14 of 15 structures are unknown to PubChem). This version replaces the Section 2.1 ligand table with structure-only descriptors, deletes the drug names and the reported-IC50 column, retains ChEMBL accessions only as unverified structure keys, and withdraws the pIC50 biology-validation (Sections 3.5/4.5) in full. The headline OMol25-paradox result and all cross-NNP agreement statistics are unaffected — they are properties of computations on fixed structures, independent of compound identity or potency. Versioned correction, not a retraction. In silico only. Machine-learning interatomic potentials (MLIPs) trained at high-level DFT are widely assumed to dominate semi-empirical and materials-trained alternatives on molecular tasks. We test this assumption on the conformational energy landscape of fifteen Zn²⁺-binding matrix metalloproteinase-1 (MMP-1) inhibitors drawn from ChEMBL, generated by Boltz-2x cofold sampling with physicality-steering (--use_potentials). Across thirteen independent Boltz-2x diffusion replicates (Boltz seeds 42–54, conditions v11–v23), we evaluate ten energy functions on 100 conformers per ligand per replicate (19,500 conformers per replicate; 253,500 conformer single points total): three semi-empirical (GFN1-xTB, GFN2-xTB, GFN-FF), one Acellera-trained NN force field (AceFF-2), two molecular NN potentials (AIMNet2-NSE, ANI-2x), three universal MLIPs trained on materials physics (Orb-v3 OMat, Orb-v3 OMol25, MatterSim-5M), and two classical force fields (MMFF94, UFF). Per-ligand intra-conformer Pearson correlations are computed within each ligand (size-invariant by construction), then averaged across the 15 ligands and 13 replicates. Three clusters emerge with high cross-replicate stability (SD ≤ 0.038). Cluster I (QM-like) tightly groups GFN1, GFN2, AIMNet2, AceFF-2, Orb-v3 OMat, and MatterSim-5M (intra-cluster r = 0.66–0.99). Cluster II (force-field) groups MMFF94 and UFF (r = 0.66), weakly anti-correlated with Cluster I (r = −0.13 to −0.25). The headline finding is the OMol25 paradox: the same Orb-v3 architecture trained on Meta FAIR's OMol25 molecular dataset (ωB97M-V/def2-TZVPD) drops to r = 0.374 ± 0.025 against GFN1-xTB, versus r = 0.886 ± 0.009 for Orb-v3 OMat — a 0.512 Pearson gap with ≈ 68σ statistical significance across 13 replicates. The paradox is corroborated by a second materials-trained MLIP, MatterSim-5M, which agrees with Orb-v3 OMat at r = 0.936 ± 0.006 (the tightest NN–NN pair in our matrix) but drops to r = 0.410 ± 0.026 against Orb-v3 OMol25. Cluster placement is set by training-data domain, not by architecture. We discuss the mechanistic explanation — OMol25's self-reported weakness on "long-range interactions" (Levine et al., arXiv:2505.08762) coincides with the dominant physics of MMP-1 hydroxamate–Zn²⁺ chemistry — and we explicitly retract a preliminary biology-validation claim that was confounded by ligand size. We recommend that practitioners choose materials-trained or broadly-trained MLIPs (Orb-v3 OMat, MatterSim, AIMNet2) over molecular-only-trained MLIPs (Orb-v3 OMol25, SevenNet OMol25) for Zn²⁺ metalloenzyme conformer-ensemble work, pending higher-level DFT validation. Word count (abstract): 388 words. v5h revision (2026-05-15): Section 2.6 added: extended-cascade pipeline timing reproducibility (chains v_v77–v_v93, 17 independent ensembles) 15-cycle SUSTAINED baseline reported for the four-stage chain backbone (SP CV 1.9%, OPT CV 0.4%, HESS CV 0.94%, GFN-FF CV 1.0%) GFN-FF cache 5-cycle plateau finding (C88–C92 = 10.7 min, C93 break to 11.0 min) Chain × concurrent-ADMET fair-share linear regression model derived from C80 (8 ADMET chains, +52% saturation cap) and C93/C94 (1 chain, +10–13% per stage): regression% ≈ 0.34 × overlap_fraction × N_ADMET_chains Section 4.6 Limitation #6 added (timing-baseline ADMET caveat) Original cross-NNP cluster results (Sections 3.1–3.9) and OMol25 Paradox conclusion UNCHANGED PDF rebuilt with typst engine to fix WeasyPrint table-row drop bug (v0.1 PDF had 14/16 ChEMBL panel rows + 3/4 timing rows missing)

CORRECTION (new version, 2026-07-18): Corrected version (2026-07-18), fabricated ligand-panel annotations removed. A primary-source audit (2026-07-16) established that the 15-ligand calibration panel used here carried fabricated compound names, potencies and literature attributions (all seven named entries are a different molecule than named; 14 of 15 structures are unknown to PubChem). This version replaces the Section 2.1 ligand table with structure-only descriptors, deletes the drug names and the reported-IC50 column, retains ChEMBL accessions only as unverified structure keys, and withdraws the pIC50 biology-validation (Sections 3.5/4.5) in full. The headline OMol25-paradox result and all cross-NNP agreement statistics are unaffected — they are properties of computations on fixed structures, independent of compound identity or potency. Versioned correction, not a retraction. In silico only. HCW is founder of HAN PREDICT, Inc. and consults for Recover Korean Medicine Clinic. No external funding for this work. In silico only. No wet-lab data and no patient data are reported. IRB approval was filed 2026-04-27 and is pending. Recover Korean Medicine Clinic opens 2026-08-15. Direct Zenodo deposit (no prior journal submission). bioRxiv/ChemRxiv door confirmed closed for in-silico-only manuscripts per 2026-04~05 rejection cohort.

Keywords

MMP-1, metalloenzyme, cofold, neural network potentials, MatterSim, Zn2+, hydroxamate, Orb-v3, Boltz-2x, drug discovery, OMol25, benchmark, in silico, machine-learning interatomic potentials

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average
Green