Researchers have developed a novel semi-supervised learning framework to predict nuclear magnetic resonance (NMR) chemical shifts, a crucial step in spectral analysis and molecular structure elucidation. This new method leverages millions of unassigned spectra extracted from scientific literature, integrating them with a smaller set of explicitly labeled data. The approach treats chemical shift prediction from literature spectra as a permutation-invariant set supervision problem, utilizing an optimal bipartite matching technique that simplifies to a sorting-based loss for stable, large-scale training. The resulting models demonstrate superior accuracy and robustness compared to existing state-of-the-art methods, with improved generalization on larger and more diverse molecular datasets. Notably, the framework is the first to incorporate solvent information at scale, capturing systematic solvent effects across common NMR solvents and highlighting the potential of literature-derived, weakly structured data for AI in scientific research. AI
IMPACT This research demonstrates a novel method for leveraging large-scale unlabeled scientific literature to train AI models, potentially accelerating discovery in chemistry and other scientific fields.
RANK_REASON Academic paper detailing a new machine learning methodology for scientific research. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX Code Finder for Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- IArxiv Recommender
- ScienceCast
- Yongqi Jin
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →