5. AI for Science: machine learning for biophysics
GNN-based score-based MCMC and evaluation of foundation-model representations for biological sequences.
Background
Machine learning is changing how physical simulations are performed and how biological data are analysed. Two questions guide our work: can a learned model make sampling of molecular configurations more efficient, and when do the internal representations of large pre-trained models for biological sequences actually contain more information than simple, classical descriptors?
What we study
- Score-based MCMC with graph neural networks. A molecule can be represented as a graph of atoms and bonds. We train graph neural networks to learn the score function—the gradient of the log-probability of a configuration—and use it to guide the proposals of a Markov chain Monte Carlo sampler. This extends our GM-MCMC work with a learned component.
- Foundation-model representations for biological sequences. Large pre-trained models for proteins and DNA produce high-dimensional embeddings that are widely used in downstream tasks. We evaluate systematically, under identical protocols, when these embeddings genuinely outperform trivial baselines such as k-mer counts, and when they do not.
- Computing environment. A GPU workstation in the laboratory, together with allocations on the supercomputer Fugaku, supports this integration of simulation and machine learning.
Support
This research is supported by the MEXT programme “AI for Science” (SPReAD) and by JSPS KAKENHI.