Efficient biomolecule modeling and drug discovery with large language models
ID:69 View Protection:ATTENDEE Updated Time:2025-03-25 14:42:39 Hits:468 Oral Presentation

Start Time:2025-03-30 10:20(Asia/Shanghai)

Duration:20min

Session:S7 前沿论坛 (基因组大数据与AI) » s7前沿论坛(基因组大数据与AI)

No files

Abstract
Large language models, which can integrate and process large amounts of data in biomedicine, have great potential in modeling complex diseases and discovering functional biomolecules for potential therapeutics. In this talk, we will first introduce the models based on protein language models to efficiently discover remote homologs and functional biomolecules from nature, such as signal peptides. With the model, we can identify remote homologs 22 times faster than PSI-BLAST and discover diverse functional peptides with sequence similarity lower than 20% against the known ones. Then, we developed an RNA language model to model the RNA sequence and structure relation, which enables us to perform RNA structure prediction and reverse design effectively. Within two months, we designed and experimentally validated 19 RNA aptamers that are structurally similar, yet sequence dissimilar, to known light-up aptamers. More importantly, 10 designed aptamers show higher fluorescence than the native Mango-I. The above projects demonstrate the great potential of large language models in promoting fundamental computational biological research and transformational development.
Keywords
Speaker
李煜
助理教授 香港中文大学

Submit Comment
Verify Code Change Another
All Comments
Important Date
  • Conference Date

    Mar 28

    2025

    to

    Mar 30

    2025

  • Apr 15 2025

    Registration deadline

Sponsored By
中国生物信息学学会基因组信息学专业委员会
Organized By
中国农业科学院农业基因组研究所
大鹏湾实验室
Contact Information
Previous Conferences