解锁酶发现:高分辨率系统发育基因簇数据库的构建
CSTR:
作者:
作者单位:

(江南大学生物工程学院 江苏无锡 214122)

作者简介:

通讯作者:

中图分类号:

基金项目:

国家重点研发计划项目(2025YFA0923700);湖南省重点研发计划项目(2024JK2150);国家自然科学基金优秀青年科学基金项目(32322069)


Unlocking Enzyme Discovery: Construction of a High-resolution Phylogenetic Gene Cluster Database
Author:
Affiliation:

(School of Bioengineering, Jiangnan University, Wuxi 214122, Jiangsu)

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    高通量测序技术的应用使全球基因数据以指数级增长(每7个月翻番),为高效酶资源的发现提供了新机遇。然而,基因组结构的高度复杂性增加了功能注释的难度,使得海量数据中高质量、精确的注释信息仍然不足。基于序列比对和同源推断的传统注释方法,仅能对高相似性序列实现有限标注,在面对多功能活性酶(如具有催化混杂性的酶)及真核生物基因组中嵌套基因、可变剪接等复杂结构时存在瓶颈。本研究创新性构建跨物种自动化基因注释管道(涵盖基因结构预测、功能域识别和代谢通路重构)和基于多位点系统发育分析的功能序列挖掘体系,成功建立了覆盖43万细菌和2万真菌基因组的高分辨率系统发育基因簇数据库。通过多基因座进化树拓扑结构解析,精准鉴定到中链脂肪酸合成r-BOX循环中的关键酶组分,包括17 135个脂肪酸β-氧化多功能酶亚基(FadB)、17 115个β-酮硫解酶(BktB)、18 985个烯酰-CoA还原酶(Ter)及18 575个3-氧代酸CoA转移酶(YdiI),其中100%的序列在NCBI数据库中未见收录。为验证数据库效能,以FadB为靶标筛选出10种未表征的同源酶,其中3条展现出活性,通过重组表达使摇瓶培养产量从0.65 g/L提高至1.1 g/L。通过高分辨率数据库不仅构建了面向稀缺高价值酶挖掘的系统化方法框架,还配套开发了开放访问的交互式平台(https://thezsd.github.io/),支持用户自主开展酶资源挖掘及基因数据上传,实现精准注释,为合成生物学和代谢工程领域提供了数据驱动的解决方案。

    Abstract:

    High-throughput sequencing has driven an exponential expansion of global gene data (doubling approximately every seven months), creating new opportunities for discovering high-performance biocatalysts. However, the high structural complexity of genomes substantially increases the difficulty of functional annotation, leaving a large fraction of sequences in massive datasets with insufficient high-quality, precise annotations. Conventional annotation strategies based on sequence alignment and homology inference provide only limited labels for highly similar sequences and face major bottlenecks when dealing with multifunctional enzymes (e.g., catalytic promiscuity) and complex eukaryotic genome architectures such as nested genes and alternative splicing. This study developed an automated, cross-species genome annotation pipeline integrating gene structure prediction, domain identification, and metabolic pathway reconstruction, together with a function-oriented sequence mining framework based on multilocus phylogenetic analyses. Using these approaches, this study built a high-resolution phylogenetic gene-cluster database covering 430 000 bacterial and 20 000 fungal genomes. By analyzing multilocus tree topologies, this study precisely identified key enzymes in the r-BOX cycle for medium-chain fatty acid biosynthesis, including 17 135 fatty-acid β-oxidation multifunctional enzyme subunit(FadB), 17 115 β-ketothiolase (BktB), 18 985 enoyl-CoA reductase(Ter), and 18 575 3-oxoacid CoA-transferase(YdiI), none of which were found to be recorded in the NCBI database. To validate the utility of the database, this study targeted FadB and screened 10 uncharacterized homologs, of which three exhibited catalytic activity. Recombinant expression increased shake-flask production from 0.65 g/L to 1.1 g/L. In addition, this study deployed an openly accessible interactive platform (https://thezsd.github.io/) that enables user-driven enzyme mining and genome data submission to support accurate annotation. Collectively, this database-and-platform framework provided a data-driven solution for enzyme discovery and functional genomics in synthetic biology and metabolic engineering.

    参考文献
    相似文献
    引证文献
引用本文

张斯顿,刘文睿,贺俊龙,刘子民,吴俊俊.解锁酶发现:高分辨率系统发育基因簇数据库的构建[J].中国食品学报,2026,(2):40-47

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-03-19
  • 最后修改日期:
  • 录用日期:
  • 在线发布日期: 2026-04-03
  • 出版日期:
文章二维码
版权所有 :《中国食品学报》杂志社     京ICP备09084417号-4
地址 :北京市海淀区阜成路北三街8号9层      邮政编码 :100048
电话 :010-65223596 65265375      电子邮箱 :chinaspxb@vip.163.com
技术支持:北京勤云科技发展有限公司

漂浮通知


×
喜报 | 《中国食品学报》入选2025年度首都科技期刊卓越行动计划中英文单刊