|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/gene-prediction-a-step-by-step-guide-to-finding-genes/
课程评论:没有评论
Coursera 基因预测:寻找基因的生物信息学协议 课程总结 本课程聚焦于基因预测,即识别基因组中编码蛋白质或非编码基因区域的关键生物信息学过程。在全球各类基因组数据库中,许多已组装但未标注(或标注过时)的基因组占据了显著比例,因此,对基因组进行重新注释至关重要。 课程提供一套实用的、分步的、带有屏幕录制的教程,演示如何使用多种基因预测软件。所有协议均经过修改和测试,确保了准确性和无误性,适用于从原核生物到真核生物的所有类型基因组,涵盖了蛋白质编码基因和非编码基因。此外,课程还教授如何评估基因预测软件的输出结果。 **核心软件与应用:** * **通用软件 (原核生物 & 真核生物):** * **tRNAscan-SE:** 用于预测 tRNA 基因,对计算资源要求不高。 * **Infernal:** 结合 Rfam 数据库,用于预测所有非编码 RNA 基因,尤其擅长 rRNA 和其他非编码 RNA。对计算资源要求不高,但运行时间较长。 * **真核生物专属软件:** * **BRAKER:** 利用从头(ab initio)方法预测蛋白质编码基因,并整合 RNA-seq 和蛋白质比对证据以提高准确性。需要高性能计算资源。 * **GeMoMa:** 通过同源性预测蛋白质编码基因,并结合 RNA-seq 比对证据。需要高性能计算资源。 * **原核生物专属软件:** * **Prokka:** 采用从头(ab initio)方法预测蛋白质编码基因和非编码基因,并利用蛋白质证据。对计算资源要求不高。 * **评估软件:** * **BUSCO:** 用于评估原核生物和真核生物中蛋白质编码基因的预测结果。对计算资源要求不高。 **技术支持与实践:** 所有软件操作教程均在 Ubuntu Linux 发行版上录制。强烈建议在您的设备上安装 Ubuntu 系统,或使用 VirtualBox 创建 Ubuntu 虚拟机,以便进行实际操作。 **学习成果:** 完成本课程后,您将掌握一套可用于数据库基因组分析的文字版基因预测协议,可直接应用于毕业设计、研究论文或海报展示中。
Gene prediction or gene finding is an essential step aimed at identifying gene regions in genomes, whether these genes are non-coding or protein-coding.There are many assembled genomes in databases used in various genetic studies, but few of these genomes have annotations.In addition, only a very small set of genomes are updated annotations.Therefore, in research dealing with genomes from databases, the step of re-annotation is essential.Therefore, we present this course with practical, step-by-step, screen-recorded content for the use of gene prediction software with more than one example.This course provides a set of protocols that have been modified and implemented without any errors.The course includes all organisms from prokaryotes to eukaryotes.The course Includes protein-coding genes and non-coding genes.In addition to evaluating the resulting prediction from gene prediction software.There are generic programs that are used with prokaryotes and eukaryotes:tRNAscan-SE is used to predict tRNA genes. It is characterized by the fact that it does not require a computer with high capabilities.Infernal is used with rfam database to predict all non-coding genes but is more accurate in predicting rRNA and other non-coding RNA genes. It is characterized by the fact that it does not require a computer with high capabilities, but it takes a long time.There are programs for eukaryotic organisms:BRAKER, which is used to predict protein-coding genes using the ab initio method, also uses extrinsic evidence from mapped RNA-seq and protein to support and increase accuracy. Requires a computer with high capabilities.GeMoMa is used to predict protein-coding genes by homology and also uses extrinsic evidence from mapped RNA-seq. Requires a computer with high capabilities.There is a program for prokaryotic organisms:Prokka is used to predict protein-coding and non-coding genes using the ab initio method and also uses extrinsic evidence from the protein. It does not require a high-powered computer.Finally, the BUSCO program is used to evaluate the prediction of protein-coding genes in prokaryotic and eukaryotic organisms. It does not require a high-powered computer.All programs have been screen-recorded on the ubuntu distribution since it is the most famous Linux distribution.It is best to install the Ubuntu distribution on your device or create virtual ubuntu using a virtual box in order to implement the course practically on your device.You will eventually get the protocols in text format that you can apply with genomes from the database in a graduation project, in a research paper, or in a poster.