Big Data Science with the BD2K-LINCS Data Coordination and Integration Center

所在平台: Coursera

课程主页: https://www.coursera.org/learn/bd2k-lincs

课程评论:没有评论

第一个写评论        关注课程

课程简介

课程名称:与BD2K-LINCS数据协调与整合中心的“大数据科学” 课程概述: 本课程由美国国立卫生研究院(NIH)共同基金项目“网络基础细胞签名库(LINCS)”提供,旨在探讨不同类型的人类细胞在多种扰动(如药物、遗传操控和微环境条件变化等)下的反应。学生将学习如何通过量测各种变量的变化(如mRNA、蛋白质和代谢物的水平,以及细胞形态等表型变化),以便深入理解这些扰动影响的分子网络。课程特别介绍了BD2K-LINCS数据协调与整合中心(DCIC)如何组织、分析、可视化和整合这些数据,并讲解元数据与本体的关联。 课程大纲: 1. LINCS程序概述:介绍LINCS项目的概念及使用LINCS L1000数据集的入门教程。 2. 元数据与本体:描述元数据和本体的基本概念及其在LINCS数据集中的应用。 3. 通过API提供数据:讲解如何通过应用程序接口(API)访问数据。 4. 生物信息学管道:描述生物信息学管道的重要概念。 5. Harmonizome项目:介绍整合基因和蛋白质知识的项目Harmonizome,并提供Web服务应用的链接。 6. 数据标准化:讲解数据标准化的数学概念。 7. 数据聚类:讨论无监督学习中的数据聚类概念及如何识别数据中的模式。 8. 期中考试:包含45道选择题,涵盖前7个模块。 9. 富集分析:介绍基因集富集分析的概念,探讨如何从基因组和蛋白质组研究中查询基因集。 10. 机器学习:讲解监督机器学习的数学概念及其预测过程。 11. 基准测试:讨论如何比较和评估生物信息学管道。 12. 互动数据可视化:提供编程示例,指导创建互动型数据可视化元素。 13. 众包项目:介绍超越课程的LINCS相关项目合作的机会。 14. 期末考试:涵盖所有模块的60道选择题,部分题目可能需要使用所学分析方法处理新数据集。 该课程旨在为学生提供大数据科学的理论基础和实用技能,鼓励随机应变以及团队合作,以推动生物医学研究的进展。

课程大纲

Name:The Library of Integrated Network-based Cellular Signatures (LINCS) Program Overview

Description:This module provides an overview of the concept behind the LINCS program; and tutorials on how to get started with using the LINCS L1000 dataset.

Name:Metadata and Ontologies

Description:This module includes a broad high level description of the concepts behind metadata and ontologies and how these are applied to LINCS datasets.

Name:Serving Data with APIs

Description:In this module we explain the concept of accessing data through an application programming interface (API).

Name:Bioinformatics Pipelines

Description:This module describes the important concept of a Bioinformatics pipeline.

Name:The Harmonizome

Description:This module describes a project that integrates many resources that contain knowledge about genes and proteins. The project is called the Harmonizome, and it is implemented as a web-server application available at: http://amp.pharm.mssm.edu/Harmonizome/

Name:Data Normalization

Description:This module describes the mathematical concepts behind data normalization.

Name:Data Clustering

Description:This module describes the mathematical concepts behind data clustering, or in other words unsupervised learning - the identification of patterns within data without considering the labels associated with the data.

Name:Midterm Exam

Description:The Midterm Exam consists of 45 multiple choice questions which covers modules 1-7. Some of the questions may require you to perform some analysis with the methods you learned throughout the course on new datasets.

Name:Enrichment Analysis

Description:This module introduces the important concept of performing gene set enrichment analyses. Enrichment analysis is the process of querying gene sets from genomics and proteomics studies against annotated gene sets collected from prior biological knowledge.

Name:Machine Learning

Description:This module describes the mathematical concepts of supervised machine learning, the process of making predictions from examples that associate observations/features/attribute with one or more properties that we wish to learn/predict.

Name:Benchmarking

Description:This module discusses how Bioinformatics pipelines can be compared and evaluated.

Name:Interactive Data Visualization

Description:This module provides programming examples on how to get started with creating interactive web-based data visualization elements/figures.

Name:Crowdsourcing Projects

Description:This final module describes opportunities to work on LINCS related projects that go beyond the course.

Name:Final Exam

Description:The Final Exam consists of 60 multiple choice questions which covers all of the modules of the course. Some of the questions may require you to perform some analysis with the methods you learned throughout the course on new datasets.

课程评论(0条)

课程详情

The Library of Integrative Network-based Cellular Signatures (LINCS) is an NIH Common Fund program. The idea is to perturb different types of human cells with many different types of perturbations such as: drugs and other small molecules; genetic manipulations such as knockdown or overexpression of single genes; manipulation of the extracellular microenvironment conditions, for example, growing cells on different surfaces, and more. These perturbations are applied to various types of human cells including induced pluripotent stem cells from patients, differentiated into various lineages such as neurons or cardiomyocytes. Then, to better understand the molecular networks that are affected by these perturbations, changes in level of many different variables are measured including: mRNAs, proteins, and metabolites, as well as cellular phenotypic changes such as changes in cell morphology. The BD2K-LINCS Data Coordination and Integration Center (DCIC) is commissioned to organize, analyze, visualize and integrate this data with other publicly available relevant resources. In this course we briefly introduce the DCIC and the various Centers that collect data for LINCS. We then cover metadata and how metadata is linked to ontologies. We then present data processing and normalization methods to clean and harmonize LINCS data. This follow discussions about how data is served as RESTful APIs. Most importantly, the course covers computational methods including: data clustering, gene-set enrichment analysis, interactive data visualization, and supervised learning. Finally, we introduce crowdsourcing/citizen-science projects where students can work together in teams to extract expression signatures from public databases and then query such collections of signatures against LINCS data for predicting small molecules as potential therapeutics.

课程标签

0人关注该课程

主题相关的课程