|
所在平台: Coursera |
课程主页: https://www.coursera.org/learn/data-results
课程评论:没有评论
课程名称:传达数据科学结果 课程概述:本课程的第二个作业涉及云中的图分析,您将使用 Elastic MapReduce 和 Pig 语言对大约600GB的中等规模数据集进行图分析。为了完成此作业,您需要使用亚马逊网络服务(AWS)。亚马逊已慷慨提供每位学习者最多50美元的免费AWS信用额度,以便您完成作业。有关如何获得这些信用额度的更多详细信息,请查阅课程欢迎信息和作业本身。请注意,亚马逊、华盛顿大学和Coursera不能对超出信用额度的费用进行报销。如果你无法或不愿意使用AWS,我们无法向未完成作业的学习者颁发课程证书。因此,如果您无法或不愿意使用AWS,请不要支付《传达数据结果》课程的证书费用,因为您将无法顺利完成课程。 有效的数据科学家不仅会进行预测,还会解释和解读结果,并准确地向利益相关者传达发现,以便为商业决策提供信息。可视化是计算机科学中研究定量结果有效沟通的领域,将感知、认知和算法结合起来,充分利用人类视觉皮层的巨大的带宽。在本课程中,您将学习如何识别、设计和使用有效的可视化。 尽管您可以做出预测并说服他人采取行动,但并不意味着您应该这样做。在本课程中,您将探讨大数据的伦理考虑,以及这些考虑如何开始影响政策和实践。您将学习利用技术保护隐私的基础局限性,以及正在出现的指导数据科学家行为的行为规范。您还将了解可重复性在数据科学中的重要性,以及商业云如何支持即使涉及大规模数据集、复杂计算基础设施或两者结合的可重复研究。 学习目标:完成本课程后,您将能够: 1. 设计和评估可视化。 2. 解释大数据和数据科学周围隐私、伦理和治理的最新发展。 3. 使用云计算以可重复的方式分析大数据集。 课程大纲: 第一部分:可视化 描述:来自大型、异构和嘈杂数据集的统计推断是毫无意义的,如果您无法与同事、客户、管理层和其他利益相关者进行沟通。学习信息可视化的基本概念,这是一项对数据科学家越来越重要的研究领域和技能。此模块由人本设计与工程系的 Cecilia Aragon 教授授课。 第二部分:隐私与伦理 描述:大数据与隐私和伦理问题密切相关:随着我们在数据上可以做的事情的限制逐渐消失,关于我们应该如何处理数据的问题变得至关重要。在案例研究的背景下,您将学习数据科学和统计分析的行为规范核心原则。您将了解在允许有用的统计分析的同时保护隐私的当前理论局限性。 第三部分:可重复性与云计算 描述:科学正面临由于不可靠的可重复性而导致的信誉危机,随着研究变得愈发计算化,这一问题似乎变得更加严重。但可重复性不仅仅是学术界的事情:无法分享、解释和捍卫其方法的数据科学家是危险的。在此模块中,您将探索可重复研究的重要性,以及云计算如何提供新的机制来共享代码、数据、环境,甚至是对实用可重复性至关重要的费用。
Part: 1
Title:Visualization
Description:Statistical inferences from large, heterogeneous, and noisy datasets are useless if you can't communicate them to your colleagues, your customers, your management and other stakeholders. Learn the fundamental concepts behind information visualization, an increasingly critical field of research and increasingly important skillset for data scientists. This module is taught by Cecilia Aragon, faculty in the Human Centered Design and Engineering Department.
Part: 2
Title:Privacy and Ethics
Description:Big Data has become closely linked to issues of privacy and ethics: As the limits on what we *can* do with data continue to evaporate, the question of what we *should* do with data becomes paramount. Motivated in the context of case studies, you will learn the core principles of codes of conduct for data science and statistical analysis. You will learn the limits of current theory on protecting privacy while still permitting useful statistical analysis.
Part: 3
Title:Reproducibility and Cloud Computing
Description:Science is facing a credibility crisis due to unreliable reproducibility, and as research becomes increasingly computational, the problem seems to be paradoxically getting worse. But reproducibility is not just for academics: Data scientists who cannot share, explain, and defend their methods for others to build on are dangerous. In this module, you will explore the importance of reproducible research and how cloud computing is offering new mechanisms for sharing code, data, environments, and even costs that are critical for practical reproducibility.
Important note: The second assignment in this course covers the topic of Graph Analysis in the Cloud, in which you will use Elastic MapReduce and the Pig language to perform graph analysis over a moderately large dataset, about 600GB. In order to complete this assignment, you will need to make use of Amazon Web Services (AWS). Amazon has generously offered to provide up to $50 in free AWS credit to each learner in this course to allow you to complete the assignment. Further details regarding the process of receiving this credit are available in the welcome message for the course, as well as in the assignment itself. Please note that Amazon, University of Washington, and Coursera cannot reimburse you for any charges if you exhaust your credit. While we believe that this assignment contributes an excellent learning experience in this course, we understand that some learners may be unable or unwilling to use AWS. We are unable to issue Course Certificates for learners who do not complete the assignment that requires use of AWS. As such, you should not pay for a Course Certificate in Communicating Data Results if you are unable or unwilling to use AWS, as you will not be able to successfully complete the course without doing so. Making predictions is not enough! Effective data scientists know how to explain and interpret their results, and communicate findings accurately to stakeholders to inform business decisions. Visualization is the field of research in computer science that studies effective communication of quantitative results by linking perception, cognition, and algorithms to exploit the enormous bandwidth of the human visual cortex. In this course you will learn to recognize, design, and use effective visualizations. Just because you can make a prediction and convince others to act on it doesn’t mean you should. In this course you will explore the ethical considerations around big data and how these considerations are beginning to influence policy and practice. You will learn the foundational limitations of using technology to protect privacy and the codes of conduct emerging to guide the behavior of data scientists. You will also learn the importance of reproducibility in data science and how the commercial cloud can help support reproducible research even for experiments involving massive datasets, complex computational infrastructures, or both. Learning Goals: After completing this course, you will be able to: 1. Design and critique visualizations 2. Explain the state-of-the-art in privacy, ethics, governance around big data and data science 3. Use cloud computing to analyze large datasets in a reproducible way.