|
所在平台: Coursera |
课程主页: https://www.coursera.org/learn/four-rare-machine-learning-skills-all-data-scientists-need
课程评论:没有评论
课程名称:数据科学家必备的四种稀有机器学习技能 课程概述:本课程覆盖了机器学习中被忽视但至关重要的四种技能,这些技能在大多数课程和书籍中几乎没有提及。 1) 引导建模(UPLIFT MODELING,又称说服建模):在建模时,您是否真的在预测正确的事物? 2) 准确性谬论(ACCURACY FALLACY):在评估模型效果时,您是否在报告正确的指标? 3) P值操控(P-HACKING):您从数据中得出的简单发现是否真实有效? 4) 集成模型的悖论(PARADOX OF ENSEMBLE MODELS):尽管它们似乎违背奥卡姆剃刀原则,您是否理解它们的工作原理? 为什么这些高级方法至关重要:每一个技能都解决了机器学习中的基本问题。对于许多项目而言,成功依赖于这些特定技能的运用。 课程特色:此课程不包括实践操作,专为技术学习者设计。它没有代码和机器学习软件的使用,而是为实践打下了概念基础。在深入实践前,理解这些先进技术及常见陷阱的概念知识是非常重要的。 课程内容中立:尽管课程包含了使用SAS产品的机器学习动态演示,但课程内容中立,面向所有适用的机器学习软件工具。学习目标和内容适用于您选择的任何机器学习工具。 课程大纲:为期一周的课程,只有一个模块,涵盖四个稀有但重要的主题: 1) 引导建模:如何在无法普遍建立因果关系的情况下优化营销,以预测对结果的影响。 2) 准确性谬论:高准确性是达不到的,且准确性本身并不是正确的评估指标。 3) P值操控:大数据可能带来的更大危害,如何确保科学发现的可靠性。 4) 集成模型的悖论:在不涉及神经网络复杂性的情况下,如何提升模型的能力和性能。
Part: 1
Title:Four Rare Machine Learning Skills All Data Scientists Need
Description:This one-week course has only one module, which covers the course's four rare yet vital topics: (1) UPLIFT MODELING: How do you optimize marketing – which is meant to persuade – if we cannot generally establish causal relationships? Put another way, how do you model and predict influence when you cannot measure influence? The special, advanced method uplift modeling (aka persuasion modeling) goes beyond predicting an outcome to actually predicting the influence that a treatment decision would have on that outcome. We'll explore the marketing applications of uplift modeling and see success stories from the likes of US Bank and President Obama's 2012 reelection campaign. (2) THE ACCURACY FALLACY: For many machine learning projects, high accuracy is unattainable – and, besides, accuracy isn't the right metric in the first place. But many projects are falsely advertised as "highly accurate." Learn to identify occurrences of the accuracy fallacy, a common misstep by which researches spread misinformation about predictive model performance. (3) P-HACKING: In what way is bigger data more dangerous? How do we avoid being fooled by random noise and ensure scientific discoveries are trustworthy? This prevalent pitfall is a huge gotcha! (4) THE PARADOX OF ENSEMBLE MODELS: Is there a way to advance model capability and performance that's elegant and simple, without involving the complexity of neural networks? Why yes there is.
This course covers the most neglected yet critical skills in machine learning, four vital techniques that are very rarely covered – most courses and books omit them entirely. 1) UPLIFT MODELING (AKA PERSUASION MODELING): When you're modeling, are you even predicting the right thing? 2) THE ACCURACY FALLACY: When evaluating how well a model works, are you even reporting on the right thing? 3) P-HACKING: Are your simplest discoveries from data even real? 4) THE PARADOX OF ENSEMBLE MODELS: Do you understand how they work, even though they seem to defy Occam's Razor? >> WHY THESE ADVANCED METHODS ARE ESSENTIAL: Each one addresses a question that is fundamental to machine learning (above). For many projects, success hinges on these particular skills. >> NO HANDS-ON – BUT FOR TECHNICAL LEARNERS: This course has no coding and no use of machine learning software. Instead, it lays the conceptual groundwork before you take on the hands-on practice. When it comes to these state-of-the-art techniques and prevalent pitfalls, there's a foundation of conceptual knowledge to build before going hands-on – and you'll be glad you did. >> VENDOR-NEUTRAL: This course includes illuminating software demos of machine learning in action using SAS products. However, the curriculum is vendor-neutral and universally-applicable. The contents and learning objectives apply, regardless of which machine learning software tools you end up choosing to work with.