Big Data Analysis: Hive, Spark SQL, DataFrames and GraphFrames

所在平台: CourseraArchive

课程类别: 其他类别

大学或机构: CourseraNew

课程主页: https://www.coursera.org/archive/big-data-analysis

课程评论:没有评论

第一个写评论        关注课程

课程简介

Yandex

课程评论(0条)

课程详情

No doubt working with huge data volumes is hard, but to move a mountain, you have to deal with a lot of small stones. But why strain yourself? Using Mapreduce and Spark you tackle the issue partially, thus leaving some space for high-level tools. Stop struggling to make your big data workflow productive and efficient, make use of the tools we are offering you. This course will teach you how to: - Warehouse your data efficiently using Hive, Spark SQL and Spark DataFframes. - Work with large graphs, such as social graphs or networks. - Optimize your Spark applications for maximum performance. Precisely, you will master your knowledge in: - Writing and executing Hive & Spark SQL queries; - Reasoning how the queries are translated into actual execution primitives (be it MapReduce jobs or Spark transformations); - Organizing your data in Hive to optimize disk space usage and execution times; - Constructing Spark DataFrames and using them to write ad-hoc analytical jobs easily; - Processing large graphs with Spark GraphFrames; - Debugging, profiling and optimizing Spark application performance. Still in doubt? Check this out. Become a data ninja by taking this course! Special thanks to: - Prof. Mikhail Roytberg, APT dept., MIPT, who was the initial reviewer of the project, the supervisor and mentor of half of the BigData team. He was the one, who helped to get this show on the road. - Oleg Sukhoroslov (PhD, Senior Researcher at IITP RAS), who has been teaching MapReduce, Hadoop and friends since 2008. Now he is leading the infrastructure team. - Oleg Ivchenko (PhD student APT dept., MIPT), Pavel Akhtyamov (MSc. student at APT dept., MIPT) and Vladimir Kuznetsov (Assistant at P.G. Demidov Yaroslavl State University), superbrains who have developed and now maintain the infrastructure used for practical assignments in this course. - Asya Roitberg, Eugene Baulin, Marina Sudarikova. These people never sleep to babysit this course day and night, to make your learning experience productive, smooth and exciting.

大数据分析:Hive,Spark SQL,DataFrames和GraphFrames:无疑,处理海量数据非常困难,但是要想翻山越岭,您必须处理很多小问题。但是为什么要紧张自己呢?使用Mapreduce和Spark可部分解决问题,从而为高级工具留出一些空间。停止为使大数据工作流程高效而高效而奋斗,请使用我们为您提供的工具。   本课程将教您如何: -使用Hive,Spark SQL和Spark DataFframe有效地存储数据。 -使用大型图,例如社交图或网络。 -优化Spark应用程序以获得最佳性能。 准确地说,您将掌握以下知识: -编写和执行Hive& Spark SQL查询; -推理如何将查询转换为实际的执行原语(无论是MapReduce作业还是Spark转换); -在Hive中组织数据以优化磁盘空间使用和执行时间; -构造Spark DataFrames并使用它们轻松编写临时分析作业; -使用Spark GraphFrames处理大型图形; -调试,分析和优化Spark应用程序性能。   还是有疑问吗?看一下这个。通过学习本课程成为数据忍者! 特别感谢: -MIPT APT部门的Mikhail Roytberg教授,他是该项目的最初审阅者,也是BigData团队一半的主管和导师。他是帮助推动这场演出的人。 -Oleg Sukhoroslov(博士,IITP RAS高级研究员),自2008年以来一直在教授MapReduce,Hadoop和朋友。现在,他领导基础架构团队。 -奥列格·伊夫琴科(MITP博士,APT系学生),帕维尔·阿克赫蒂亚莫夫(Pavel Akhtyamov)(MITP系APT系硕士)和弗拉基米尔·库兹涅佐夫(PG德米多夫·雅罗斯拉夫尔州立大学的助教),他们已经开发并维护了用于本课程中的实际作业。 -Asya Roitberg,Eugene Baulin,Marina Sudarikova。这些人日夜不睡觉,这会让您的学习体验富有成效,顺畅而令人兴奋。

课程标签

1人关注该课程

主题相关的课程