|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/spark-python-crush-course/
课程评论:没有评论
COURSE TITLE: 超速入門!【データサイエンスへの最初の一歩】PythonとSparkで学ぶデータ分析のための前処理と分散処理 一気見講座 COURSE OVERVIEW: This course, taught by an active data engineer, focuses on data engineering, the often time-consuming process of preparing and managing data for AI and machine learning, which accounts for over 80% of a data scientist's work. The course specifically covers data engineering using Apache Spark, a distributed processing engine increasingly becoming the standard for big data. It provides a comprehensive guide to using PySpark, the Python interface for Spark, for practical, real-world operations. This course is geared towards data engineers, not those seeking deep scientific or mathematical concepts. It is ideal for individuals already familiar with Python who aim to become engineers in the AI or big data fields and want to master data manipulation. Source code and explanations are available in a GitHub repository, with video lectures offering supplementary explanations alongside the GitHub materials. SYLLABUS: Not provided.
現役のデータエンジニアがレクチャーします!AIや機械学習を行う際に最も時間のかかる作業は、データの準備とそれらの管理です。これらの作業のことをデータエンジニアリングと呼びます。実に80%以上の時間をデータエンジニアリング(データサイエンスのための前処理など)に割いてるのが現状です。本コースではApache Sparkを使ったデータエンジニアリングについて学びます。ポイント:本コースでは分散処理のデファクトとなりつつあるSparkについて学びます。Apache Sparkはビッグデータ処理で多く使われている分散処理エンジンです。今回はPythonと組み合わせた実際の現場で使われるPySparkを使った操作を一挙にまとめました。特徴:データエンジニアリングよりの講座です。難しいいサイエンスや数学は出てきませんが、データの3職種のうちの一つである「データエンジニア」のためのコースです。普段Pythonを使っている方やこれからAIやビッグデータの分野にエンジニアとして参画してデータを自在に操りたいという方にはぴったりですソースコードや解説は以下のGitHubリポジトリにあります。動画内ではGitHubの資料に加え補足をしながら解説を進めています。