Distributed processing of Wikipedia history files using Hadoop and Spark
-
Updated
Jan 6, 2019 - Scala
Distributed processing of Wikipedia history files using Hadoop and Spark
An end-to-end data engineering and analysis project to process a large-scale movie dataset, derive actionable business insights using Apache Spark, and build a content-based recommendation system.
To associate your repository with the distributed-processing topic, visit your repo's landing page and select "manage topics."