Apache Spark for Data Science Cookbook

Apache Spark for Data Science Cookbook

by Padma Priya Chitturi
English | 2016 | ISBN: 1785880101 | 388 Pages | True PDF | 4.5 MB

Spark has emerged as the most promising big data analytics engine for data science professionals. The true power and value of Apache Spark lies in its ability to execute data science tasks with speed and accuracy. Spark's selling point is that it combines ETL, batch analytics, real-time stream analysis, machine learning, graph processing, and visualizations. It lets you tackle the complexities that come with raw unstructured data sets with ease.

This guide will get you comfortable and confident performing data science tasks with Spark. You will learn about implementations including distributed deep learning, numerical computing, and scalable machine learning. You will be shown effective solutions to problematic concepts in data science using Spark's data science libraries such as MLLib, Pandas, NumPy, SciPy, and more. These simple and efficient recipes will show you how to implement algorithms and optimize your work.

What you will learn:

- Explore the topics of data mining, text mining, Natural Language Processing, information retrieval, and machine learning.
- Solve real-world analytical problems with large data sets.
- Address data science challenges with analytical tools on a distributed system like Spark (apt for iterative algorithms), which offers in-memory processing and more flexibility for data analysis at scale.
- Get hands-on experience with algorithms like Classification, regression, and recommendation on real datasets using Spark MLLib package.
- Learn about numerical and scientific computing using NumPy and SciPy on Spark.
- Use Predictive Model Markup Language (PMML) in Spark for statistical data mining models.



[Fast Download] Apache Spark for Data Science Cookbook

Ebooks related to "Apache Spark for Data Science Cookbook" :
Tableau 10 Business Intelligence Cookbook
SAP Data Services 4.x Cookbook
Chinese Computational Linguistics and Natural Language Processing Based on Naturally Annotated Big D
Shared Encounters
Learning Social Media Analytics
Business Intelligence with SQL Server Reporting Services
Regression Analysis with Python
OCA/OCP Oracle Database 11g All-in-One Exam Guide: Exams 1Z0-051, 1Z0-052, 1Z0-053
Handbook on Data Centers
Complete MongoDB Course Series : Part 1- Introduction to NoSQL Databases & Installation
Copyright Disclaimer:
This site does not store any files on its server. We only index and link to content provided by other sites. Please contact the content providers to delete copyright contents if any and email us, we'll remove relevant links or contents immediately.