UCSanDiegoX: Big Data Analytics Using Spark

Learn how to analyze large datasets using Jupyter notebooks, MapReduce and Spark as a platform.

4.4|Reviews (11)
radio_button_checkededX Subscription
₹12,510/yr
Get this course and thousands more with a edX Subscription subscription.
radio_button_uncheckedIndividual Course
290500
Keep the course forever with lifetime access and receive a certificate.
TeamsBilled annually • 14-day free trial
₹33,026
EnterpriseFor organizations with 50+ learners
Custom Pricing

Why choose Essentials (Small Teams)?

check
1 Academy,16 Self-Paced Courses
check
Self-Paced Learning ,Employee Upskilling
verifiedGet this course for free with the Coursera Plus subscription.
✓ Compare courses before making a decision
Check Latest Price →
Price may vary. Check latest price on provider site.

Course Insight

Designed for advanced/expert practitioners. Designed for experienced practitioners. We recommend having a solid grasp of Computer Science fundamentals before starting this specialization.

Advanced LevelCertification IncludedSelf-Paced LearningHands-On Learning

SKILLS TO
MASTER

Computer Science Basics
Fundamental principles and concepts
Practical ApplicationTrending
Real-world project implementation
Best Practices
Industry standard workflows and guidelines
Problem Solving
Core Concepts
Implementation
Workflow Integration
Optimization
Careers:Data Scientist, Data Analyst, Machine Learning Engineer.

Quick Facts

Below sections are verified from last major sync. For real-time updates and today's latest lectures, Check official page here.

What You’ll Learn

  • Programming Spark using Pyspark.
  • Identifying the computational tradeoffs in a Spark application.
  • Performing data loading and cleaning using Spark and Parquet.
  • Modeling data through statistical and machine learning methods.
See side-by-side differences in what you’ll learn

Description

In data science, data is called "big" if it cannot fit into the memory of a single standard laptop or workstation.

The analysis of big datasets requires using a cluster of tens, hundreds or thousands of computers. Effectively using such clusters requires the use of distributed files systems, such as the Hadoop Distributed File System (HDFS) and corresponding computational models, such as Hadoop, MapReduce and Spark.

In this course, part of the Data Science MicroMasters program, you will learn what the bottlenecks are in massive parallel computation and how to use spark to minimize these bottlenecks.

You will learn how to perform supervised an unsupervised machine learning on massive datasets using the Machine Learning Library (MLlib).

In this course, as in the other ones in this MicroMasters program, you will gain hands-on experience using PySpark within the Jupyter notebooks environment.

See how this course compares with alternatives

FAQs

Top Alternatives

Highly-rated courses worth your attention

Managing an Agile Team
4.7· Flexible schedule
Intermediate
COURSERA PLUS
₹7,499/yr₹13,99946% OFF|₹2,099/mo
Managing Google Workspace
4.7· 5 Hrs (approximately)
Beginner
COURSERA PLUS
₹7,499/yr₹13,99946% OFF|₹2,099/mo
Python Data Analysis
4.7· 9 Hrs (approximately)
Beginner
COURSERA PLUS
₹7,499/yr₹13,99946% OFF|₹2,099/mo
SAS Macro Language
4.8· 20 Hrs (approximately)
Intermediate
COURSERA PLUS
₹7,499/yr₹13,99946% OFF|₹2,099/mo
APIs
coursera
APIs
4.4· 20 Hrs (approximately)
Intermediate
COURSERA PLUS
₹7,499/yr₹13,99946% OFF|₹2,099/mo
Testing and Debugging Python
4.3· 55 minutes
Intermediate
COURSERA PLUS
₹7,499/yr₹13,99946% OFF|₹2,099/mo
UCSanDiegoX: Big Data Analytics Using Spark
4.4(11+ learners)