Course Logistics Help Center Learn more.

Course Overview

Commerce and research is being transformed by data-driven discovery and prediction. Skills required for data analytics at massive levels – scalable data management on and off the cloud, parallel algorithms, statistical modeling, and proficiency with a complex ecosystem of tools and platforms – span a variety of disciplines and are not easy to obtain through conventional curricula. Tour the basic techniques of data science, including both SQL and NoSQL solutions for massive data management (e.g., MapReduce and contemporaries), algorithms for data mining, and basic statistical modeling.

We are designing this course to have minimal prerequisites, yet still cover some advanced topics spanning a variety of areas: large-scale systems, important analysis techniques, visualization, and some topics in the curation, management, and ethics of big data.

Our approach is to focus on concepts and abstractions that come up over and over again, as well as teach a set of concrete, practical techniques that every data scientist should know. We are less focused on teaching the details of specific software tools (although the assignments will be hands-on.)

After completing this class, you can expect to be an "advanced beginner" in the area of data science, fully prepared to dive in to deeper concepts or to apply the material in practice.

Final grades will be determined by scores on the assignments, with a component of optional "make up credit" that can be earned by completing the optional assignments, completing the real world project, or participating constructively in the forums.

A statement of accomplishment will be give to those students who earn 70% of the points available from the assignments, either through the assignments themselves or through earning make-up credit.  

Learning Objectives

By the end of this course, students should be able to: 

  • Explain the major trends in technology, business, and science behind the terms data science and big data. 
  • Explain the relative strengths and weaknesses of the scalable data analytics platforms in use today, including databases, MapReduce-based platforms, NoSQL solutions, cloud services. 
  • Given a specification of a large-scale analytics problem, propose a reasonable approach to solve it. 
  • Describe a selection of data mining techniques and how they can be implemented scalably. 
  • Identify and articulate the recurring concepts and abstractions independently of specific technologies.

Course Structure

The course will consist of video lectures and hands-on assignments. Readings will also be associated with each week of material.

Student Assessment

Satisfactory completion will depend on completion of all required assignments. We hope you will augment your experience with active participation in the forum discussions.

Assignments

There will be four programming assignments: two in Python, one in R, and one in SQL. You are not expected to have proficiency in these languages, and the assignments will not exercise any esoteric features. You are expected to have programming experience in some language.

You will be provided with a virtual machine with all necessary software installed, though you are free to install software yourself if you prefer.

We will use Python 2.7 (not 3.3!), which you can download from the Python website.

In addition to the programming assignments, there will be two assignments graded through peer assessment. One will entail constructing a visualization using a product called Tableau. The other and an open-ended mini-project in which you will compete in a kaggle.com competition.

There will also be two optional assignments: a project where you will work on a real data science problem proposed by a third party organization with real needs, and a large-scale data analysis problem in which you will process ~1TB of data using Amazon Web Services. (The cost for the AWS resources is typically less than $20, but must be paid by those students who choose to complete the assignment -- we can't afford to pay for 45000 students!)

Textbook

None. All reading materials will be on the website. We will use a combination of relevant research papers, documentation from relevant systems, and readings from the web.

Prerequisites

You are expected to have intermediate proficiency in programming, databases, and basic statistics.

If you are completely new to Python, a good place to start is codeacademy.com, focusing on lessons 1-9, plus lesson 12 (file IO):

You may also find the tutorial on the Python website to be valuable.

We will use R primarily for its libraries in statistics and machine learning rather than as a general purpose programming environment. This class is less about specific tools and more about concepts, abstractions and techniques; we will not cover the details of the R libraries or the language itself.

You can download R from the website.

A good introductory book is also available on the website:


Created Tue 24 Jul 2012 10:31 AM CEST
Last Modified Mon 7 Jul 2014 1:11 AM CEST