Learn more
Learn more
Browse all discussions »
The Kaggle evaluation due date has been extended by at least 48 hours due to a problem with the evaluation page.
We've sent a note to coursera.
The final set of lectures on graph analytics has been released!
I'll also take this opportunity to remind you all of the outstanding assignment deadlines, including the submission for the external project (due tomorrow), the peer reviews for the Kaggle assignment (to open on September 3rd), and the submission for the Tableau assignment (due September 5th).
The (correct) visualization assignment using Tableau has been released. This is the last assignment of the course, and is optional.
(There was an incomplete programming assignment inadvertently made visible for the last couple of days -- VERY sorry for the confusion this may have caused. It has been removed and the correct assignment has been posted.)
The assignment will be graded by peer review.
You can run Tableau directly on a Windows or a Mac. If you use linux, the assignment provides several options for obtaining access to a Windows VM.
Everyone will have free access to Tableau until October, thanks to Tableau's support of this assignment.
As usual, "optional" means that you can get full credit for the course without completing this assignment, but that your score on this assignment could still help improve your grade if you lost points on other assignments.
In this assignment, you will create a series of visualizations and use them to explore a dataset.
While statistical programming environments such as R, MATLAB, Python (with appropriate llibraries such as matplotlib), Octave, SAS, and SPSS offer extensive visualization capabilities and offer some support for interactivity, they were originally designed to provide maximum control in producing publication-quality static images. In contrast, as a data scientist, you will often be performing exploratory visual analytics -- using visualization not as a presentation tool, but as an analysis tool.
Moreover, you will often be asked to produce interactive data products -- for example, dashboards -- that your stakeholders can use to answer their own questions.
In these scenarios, many find that Tableau and other visual analytics tools can dramatically improve productivity relative to low-level programming environments where details must be coded by hand.
For the purposes of this class, Tableau offers another benefit in that it is relatively easy to get started using regardless of your background, allowing you to spend more time considering the visualization principles being employed and less time wrestling with code.
The learning objectives for this assignment are to gain experience with Tableau, gain experience using visualization for data exploration, and gain experience using interactive visualizations to tell a story.
Tableau is also a lot of fun -- you can create some sophisticated visualizations in just a few seconds.
The R assignment is now available.
The assignment is implemented as a QUIZ rather than a programming assignment; make sure you look under "Quizzes" in the navigation bar
As always, perform a git pull on the course repository when you begin the assignment.
This is the first time we've offered this assignment, so please report problems!
We will release a new R assignment tonight or tomorrow where you asked to classify ocean microbes based on optical properties. The data is very real and the problem is very real, so we think it will be an engaging assignment. You will be using methods that we discussed in class, and reasoning about their results. As always, the goal is not to teach R the language, but to provide an opportunity to get some hands-on experience with popular methods and tools.
Sorry for the delay in releasing the assignment; the due date will be appropriately delayed as well!
As you may have noticed, there was no assignment posted this week.
Spend some time on the optional assignments and projects!
Tonight we released the peer review rubric for the optional external real-world project. You can find the assignment under "Peer Assessments" from the course web page.
Like the Kaggle assignment, you are asked to describe the problem and solution and submit them for peer review.
As an optional assignment, your score will not affect your ability to earn the certificate, except as a one of several indications of participation. That said, if you do choose to participate, do take the assignment seriously and complete the submission to the best of your ability/
As we've discussed, data science can be considered in terms of three steps: preparing the data, running some statistical model, and communicating the results. The latter step is the most important -- nobody is going to make business decisions using your results if you can't explain the argument properly. This assignment and the Kaggle assignment both ask you to focus on written communication to convey technical results -- do not underestimate the importance of this skill!
In the assignment released tonight, you are asked to participate in a Kaggle competition and write up your solution for peer review.
The goals for this assignment are to have you participate in a real prediction task with a concrete outcome, try out Kaggle, and earn some experience writing up your technical results with enough clarity that others can reproduce them.
Your peer reviewers will not actually be expected to reproduce your results, but you should give them enough information to be able to do so!
Reproducibility is paramount in data science, as it is in all science --- too often results may be subtly dependent on undocumented steps you took to prepare the data rather than on the method itself.
Have fun with this assignment -- Kaggle is a blast.
We've extended the deadline for the Twitter assignment since the grader was not available for the last several hours.
Thanks again for all your hard work on the assignments!
Yesterday, we released lectures on MapReduce as programming model, with some background and context on scalable algorithms. We started from the basics, not wanting to assume all students have been asked to think about algorithms, parallel or otherwise. We do not focus too much on MapReduce as a system; the emphasis is on "thinking in MapReduce" in order to understand how to develop algorithms that use it.
Today, we released the next assignment, which asks you to write MapReduce programs directly in Python. In this assignment, there's no need to worry about parallelism, large-scale data, java, cluster configuration, or any other details -- the goal is just to learn to use the MapReduce abstraction to manipulate data.
Assignment 2 has been released.
This assignment will be due in two weeks, with an additional two weeks before the hard deadline.
Good luck!
Hope you enjoyed the first week of the course. The discussions in the forums have been a fantastic read --- there are some incredibly thoughtful solutions to the assignment being explored!
We've released new lectures for week 2 on relational models and languages. You may be familiar with relational databases, but the emphasis here is on the underlying foundations behind the technology, especially the relational algebra, and less on databases themselves. Why? Because these foundations have begun to be applied in all sorts of situations, including those that have little to do with databases themselves.
Consider that the R library dplyr, the Python library pandas, several languages over Hadoop, and a number of google systems all support features and interfaces derived from the relational algebra. The relational algebra is emerging as a fundamental concept in data science, one that transcends any particular tool, language or implementation -- learn to recognize it and use it!
The next assignment will be released tomorrow afternoon, but note that you still have a week until the deadline of the first assignment (and another two weeks after that to earn 80% credit before the hard deadline after which solutions will not be accepted.)
Thanks again for your hard work,
Introduction to Data Science staff
We hope you've enjoyed the first week of the course. As mentioned earlier, we're including a component of the course that will help organizations and students to come together and work on real-world data science projects. The purpose of these projects is to provide students a chance to apply what they learn throughout this course to real projects and for organizations to get a little more help in making sense of their data.
To organize this component of the course, we have partnered with Coursolve. Coursolve is a company working to facilitate collaborations between students seeking real-world experience and organizations with real-world needs. To participate in the optional real-world project, we recommend you join Coursolve as a first step -- you will have better access to projects and better support during the process. If you would prefer not to join Coursolve, you can still participate by browsing the forum here on the course website to identify potential projects, or source your project offline.
If you are interested in posting a real-world project on behalf of your organization, you will fill out the form on the Coursolve website, making sure to select "Introduction to Data Science" at the bottom to ensure other students can find your challenge. (You may also choose to post your project on the Organizations Seeking Assistance sub-forum.)
Some collaborations from last year's course included mining Twitter data to improve the US healthcare system and visualizing vehicle data in collaboration with a large motor company.
When posting projects, please ensure that you have adequate capacity to work closely with the students that express interest in your project.
If you are interested in working on a real-world project as a student, you can first select "work on this course project" on the course project page and then browse needs to help find a project that interests you. (Again, if you would prefer not to join Coursolve, you can still participate by browsing the Organizations Seeking Assistance forum here on the course website to identify potential projects, or source your project offline.)
We hope you will enjoy using what you learn to conceptualize and solve real-world data science problems!
--Introduction to Data Science Staff
The grader for assignment 1 has been deployed, and submissions are now being processed! Sorry for the delay and the silence, and thanks for your persistence.
For Part 1, the assignment asks for the file name to be "output.txt" in the submission form, but asks for the name to be "problem_1_submission.txt" in the assignment instructions -- sorry for the inconsistency. The grader expects "problem_1_submission.txt" and I've updated the assignment submission form to reflect this.
The system should now be reporting a useful error message for this case. Previously, the system appeared to be reporting "Could not run submission tests," which was utterly unhelpful.
The first assignment has been released, with a due date of July 15 at 5:00 pm. The hard deadline is two weeks after that; submissions after the due date and before the hard deadline receive a 20% penalty. Submissions after the hard deadline will not be accepted.
You'll have two weeks to work on it, but be aware that the next assignment will be released next Tuesday. So on any given week except the first and last, there will be two assignments underway.
The grader for this assignment will not be live until tomorrow, however --- you won't receive feedback on submissions until then.
In this assignment, you'll be using Python to work with the live stream from Twitter.
You are free to discuss the assignment on the forums. Please don't paste solutions!
The assignment will be graded automatically by comparing results with prepared files. You will not need to reproduce our exact solution; there are many right answers, and some of the problems are open-ended. You will need to make sure your solution produces sensible results -- we think we have a pretty good way of checking this.
We have several learning objectives for this assignment:
1) We want you to gain experience working with Twitter data, which is useful in its own right for a broad array of data science applications. It's potentially a nice line item to put on your resume/CV.
2) We want you to get accustomed to open-ended problems, and being resourceful in solving them. So the tasks to complete this assignment won't be spelled out step by step, and you shouldn't get discouraged if you feel like you don't know exactly what you should do next. Explore the data, explore Python if it is new to you, ask questions on the forum, and be bold.
3) We want you to become fearless about digging into research papers. These papers can sometimes be dense, but the underlying ideas may often be pretty simple. So one of the assignment problems cites a research paper that provides a solution to one of the problems (though not the only solution; you can ignore the paper altogether and still complete the assignment!)
A few hints:
Problem 3: You don't necessarily need to use the formula in the paper. The idea is to determine the sentiment of a tweet using the known words, then determine the sentiment of the other words using the tweet.
Problem 5: There are a few different ways you might try to assign a "state" to a tweet. Don't worry about getting the "right" way. The user is associated with location information, and the tweet is also associated with location information. Remember that real data is dirty -- not every tweet will have every field filled out. That's ok.
Problem 6: The hashtags are already parsed out for you in the data structure; you just need to find them by looking in the API documentation.
Good luck!
Reminders
Course Calendar ICS
If you use a calendar that accepts .ics files (ex: google, ical), then import the URL below to see due dates for all quizzes and assignments
Upcoming Deadlines
Recent Discussions
Thread title |
|---|
|
Announcements
Kaggle Evaluation Due Date extended
The Kaggle evaluation due date has been extended by at least 48 hours due to a problem with the evaluation page.
We've sent a note to coursera.
Mon 8 Sep 2014 8:25 AM CEST
Final lecture set released
The final set of lectures on graph analytics has been released!
I'll also take this opportunity to remind you all of the outstanding assignment deadlines, including the submission for the external project (due tomorrow), the peer reviews for the Kaggle assignment (to open on September 3rd), and the submission for the Tableau assignment (due September 5th).
Wed 27 Aug 2014 8:55 AM CEST
Visualization Assignment released (and clarification)
The (correct) visualization assignment using Tableau has been released. This is the last assignment of the course, and is optional.
(There was an incomplete programming assignment inadvertently made visible for the last couple of days -- VERY sorry for the confusion this may have caused. It has been removed and the correct assignment has been posted.)
The assignment will be graded by peer review.
You can run Tableau directly on a Windows or a Mac. If you use linux, the assignment provides several options for obtaining access to a Windows VM.
Everyone will have free access to Tableau until October, thanks to Tableau's support of this assignment.
As usual, "optional" means that you can get full credit for the course without completing this assignment, but that your score on this assignment could still help improve your grade if you lost points on other assignments.
In this assignment, you will create a series of visualizations and use them to explore a dataset.
While statistical programming environments such as R, MATLAB, Python (with appropriate llibraries such as matplotlib), Octave, SAS, and SPSS offer extensive visualization capabilities and offer some support for interactivity, they were originally designed to provide maximum control in producing publication-quality static images. In contrast, as a data scientist, you will often be performing exploratory visual analytics -- using visualization not as a presentation tool, but as an analysis tool.
Moreover, you will often be asked to produce interactive data products -- for example, dashboards -- that your stakeholders can use to answer their own questions.
In these scenarios, many find that Tableau and other visual analytics tools can dramatically improve productivity relative to low-level programming environments where details must be coded by hand.
For the purposes of this class, Tableau offers another benefit in that it is relatively easy to get started using regardless of your background, allowing you to spend more time considering the visualization principles being employed and less time wrestling with code.
The learning objectives for this assignment are to gain experience with Tableau, gain experience using visualization for data exploration, and gain experience using interactive visualizations to tell a story.
Tableau is also a lot of fun -- you can create some sophisticated visualizations in just a few seconds.
Fri 22 Aug 2014 8:00 PM CEST
R assignment available (under Quizzes)
The R assignment is now available.
The assignment is implemented as a QUIZ rather than a programming assignment; make sure you look under "Quizzes" in the navigation bar
As always, perform a git pull on the course repository when you begin the assignment.
This is the first time we've offered this assignment, so please report problems!
Thu 14 Aug 2014 12:28 PM CEST
R Assignment to be released tonight
We will release a new R assignment tonight or tomorrow where you asked to classify ocean microbes based on optical properties. The data is very real and the problem is very real, so we think it will be an engaging assignment. You will be using methods that we discussed in class, and reasoning about their results. As always, the goal is not to teach R the language, but to provide an opportunity to get some hands-on experience with popular methods and tools.
Sorry for the delay in releasing the assignment; the due date will be appropriately delayed as well!
Wed 13 Aug 2014 7:45 PM CEST
No Assignment this Week!
As you may have noticed, there was no assignment posted this week.
Spend some time on the optional assignments and projects!
Sat 9 Aug 2014 2:45 AM CEST
External Project Peer Review Released
Tonight we released the peer review rubric for the optional external real-world project. You can find the assignment under "Peer Assessments" from the course web page.
Like the Kaggle assignment, you are asked to describe the problem and solution and submit them for peer review.
As an optional assignment, your score will not affect your ability to earn the certificate, except as a one of several indications of participation. That said, if you do choose to participate, do take the assignment seriously and complete the submission to the best of your ability/
As we've discussed, data science can be considered in terms of three steps: preparing the data, running some statistical model, and communicating the results. The latter step is the most important -- nobody is going to make business decisions using your results if you can't explain the argument properly. This assignment and the Kaggle assignment both ask you to focus on written communication to convey technical results -- do not underestimate the importance of this skill!
Thu 31 Jul 2014 8:00 AM CEST
Kaggle Assignment Released
In the assignment released tonight, you are asked to participate in a Kaggle competition and write up your solution for peer review.
The goals for this assignment are to have you participate in a real prediction task with a concrete outcome, try out Kaggle, and earn some experience writing up your technical results with enough clarity that others can reproduce them.
Your peer reviewers will not actually be expected to reproduce your results, but you should give them enough information to be able to do so!
Reproducibility is paramount in data science, as it is in all science --- too often results may be subtly dependent on undocumented steps you took to prepare the data rather than on the method itself.
Have fun with this assignment -- Kaggle is a blast.
Thu 31 Jul 2014 7:28 AM CEST
Lectures, Assignments, and more
Quick update on recent content:
Last week, we released new lectures on NoSQL with a goal of giving an overview of what's going on in the space and how the different buzzwords and systems fit together.
We also released an optional AWS assignment that gives you the opportunity to process a large graph dataset consisting primarily of social network information.
Yesterday, we released some lectures on selected topics in statistical analysis. Our goal here was not to try and give a compressed treatment of a stats 101 course -- besides not being feasible in the time allotted, we thought this approach would be a bit dry. Instead, we tried to identify topics that are not typically covered in early statistics courses, but should be.
The assignment we will release today or tomorrow will involve participating in a Kaggle competition; your peers will review an English description of your solution for comprehensibility and soundness, but your actual score on the competition will not affect your score.
Last week, we released new lectures on NoSQL with a goal of giving an overview of what's going on in the space and how the different buzzwords and systems fit together.
We also released an optional AWS assignment that gives you the opportunity to process a large graph dataset consisting primarily of social network information.
Yesterday, we released some lectures on selected topics in statistical analysis. Our goal here was not to try and give a compressed treatment of a stats 101 course -- besides not being feasible in the time allotted, we thought this approach would be a bit dry. Instead, we tried to identify topics that are not typically covered in early statistics courses, but should be.
The assignment we will release today or tomorrow will involve participating in a Kaggle competition; your peers will review an English description of your solution for comprehensibility and soundness, but your actual score on the competition will not affect your score.
Tue 29 Jul 2014 11:00 PM CEST
Twitter Assignment deadline extended
We've extended the deadline for the Twitter assignment since the grader was not available for the last several hours.
Wed 16 Jul 2014 5:45 PM CEST
Lectures released on Scalability and MapReduce
Thanks again for all your hard work on the assignments!
Yesterday, we released lectures on MapReduce as programming model, with some background and context on scalable algorithms. We started from the basics, not wanting to assume all students have been asked to think about algorithms, parallel or otherwise. We do not focus too much on MapReduce as a system; the emphasis is on "thinking in MapReduce" in order to understand how to develop algorithms that use it.
Today, we released the next assignment, which asks you to write MapReduce programs directly in Python. In this assignment, there's no need to worry about parallelism, large-scale data, java, cluster configuration, or any other details -- the goal is just to learn to use the MapReduce abstraction to manipulate data.
Wed 16 Jul 2014 10:13 AM CEST
Assignment 2 Released
Assignment 2 has been released.
In this assignment, you will be writing SQL queries to perform tasks that are not typically associated with SQL --- simple text analytics and linear algebra. Take a look at the lectures, the readings, and the assignment itself to understand why this might be a good idea!
You will use sqlite for this assignment, which is widely available and very easy to install. It has some limitations, but we won't bump into those in this assignment.
This assignment will be due in two weeks, with an additional two weeks before the hard deadline.
Good luck!
Wed 9 Jul 2014 3:00 AM CEST
Lectures Released: Relational Models and Languages
Hope you enjoyed the first week of the course. The discussions in the forums have been a fantastic read --- there are some incredibly thoughtful solutions to the assignment being explored!
We've released new lectures for week 2 on relational models and languages. You may be familiar with relational databases, but the emphasis here is on the underlying foundations behind the technology, especially the relational algebra, and less on databases themselves. Why? Because these foundations have begun to be applied in all sorts of situations, including those that have little to do with databases themselves.
Consider that the R library dplyr, the Python library pandas, several languages over Hadoop, and a number of google systems all support features and interfaces derived from the relational algebra. The relational algebra is emerging as a fundamental concept in data science, one that transcends any particular tool, language or implementation -- learn to recognize it and use it!
The next assignment will be released tomorrow afternoon, but note that you still have a week until the deadline of the first assignment (and another two weeks after that to earn 80% credit before the hard deadline after which solutions will not be accepted.)
Thanks again for your hard work,
Introduction to Data Science staff
Tue 8 Jul 2014 4:10 AM CEST
Connecting with others to address real-world data science needs
We hope you've enjoyed the first week of the course. As mentioned earlier, we're including a component of the course that will help organizations and students to come together and work on real-world data science projects. The purpose of these projects is to provide students a chance to apply what they learn throughout this course to real projects and for organizations to get a little more help in making sense of their data.
To organize this component of the course, we have partnered with Coursolve. Coursolve is a company working to facilitate collaborations between students seeking real-world experience and organizations with real-world needs. To participate in the optional real-world project, we recommend you join Coursolve as a first step -- you will have better access to projects and better support during the process. If you would prefer not to join Coursolve, you can still participate by browsing the forum here on the course website to identify potential projects, or source your project offline.
If you are interested in posting a real-world project on behalf of your organization, you will fill out the form on the Coursolve website, making sure to select "Introduction to Data Science" at the bottom to ensure other students can find your challenge. (You may also choose to post your project on the Organizations Seeking Assistance sub-forum.)
Some collaborations from last year's course included mining Twitter data to improve the US healthcare system and visualizing vehicle data in collaboration with a large motor company.
When posting projects, please ensure that you have adequate capacity to work closely with the students that express interest in your project.
If you are interested in working on a real-world project as a student, you can first select "work on this course project" on the course project page and then browse needs to help find a project that interests you. (Again, if you would prefer not to join Coursolve, you can still participate by browsing the Organizations Seeking Assistance forum here on the course website to identify potential projects, or source your project offline.)
We hope you will enjoy using what you learn to conceptualize and solve real-world data science problems!
--Introduction to Data Science Staff
Mon 7 Jul 2014 6:02 AM CEST
Assignment 1 grading update
The grader for assignment 1 has been deployed, and submissions are now being processed! Sorry for the delay and the silence, and thanks for your persistence.
For Part 1, the assignment asks for the file name to be "output.txt" in the submission form, but asks for the name to be "problem_1_submission.txt" in the assignment instructions -- sorry for the inconsistency. The grader expects "problem_1_submission.txt" and I've updated the assignment submission form to reflect this.
The system should now be reporting a useful error message for this case. Previously, the system appeared to be reporting "Could not run submission tests," which was utterly unhelpful.
Thu 3 Jul 2014 8:05 AM CEST
First Assignment Released
The first assignment has been released, with a due date of July 15 at 5:00 pm. The hard deadline is two weeks after that; submissions after the due date and before the hard deadline receive a 20% penalty. Submissions after the hard deadline will not be accepted.
You'll have two weeks to work on it, but be aware that the next assignment will be released next Tuesday. So on any given week except the first and last, there will be two assignments underway.
The grader for this assignment will not be live until tomorrow, however --- you won't receive feedback on submissions until then.
In this assignment, you'll be using Python to work with the live stream from Twitter.
You are free to discuss the assignment on the forums. Please don't paste solutions!
The assignment will be graded automatically by comparing results with prepared files. You will not need to reproduce our exact solution; there are many right answers, and some of the problems are open-ended. You will need to make sure your solution produces sensible results -- we think we have a pretty good way of checking this.
We have several learning objectives for this assignment:
1) We want you to gain experience working with Twitter data, which is useful in its own right for a broad array of data science applications. It's potentially a nice line item to put on your resume/CV.
2) We want you to get accustomed to open-ended problems, and being resourceful in solving them. So the tasks to complete this assignment won't be spelled out step by step, and you shouldn't get discouraged if you feel like you don't know exactly what you should do next. Explore the data, explore Python if it is new to you, ask questions on the forum, and be bold.
3) We want you to become fearless about digging into research papers. These papers can sometimes be dense, but the underlying ideas may often be pretty simple. So one of the assignment problems cites a research paper that provides a solution to one of the problems (though not the only solution; you can ignore the paper altogether and still complete the assignment!)
A few hints:
Problem 3: You don't necessarily need to use the formula in the paper. The idea is to determine the sentiment of a tweet using the known words, then determine the sentiment of the other words using the tweet.
Problem 5: There are a few different ways you might try to assign a "state" to a tweet. Don't worry about getting the "right" way. The user is associated with location information, and the tweet is also associated with location information. Remember that real data is dirty -- not every tweet will have every field filled out. That's ok.
Problem 6: The hashtags are already parsed out for you in the data structure; you just need to find them by looking in the API documentation.
Good luck!
Wed 2 Jul 2014 2:01 AM CEST
Course Logistics
Welcome again to Introduction to Data Science!
The first set of lectures went live yesterday, and the first assignment will be available soon.
We will have one assignment per week for each of the first seven weeks, though some of them will be optional. Here is the schedule:
Each assignment will be due two weeks after it is released, with a hard deadline (and a 20% penalty) two weeks after the due date. Many courses make all assignments due at the end of the quarter; we've found it useful to stagger the deadlines to encourage continuous effort throughout the quarter and to avoid a backlog of work for you, me, and the autograder as the course winds down.
There will also be an optional External Real-World Project administered through Coursolve. We will send another announcement about these projects soon.
There will be no traditional quizzes nor a final exam --- we hope keep this material as hands-on as possible.
Final grades will be determined by scores on the assignments, with a component of optional "make up credit" that can be earned by completing the optional assignments, completing the real world project, or participating constructively in the forums.
Stay tuned for more information on the assignments!
Bill PS: If you are new to Python, now is the time to get started on some tutorials!
The first set of lectures went live yesterday, and the first assignment will be available soon.
We will have one assignment per week for each of the first seven weeks, though some of them will be optional. Here is the schedule:
- Week 1: Twitter Sentiment Analysis (Python)
- Week 2: In-Database Analytics (SQL)
- Week 3: MapReduce Concepts and Algorithms (Python)
- Week 4: (Optional) Large-scale data processing in the cloud (Pig, Hadoop, AWS)
- Week 5: Supervised Learning Roundup (R)
- Week 6: Visualization (Tableau and/or Javascript/D3)
- Week 7: Kaggle Competition
Each assignment will be due two weeks after it is released, with a hard deadline (and a 20% penalty) two weeks after the due date. Many courses make all assignments due at the end of the quarter; we've found it useful to stagger the deadlines to encourage continuous effort throughout the quarter and to avoid a backlog of work for you, me, and the autograder as the course winds down.
There will also be an optional External Real-World Project administered through Coursolve. We will send another announcement about these projects soon.
There will be no traditional quizzes nor a final exam --- we hope keep this material as hands-on as possible.
Final grades will be determined by scores on the assignments, with a component of optional "make up credit" that can be earned by completing the optional assignments, completing the real world project, or participating constructively in the forums.
Stay tuned for more information on the assignments!
Bill PS: If you are new to Python, now is the time to get started on some tutorials!
Tue 1 Jul 2014 10:00 PM CEST
Welcome to Introduction to Data Science!
Welcome to Introduction to Data Science!
The course website is now live, and we've posted the first set of videos and some other initial materials. We're thrilled to have you join us for this second session!
Before you do anything else, we ask that you complete the Pre-course Experience Survey, which is available from the Surveys tab on the navigation bar. This survey will not affect your completion of the course.
Course materials now available:
Welcome again to the course -- we hope you will work hard, enjoy the material, and learn a lot! Bill -- Bill Howe, PhD Associate Director, eScience Institute University of Washington
The course website is now live, and we've posted the first set of videos and some other initial materials. We're thrilled to have you join us for this second session!
Before you do anything else, we ask that you complete the Pre-course Experience Survey, which is available from the Surveys tab on the navigation bar. This survey will not affect your completion of the course.
Course materials now available:
- The first set of video segments are available providing some examples of data science, background on the terminology around data science, and some basic course logistics. We will spend a bit more time on this introductory material than we might in a typical course -- the field is relatively new, and we want to accommodate students from a variety of backgrounds.
- A reading list is available on the syllabus. This list is intended as an additional resource to complement the lectures and may change over time.
- Some instructions have been posted for accessing the class virtual machine and the class github repository. Both of these options are available from the left navigation bar. The github repository includes starter code and datasets for some assignments, but the instructions on how to complete each assignment will only be released when the assignment goes live.
- The first assignment is in Python, and will have you accessing and analyzing the live 1% twitter stream. We will provide a video walkthrough with the assignment to help you get started. Both the assignment and the walkthrough will be released tomorrow
Tuesday July 1 at 5:00 pm PT to be due onTuesday July 14, 2014 at 5:00 pm PT . - In this course, you will have the option to participate in a project with an external company to solve a real-world problem. Take a look at the initial description of this project and think about whether you'd be interested in participating as a student, or perhaps representing your organization by proposing a project. Participants will be asked to complete a survey.
Welcome again to the course -- we hope you will work hard, enjoy the material, and learn a lot! Bill -- Bill Howe, PhD Associate Director, eScience Institute University of Washington
Tue 1 Jul 2014 2:01 AM CEST