{
    "links":{},
    "photo":"https://s3.amazonaws.com/coursera/topics/datasci/large-icon.png",
    "courseFormat":"The class will consist of lecture videos about 8 to 10 minutes in length.\nThese will contain 1-2 integrated quizzes per video. Some of these videos\nwill be given by guest lecturers from the data science community.\n<br>\n<br>There will be no formal exams or standalone quizzes.\n<br>\n<br>There will be eight total assignments of which two are optional.\n<br>\n<br>We will provide a virtual machine equipped with all necessary software,\nbut you are permitted (and encouraged) to install software in your own environment as well.\n<br>\n<br>There will be four structured programming assignments: two in Python,\none in SQL, and one in R.\n<br>\n<br>There will also be two open-ended assignments graded by peer assessment:\none in visualization, and one in which you will participate\nin a Kaggle competition.\n<br>\n<br>Finally, there will be two optional assignments: One involving an open-ended\nreal-world project submitted by external organizations with real needs,\nand one involving processing a large dataset on AWS.",
    "smallIcon":"https://d1z850dzhxs7de.cloudfront.net/topics/datasci/small-icon.hover.png",
    "universityLogo":"",
    "video":"",
    "smallIconHover":"https://d1z850dzhxs7de.cloudfront.net/topics/datasci/small-icon.hover.png",
    "shortDescription":"Join the data revolution. Companies are searching for data scientists. This specialized field demands multiple skills not easy to obtain through conventional curricula. Introduce yourself to the basics of data science and leave armed with practical experience extracting value from big data. #uwdatasci",
    "id":106,
    "estimatedClassWorkload":"10-12 hours/week",
    "previewLink":"https://class.coursera.org/datasci-001/lecture/preview",
    "universityLogoSt":"",
    "targetAudience":1,
    "courseSyllabus":"<i>Part 0: Introduction</i>&nbsp;<br>\n<ul>\n    <li>Examples, data science articulated, history and context, technology\n        landscape</li></ul>\n<i>Part 1: Data\nManipulation at Scale</i><br>\n<ul>\n    <li>Databases and the relational algebra&nbsp;\n        <br>\n    </li>\n    <li>Parallel databases, parallel query processing, in-database analytics&nbsp;</li><li>MapReduce, Hadoop, relationship to databases, algorithms, extensions,\n        languages &nbsp;</li><li>Key-value stores and NoSQL; tradeoffs of SQL and NoSQL</li></ul>\n<i>Part 2: Analytics\n</i><br>\n<ul>\n    <li>Topics in statistical modeling: basic concepts, experiment design, pitfalls<br>\n    </li>\n    <li>Topics in machine learning: supervised learning (rules, trees, forests, nearest neighbor, regression), optimization (gradient descent and variants), unsupervised learning</li></ul>\n<i>Part 3: Communicating Results </i><br>\n<ul>\n    <li>Visualization, data products, visual data analytics&nbsp;\n        <br>\n    </li>\n    <li>Provenance, privacy, ethics, governance&nbsp;</li></ul><i>Part 4: Special Topics</i><ul>\n    <li>Graph Analytics: structure, traversals, analytics, PageRank, community detection, recursive queries, semantic web</li><li>Guest Lectures</li>\n</ul>",
    "aboutTheCourse":"Commerce and research are being transformed by data-driven discovery and\nprediction. Skills required for data analytics at massive levels \u2013 scalable\ndata management on and off the cloud, parallel algorithms, statistical\nmodeling, and proficiency with a complex ecosystem of tools and platforms\n\u2013 span a variety of disciplines and are not easy to obtain through conventional\ncurricula. Tour the basic techniques of data science, including both SQL\nand NoSQL solutions for massive data management (e.g., MapReduce and contemporaries),\nalgorithms for data mining (e.g., clustering and association rule mining),\nand basic statistical modeling (e.g., linear and non-linear regression).<br>",
    "largeIcon":"https://d15cw65ipctsrr.cloudfront.net/15/f86bc0352d11e485649b7db944a974/large-icon.png",
    "suggestedReadings":"There will be selected readings each week. &nbsp;<br><br>We recommend, but do not require, that students refer to the book <a href=\"http://books.google.com/books?id=OefRhZyYOb0C&amp;hl=en\" target=\"_blank\">Mining of Massive Datasets by Anand Rajaraman and Jeff Ullman</a>",
    "videoId":"",
    "faq":"<b><br>Will I get a Statement of Accomplishment after completing this class?</b>&nbsp;\n<br>\n<br>Yes. Students who successfully complete the class will receive a Statement\nof Accomplishment signed by the instructor.&nbsp;\n<br>\n<br><b>What resources will I need for this class?\n<br></b>\n\n<br>For this course, you will need an Internet connection and either a) the\nability to run a virtual machine locally or b) the ability and knowledge\nto install the appropriate software yourself. &nbsp;The software will include Python 2.7 (including various libraries), R, SQLite (or another database you are comfortable using). &nbsp;You will also have the opportunity to install and work with Hadoop, but for logistics reasons, we will not require its use in an assignment. &nbsp;Some assignments will be open-ended.<br>\n<br><b>What level of programming experience should I have? </b>\n\n<br>\n<br>We expect intermediate programming experience in some language and some familiarity\nwith database concepts. &nbsp;There will be programming assignments, but\nthese are not designed to test knowledge of the language itself and will\nnot involve using any esoteric features. &nbsp;The languages we will use\nare Python, R, and SQL.",
    "shortName":"datasci",
    "instructor":"Bill Howe",
    "name":"Introduction to Data Science",
    "subtitleLanguagesCsv":"en",
    "recommendedBackground":"We expect you to have intermediate programming experience and familiarity with\ndatabases, roughly equivalent to two college courses. &nbsp;We will have four programming assignments: two in Python, one\nin SQL, and one in R. The target audience is undergraduate students across\ndisciplines who wish to build proficiency working with large datasets and a range of tools to\nperform predictive analytics.<br><br>After taking this course, you may be interested in participating in the three-course&nbsp;<a href=\"http://www.pce.uw.edu/certificates/data-science.html\" target=\"_blank\" title=\"Link: http://www.pce.uw.edu/certificates/data-science.html\">Certificate in Data Science</a>&nbsp;offered through the <a href=\"http://www.pce.uw.edu/\" target=\"_blank\" title=\"Link: http://www.pce.uw.edu/\">University of Washington Professional and Continuing Education program</a>. &nbsp;This online course will provide an overview and introduction to the more extensive material covered in that program, which offers classroom-based instruction by data scientists from Microsoft and other Seattle players, networking opportunities with peers, case studies from the \"front lines,\" and deep dives into selected topics.",
    "aboutTheInstructor":"<img class=\"coursera-instructor-thumb\" style=\"width: 250px;\" src=\"https://s3.amazonaws.com/coursera/topics/datasci/instructor-1.png\">Bill Howe is the Director of Research for Scalable Data Analytics at the UW eScience Institute and holds an Affiliate Assistant Professor appointment in Computer Science & Engineering, where he leads a group studying data management, analytics, and visualization systems for science applications. Howe has received awards from Microsoft Research and honors for papers in scientific data management, and serves on a number of program committees, organizing committees, and advisory boards in the area, including the advisory board of the Data Science certificate program at UW. He holds a Ph.D. in Computer Science from Portland State University and a Bachelor's degree in Industrial & Systems Engineering from Georgia Tech."
}