[MUSIC]. Okay. So, the next few segments I want to spend talking about NoSqL. And so, these systems are typically associated with building very large scalable web applications as opposed to analyzing data. Which is really the focus of this course. However, I think it's important to cover this topic for a few reasons. One is, you know, as a data scientist, you'll be manipulating data that is increasingly found in one of these variants of a NoSQL system. but, but also the, the systems and the terminology in this space are really influencing peoples thinking about how to deal with large scale data. And so, being cognizant of the landscape here, being cognizant of the major trends in history allow you to make informed decisions about this. You know, as a data scientist, you may be asked to make recommendations about what kind of platform to go with, to do your kind of ana, to do your analysis. And so, understanding the pros and cons of some of these systems and how they work can be pretty important, okay? You know? And then maybe third, the same concepts that we've been discussing in other segments, you know, sort of relational algebra. logical data dependence, you know, simple analytics going a long way. Simple scalable analytics indexing these ideas come up in this space as well. And so, it's yet another application of these concepts. And then maybe, finally the data scientists may be called upon to actually build some of the large scale systems as part of their work. It's not unheard of. When I mention in the first few segments that building data products is important part of data science. And so, when manifestation of data products is some of these large scale well application, so NoSQUl may or may not be a part of that. Fine, so let's get started. Alright, so where we are so far is we mentioned that data science may be has these three steps, data preparation at large scale, alright? And we talked about manipulation and munging, and data jujitsu, and so on. And that's all these step one. And then step two is analytics or actually running the model. And we're going to get there at, in the next week. And then third is communicating the results, interpreting and communicating the results. And this will be visualization will be a big part of the way we talk about communication, okay? So, we're still in the, this data preparation phase. And we're spending a little bit longer on that. Then, you know, say a third of the course, more than a third of the course. For the reason that I gave in the first few segments, that, really, this is the part that keeps people up at night. So, I want to make sure that you're armed with how to use databases to do this data munging. How to use MapReduce to do this data munging. also make the point that a lot of times even analytics itself can be pushed into these systems, which you saw in, in hopefully in, in assignment two in the database assignment, okay? So then, some of the key ideas from databases that we took away where this concept of the relation algebra comes up even extern, even outside of databases. It's not only found, you know, with SQL systems. And we'll see that again. And this notion of physical logical data independence comes up over and over again, maybe indexing. And then we talked about MapReduce, and we gave a lot of examples of some basic operations of MapReduce. And we started an assignment involving writing your own MapReduce programs, at least at the programming model level. And so, here we saw that part of the advantages of this system itself first at Google, and then in the implementation Hadoop was volt tolerance at scale. Another one was you didn't have to load the data. You could just work with it as is unlike data bases where you have to actually impose some sort of a scheme on it in order to even touch it for the first time. And so, that first step can be a dozy, okay? Extending this point, you're direct programming on in situ data. So, anything that comes at you, you can sort of, if you can write a program and process it. You can probably write a MapReduce program that can process it at scale. And that's a very powerful thing. I'll give you another way to put this is sort of single developer, right? You're kind of up and running within the hour, with MapReduce, and that was never really a property. The databases had, it was always sort of a difficult, it was a significant project to get a database installed and running, fine. So that the background, but we haven't talked about is this NoSQL system. And so, I want to use this table as a way to sort of organize the road map for this discussion. And so, what I've done here is tried to list out by feature, a bunch of relevant systems in this space. and right now, they're sort of sorted by time. And so, these features are admittedly selected by me for what the important ones are. But I don't think they'd be too controversial, and I don't think there's anything obviously missing here either. Okay, so going through these briefly. And a major one here is that it needs to scale to sort of thousands of machines. Then, there are some need to look up by a primary index. What I mean by that is by some key value, right? You can look up by some record ident, identifier. And then another feature that they may or may not have is the look at my secondary indexes. So, I mean here you can look up by some attribute that is not that key, okay? So, databases have this. For example, you can build an index on any attribute you want and the optimizer will take advantage of it. the third one, and this is the one we'll spend a lot of time on, it's sort of pretty fundamental to the motivation for y nodes equal systems. Sort of earn their own moniker or, or are different or, trans, transactions, okay? And you can also see that there's some complexity here that I'll try to explain as we go through this. So, it's not just a yes or no. It's kind of been a case by case, okay? And then, this field is whether or not these system support, essentially join. But I generalize that to analytic, and I'm perhaps guilty of making these two things almost synonymous. If you, if you can do some, if you can do joins, then there's a whole space of different kind of analytic things you can do. And if you can't do join, then you're leaving all that work up to the client and all you can do is retrieve data. And so, that's really cannot do join a key indicator of how much computation you can push into the system. And how much you have to sort of bring the data back out to, to the client. Okay, and then integrity constraints. I debate about calling these schemas of integrity constraints. But as I think, we'll see arguably some of these systems do indeed have a schema, but may or may not actually enforce that schema. And so, it's just to avoid the confusion, I'm going to call this integrity, okay? So, this is kind of hard schemas, if you will, okay? And in views, which were if you remember is I'm a declaring that to be synonymous with logical data and dependence. And then finally, is there some sort of decorative language or algebraic way of programming this thing or is it really just sort of a low level simple operation API? Okay. All right, so a couple of caveats here. One is I haven't included any parallel data bases on this list at all. Although there's absolutely no reason why you couldn't include them here. They'd tend to have a lot of check marks across this row but since the focus is NoSQL systems, I am leaving those out. And then the other category I made is that individuals sales in this table may be debatable depending on how you interpret the columns. So, this isn't necessarily meant to be Hard and fast rules. But again, I don't think they'll be wildly controversial either, okay? So, one of the first stories you can tell by staring at this table is that relational databases, you know, have been around for quite a while. And have check marks sort of everywhere and all these new features. Except they weren't really ever shown to scale to lots and lots and lots of computers, right? Everything was sort of in the order of tens of machines, okay? So, why don't they scale? Well, we talked in the MapReduce segment about sort of re-performance and related, and sort of analyzing experimental results from a paper in 2009. Comparing sort of the benefits and strengths of MapReduce versus databases. So, maybe on re-performance there's an argument to be made that they did scale. But one, one area where they certainly didn't or least certainly weren't shown to scale it to, to this level was in updates. So not just the read, not just the analytic workload, but the transaction processing workloads. Alright.