[MUSIC]. Okay. Yes, so going back to the grid we can highlight those systems that were based on MapReduce itself. You know, the MapReduce paper itself in 2004 and these language layers on top Pig and Hive in 2008, where Hive is SQL. And Pig is a relational algebra looking language that we'll talk about in some detail in the next few segments and Tenzing which is also SQL. And Impala, which is also SQL, or Tenzing is from Google and Impala is from a company called Cloudera, that's a pretty eager evangelist of MapReduce and Hadoop, based technologies in general, okay. So, one trend I think you see is that these declarative languages on top of the parallel processing primitive of MapReduce are really here to stay, right. So, people that were relatively against these kind of languages certainly doing it. Now, it's also fair to say that the enterprises in general have made a pretty significant investment in SQL expertise. So, even if they're attracted to the, to the advantages that Hadoop might bring, they're, they're pretty much demanding SQL. So, this may be a response to this inertia from having invested in SQL in the past. I think that is certainly true, however, it is also true that the desire for declarative languages reasonably well founded for reasons we've already talked about. Okay. So, you can put these systems on this time line and the only point I want to make about this, is that there is a bit of a gap between the paper in 2004 and the systems in 2008. But as soon as we had Hadoop the system itself, developed at Yahoo and, and released as an Apache open sourced project. You sort of immediately see an ecosystem start to emerge with extensions to it, that add these, in particular, add these languages on top. And so I think that the, the need for high level interface is motivated by how quickly they came around as soon as Hadoop wa, was, was out. And again, it, it didn't stop with these later systems a few years later. Okay. And actually you know, not on this page there's potentially hundreds of, if you include you know, research projects based on extensions to MapReduce. There are really, really a lot okay and so this is some of the most popular ones alright. So, another subset of this grid that you can look at is just the no SQL system. Now the whole last few segments have sensibly been about NoSQL but I've also included these kinds of analytic systems in here. MapReduce base systems and a few others for example Dremel and Spark and Shark. Dremel is a system from Google that is the back end of a query, as a service system called Google BigQuery, which is pretty nice. I, I recommend taking a look at it, you can sort of upload data and put it in there and it doesn't matter how big it is and you can kind of query it at, at very low latency speed. Spark and Shark, come from the AMPLab at Berkeley and are part of the Berkeley Data Analytics Stack, BDAS or BDAS. And Spark is a language label on top of, it's not MapReduce, but on top of a parallel processing system and Shark is a SQL layer even on top of that, okay. And so a couple of distinguishing features of Spark is it loads everything to memory. Process everything there when possible. Writing things out to disk only for fault tolerance reasons, so much, much, less often than MapReduce. and it also supports iterative processing which is pretty important and we're going to come back to that a little later in the course. And then Shark again is just SQL on top of, on top of this. All right. So, within these NoSQL systems one thing you can look at, there's not much I want to, excuse me. There's not much I want to say about this diagram. except that, you know, to point out that there's been sort of of a cambrian explosion. So, first, you had sort of memcached, which is again just a caching layer for really the real system. And the real system was a bunch of MySQL databases that weren't really working all that well for the requirements, they're having them used for. But you can bring this in the memory and keep it there, looking it up by name. And it was just like a performance enhancement, a free performance enhancement, if you invest in this system. The real you know, approach of throwing out everything you have and replacing it with a NoSQL system came a little later. So the only couple points I want to make is that it's been kind of a cambrian explosion of a different systems around this time. And that you know, this space in here is nowhere near as empty as it looks. I've picked out a few systems here, but really the emergence of new systems in the space hasn't really slowed down much at all. So, really filling out the design space here since around 2006. And the only other point I made is that these, these quite popular document oriented data models systems, CouchDB and MongoDB, have actually been around for quite a while. So, they've been, you know, they were released in 2005 and 2000 seven. So, sometimes I think they seem like newer systems, but, but they've have some maturity. Okay. So, then finally one, one other point I want to make is about this column called, that I've called Joins and Analytics. And so, one of the distinguishing features of NoSQL systems is that they, they typically don't support any notion of joins. And you know, they would, many of the proponents of these systems would argue that that's okay, that their joining aren't really necessary. So, the argument about why you can get away from joins, goes something like this. When you're joining two tables you sort of have one record and a bunch of corresponding records in another table. Well, if I put all those corresponding records and co-locate them with the parent record, right? Then I can access everything all at once and I don't actually need to compute the join, okay. So, that works and it should sound somewhat sort a familiar, because it was part of these network and hierarchical data models that we talked about in the discussion of how we motivated relational databases. And if you remember, the down side of this was that you had kind of one dominant access path into the data. But any other access path was either not supported at all or was inefficient, okay. And if you want to reorganize your data to support a different access path or because your, your needs changed well, a lot of your code might of broken and you would need to sort of rewrite it. Okay. And so this tension between buying yourself a little bit of performance by organizing the data in a, in a pretty rigid way. Versus the development time that you save by not having to rewrite your code whenever you want to reorganize the data. That equation sort of balance a certain way in the early 70s and I would claim that it still balances the same way now. Now, this is not to say that commercial databases, relational databases, as they stand today, aren't necessarily meeting everybody's needs. I think they're clearly not for a variety of reasons, but to sort of thorough out what we know. And, you know, give up on this flexibility that was earned from the relational data model, in favour of upfront decisions about one dominant way of organizing the data. I'm not sure that's the right one either. In order to, so that's one point I want to make. Another point is, that there is not necessarily one right way of decomposing things, or one right way of even evaluating the join even if you support them. Okay, and there's a point I made before but I want to bring it up again in the context of NoSQL. So, in a, in a pretty classical web application scenario, if you want to show all comments by a user named Sue associated with any blog post by a user named Jim. There's a couple different ways you could do this, you could look up all blog posts associated with Jim, and then fetch all corresponding comments filtering for Sue. Or you could go the other way, you could find all comments by Sue and then for each of those comments, look up all the blog posts by Jim. And either one of these ways may or may not be available to you if you have already organized your data in a particular way. But even if they are available to you, it's not clear which one's the right one. And there may even be a third method that sounds a little wild but it's filter all comments by Sue and then independently filter all posts for Jim. Sort them, by some sort of blog id, and then pull one from each list and walk over the data that way. So, that sounds a little wacky, but that's just a sort merge join that you may or may not be familiar with, and if you're not that's okay. But the point is all three of these are perfectly valid. And the right one depends on the details of the data that you use a programmer, as an application programmer may or may not have access to and that may change from time to time. Time. And so, really only the system is equipped to make this make this decision. And so being, you know, sort of succumbing to the tyranny of your, of the initial design decision at the time the database was built or designed is one problem. Another problem is that even if you have some flexibility, leaving it up to the programmer to make the choice over the right way to access the, access the data. is asking him to do, to make a decision for which they aren't equipped to, they aren't equipped to make. Okay, and this is something that does not happen with re, relational databases. Neither one of these problems exists is in the same way. Okay, so, I think trying to inject some of that smarts back into these NoSQL system is, is a good idea and we see that trend happening. Okay. So, that's maybe the takeaway that I want you to have. Okay. So, that's all I want to say about this grid, but let me give you a little bit of a character of a response to NoSQL by Mike Stonebraker, in communications with the ASM. And there's a couple of different blog posts here, that talk about two different arguments in favour of NoSQL. And they responded to each one of those arguments and so I'm mostly going to focus on the second one. So, there's two pro NoSQL arguments are these. So [UNKNOWN] points out that there's two value propositions offered by the NoSQL community. One is performance and story as Mike says, that I more or less agree with is that you know, these people started out with a MySQL deployment of some kind. And they had a hard time scaling it out in a distributed environment and then they had two choices. You know, either they could invest in a large scale relational database and pay of course, many license fees or they could do something different, use one of these NoSQL systems. Alright. And then the flexibility argument is like, well look, my data doesn't conform to a rigid schema, so in the performance argument, I'm not going to spend too much time on this. Because he talked about the Trade-offs associated with different choices in transactional guarantees in these various systems. And the other part of the argument that that Mike makes, has more to do with database internals than we'd initially covered, so let me focus on the flexibility argument. So, an observation that he makes that I think I probably agree with is that you know, who are the customers of these NoSQL systems? And so it's a lot of start ups, a lot of web start ups. And it's not quite such the same penetration in the enterprise. Okay? At least not yet. So why is this? Well, one argument is that most of the applications in the enterprise are traditional OLTP. And if you don't know, OLTP means Online Transaction Processing. So these are sort of you know, bank records alright things were doing the transactions right really matters okay. And those further results would be on structured organized data and so there is a few other applications around the edges, but they're perhaps considered less important. Right, it's okay to take a high risk system because the application itself is, of less interest to, you know, executives. Okay. And so really no, no asset compliance, you know, no transactions is equivalent to not much interest in the system. You know it's not okay to screw up mission critical data, but for that, you know, for some other application on the edge, maybe it's okay to experiment with a NoSQL system. Okay, and another, and a second point he makes is that you know, relying on these low level query interfaces is a, a real tough cell. And, you know, he calls up CODASYL, which is an early data manipulation language that pre-dated the declarative language that we've been talking about. But we've sort of been down that road, and it's tough, and it's why these high level languages were invented. And again, we see the, we see it with MapReduce, which is middle [INAUDIBLE] analytics, as opposed to, to NoSQL. But we see that the value of these high level interface pays off, okay. And the third point he makes is that, you know, NoSQL means, we, all bets are off, right. There's, there's no, there's no sort of homogeneity at all between all the different deployments. And so if, in a, in a typical enterprise you have, you know maybe 10,000 databases. You already have enough trouble trying to integrate data from these databases because of the heterogeneity with in their schemas. But at least now that your always working with rows and columns, and at least have kind of a standard interface to manipulating them, okay? And so having, you know, the number of design decisions that you have to make to encode your data in one of these new SQL systems. You know, what becomes the key and what becomes the value is, are the blog posts nested under the comments or are the comments nested under the blog posts. Right, do users keep their own wall or do the you know, the, the messages on their, on their front page or does the person who wrote the message keep access to it, or both? Right. All these different design decisions of nesting layers and so forth complicate integration and complicate standards. Okay. So, it's a, it's a, it's a tough sale. You know, the other, the other point that I guess I'd like to make that's related to this, this third one here is that, you know, there's no real free lunch. Either the complexity's going to be in the system that you use to model the data, or you're going to sort of hand it over to the application. But in some sense, they're always, these schemas, right, this application business model is going to be encoded somewhere. So, having it centralized in, in the data system, as opposed to hidden more than once in various applications that access the data system, seems like a good idea.