[MUSIC]. Okay, so I want to spend a little bit more time on the details of MapReduce versus our relation database. beyond just the sort of how the pair, query processing happens. We saw that parallel query processing is largely the same. Some of the, many of the algorithms are sort of shared between them. There's a ton of details here that I'm not going to have time to go over. but the takeaway is that, that the basic strategy for, for performing parallel processing is the same between them. But there's other features that relational databases have and I've listed some of them here. So we mentioned declarative query languages, and we mentioned that those start to show up in Pig and especially HIVE. Now there's a notion of a schema in a relational database that we didn't talk too much about. But this is you know, structure on your data that is enforced at the time of the data being presented to the system. So when your data does, does not conform to the schema, can be rejected automatically by the database. This is a pretty good idea because it helps keep your data clean. It's not real feasible in many contexts, because you know the data's fundamentally dirty and so saying that you have to clean it up before you're allowed to process it, just isn't going to fly. Right. And so this is one of the reasons why MapReduce is attractive, is that it doesn't require that you enforce a schema before you were allowed to work with the data. However, you know, it doesn't mean that schemas are a bad idea when they're, when they're available. And in fact, really, you know, even with MapReduce, a schema is really there, it's just that it's hidden inside the application. Right? So when you read a record, you're assuming that the first record, the first element in the record is going to be an integer and the second record is going to be a date, and the third record is going to be a string. So that schema is really present, it's just present in your code as opposed to pushed down into the system itself. And there's a lot of great empirical evidence over the years that suggest it's better to push it down into the data itself, when where possible. And in fact you're starting to see this. So HIVE and Pig again have some notion of schema, as does DryadLINQ, as does some, some emerging system. there's, there's a system called Hadapt, that won't talk about really at all. But combines sort of Hadoop level query processing for parallelism and on the individual nodes there's a relational database operating. And one of the reasons among many is to have access to schema constraints. fine, logical data independence. This actually, you don't see quite so much, so this is the notion of views, right? Does the system support views or not and you haven't seen quite as many instances of Hadoop systems that support views but I predict they'll be coming. Indexing is another one. So, we talked about how to make things scalable that one way to do it is, to. Who've derived these indexes to support sort of logarithmic time access to data. that's not available in vanilla MapReduce. Every time you write a MapReduce job, you're going to touch every single record on the input. You're not going to be able to zoom right in to a particular record of interest. That's wasteful and it was recognized to be wasteful and so one of the solutions. You, you see people adding indexing features to Hadoop. And a H base is an open source implementation of a, another proposal by Google for a system called Big Table. That among other things provides, kind of quick access to individual records. And H based is designed to be sort of compatible with Hadoop. And so now you can design your system get the best of both worlds. Okay. See, you can't get some indexing along with your MapReduce style [UNKNOWN] in your face. And once again, I'll mention Hadapt here as well. One of the motivations for Hadapt to be able to provide index young individual nodes. Okay. Fine, so I'll skip caching/materialized views. This is the same same as logical data [INAUDIBLE] accept you can actually pre-generate the views as opposed to evaluate them all at run time. But we're not going to [INAUDIBLE] too much about that. And then transactions which I'll talk about in a couple of segments in the context of, of no sequel. But while, databases are, so databases are very good at transaction. They were thrown out the window among, among other things in this kind of context of MapReduce and their sequel. And they're starting to come back. Okay. But remember you know, what MapReduce did provide was very, very high scalability so this is you know, thousands and up, thousands machines and up. And it also provided this notion of fault tolerance. So relational databases didn't unders, didn't really treat fault tolerance this way. They were unbelievably good at, you know, recovery, right? If you were, because of this notion of transactions, if you were sort of operating on the database and everything went kaput. given some time, it would figure everything out and recover. and you will, you, you can be guaranteed to have lost no data, okay. That's fine but that's not the same thing as saying during query processing while a single query is running what if something goes wrong. Do I always have to start back over from the beginning or not? And the sort of the implict assumpti-, assumption with relational databases was that your queries aren't taking long enough for that to really matter. But in the era of big data, of massive data analytics of course you have queries that are running from many many hours. Right? And so having to restart this, in the course they're running on many many machines where failures are bound to happen. And so that context is something that MapReduce sort of really motivated and now you're, you see modern parallel databases. Capturing some of [UNKNOWN] tolerance in general. Okay. So this is sort of a list of some kind of partialistic contribution for relational databases. And this is a partialistic contribution, maybe a completeness of contributions from MapReduce. And my point is that you see a lot of mixing and matching going on but the design space is being more fully explored. It used to be sort of all about relational databases with their choi-,, their choice in the giant space, and then MapReduce kind of rebooted that a little bit. And now you see kind of a more fluid mix, people sort of cherry picking features. Okay fine. And then the last one I guess I didn't talk about here is, what I think was really, really powerful about MapReduce is it turned, you know, it turned, it turned the. Army of JAVA programmers that are out there into distributive systems program. A mere mortal JAVA programmer could all of a sudden be productive processing hundreds of terabytes without necessarily having to learn anything about the distributive systems. That was really, really powerful. Right for, with the, the analog in databases was, I mean, you had to become a database expert to be able to use these things. Okay. And so, I think that that impact is hard to overstate. Right, the ability of one person to get work done that used to be, you hire a massive team and six months of work was significant, alright.