[MUSIC]. Okay so, Google BigTable again, had a lot of influence and one of the major outcomes is just like MapReduce it was taken up and implemented as an open source project called HBase. And so it, where it, where BigTable's compatible with MapReduce. HBase was compatible with Hadoop and in one slide there's not much difference here. But I just want to mention the terminology so you've seen it. That there's table at the top level and then region store and then Mem Store and Store File. And so the exact names of these structures are a little bit different, and they did insert one more layer of abstraction which is this region. Fine. And then, how this sort of is compatible with Map Reduce is that, you know, each, each one of your Map functions will process a single tablet. And so in, sort of, one to one with the blocks of data that we talked about when we talked about MapReduce. And then, I just sort of ask this question, sort of open ended, is that there's just no such speculative execution in MapReduce that we talked about whereby for fault tolerance reasons it might kick off the same map cast twice on two different replica of the data. And the reason for this is that if one fails, well you have the other one. Right? So you just don't start over from scratch. But, you know, in this environment where you're now working on data that is being actively updated. And that's what HBase and BigTable were designed to support is, is updates. It's not quite as clear to me what's going to happen when you have, you know, it's possible because of eventual consistency that these two tablets will not always agree instantly. On the same record they'll agree, but different records within the Tablet may not. Okay. And so, just an example of when you sort of mix these 2 systems and there are no, sort of, system wide transactional guarantees. Or system wide even properties, that you can run into, run into trouble. And I think that a general theme here, with these, sort of, no sequel systems, including the ones that are designed by Google, is that you kind of offloading some of those responsibility to. The application to the program that will sort of sort this out and make sure it's okay. Okay and we're going to come back to that in, in just a couple minutes. All right. So after BigTable, several years later there's a paper by a bunch of folks in Google about a system called megastore that I'm not going to spend a lot of time on. But it's basically you know, they found this point I have sort of just made is that these loose consistency models can complicate application programming. And what they want to do is provide a little more system support for certain kinds of safe updates. Okay, so here instead of full transactions being safe within a individual record as they are in big table. They've extended it with this notion of Entity Groups. And so an Entity Group, it should be on this slide. An Entity Group is a set of records that tend to go together, tend to be accessed together. Okay. so maybe again this is the blog and all of its comments for example maybe each one of these is a record is a 6 interval record store so it's okay for them to have different schemas but they all tend to go together okay. And so what they does extend transaction support over an entire entity group, you know, a set of, a set of related records. Okay, so they still get the scalibility by not requiring full system wide global, you know, synchrony. But they allow you to sort of, they, they get away from this problem of you know, I, very frequently I might need to update one record and then update all of its sort of children records at the same time. And I can't do that any kind of safe way. Okay. So fine. Fast forward one more year Alright, and so there's a 2012 paper on a system called spanner. And I just want to mention these quotes and then we'll talk a little bit about the, the system, and this one is still sort of being explored by the online community. It's not available actually for use, but the paper's being explored, and the idea's being explored, so for example, you don't see an open source, actually, that's not true. You do see, there has been a couple open source implementations of the ideas in Spanner, but they're not quite as popular, some of the open source Googledations of the other Google systems. Okay so you know it says that even though many projects happily use BigTable we have also consistently received complaints from users that BigTable can be difficult to use for certain kinds of applications. Those that have complex evolving schemas. Or those that want strong consistency in the presence of wide area replication, okay. And then we go on to say, we believe it's better to have application programmers deal with performance problems due to overuse of transactions as the bottlenecks arise, rather than always coding around the lack of transaction. And so this, you know, the database community could have said, well sure, [LAUGH] you know, well done. [LAUGH] Alright. That's exactly the point is that system supply and support for transactions is, is always a win. Right. Because it's a, it's difficult, error prone, expensive to try to do this at the application level. And, more importantly, it's fundamentally wrong, in some sense, to do it at the application level because it does, the application doesn't have global knowledge of what's going on. Alright. Only the system does. So, fine. So, although Spanner is scalable in the number of nodes the final quote here, the node-local data structures have relatively poor performance on complex SQL queries, because they were designed for simple key-value accesses. And then, algorithms and data structures from the database literature could improve single node performance a great deal. Again, you know, it's, it's, somewhat of a, of a Google-style approach to the problem of reboot everything, rebuild it all from scratch and then sort of cherry-pick and bring things in. So this has been working pretty well. And they have fantastic impact in the community. but there's a lot out there in the database literature and in the database system that could have been used from the start, in fact trying to start from the beginning and just say were going build a big Google style parallel database may have been a good choice. Rather than sort of getting completely away from it and then coming back incrementally and finding yourself in, in a SQL system. Now I sort of skipped over what Spanner is but it's, it's a planet scale database system, there is a SQL like language should just go back to our, I mean I know what I'm missing. I'm missing our, our, our table here. Let me flip back a few. So here it is down here. So I'm missing, I'm missing this slide here where I showed it. So really big scale. Primary accesses you can't access by other attributes. There are transactions in effect their global this time their real, real asset transactions. it's not clear to me whether joins are supported I suspect they are cause if you talk about sequel but I couldn't find an example of whether there is or not. there is a notion of schema and they do sort of protect against data that doesn't perform to the schema. There is some notion of logical data independence, although they, they don't talk about it much. There is a sequel like decorative language on top of it. I didn't see much evidence that they're doing a whole lot of fancy optimization. And I did just show you that quote of where they say that they performance is sort of poor on complex analytic queries. Bus as soon as the problem would come along somewhere quickly okay. So fine, so that's, the spanner a high level. Let me give you a couple more details about what this system does so the data model here is discussed in directories and these are a set of continuous keys, with a shared prefix. So you can think of it kind of like a tablet was in BigTable, but now they have this notion of multiple logical tables that are sort of interweaved. And so if you're not used staring at this syntax, don't worry too much, but those of you that, who are thinking in terms of DDL in, in a relational database, they have kind of a create table language that looks like this. You create a table, Users, with two columns, and then you give it this key word, directory. And then you create table albums with some columns, and you have this key word, interleave in parent users. And what you end up with is something like this, where, there's a user with all of its albums and a user with all of its albums. As you can see here that, you know, what we've been talking about, all these different systems are experimenting with ways of getting these nested data structures. Hierarchical data structures that look a lot like what we saw way back in the 60s, right? And they're motivation is the same as it was then. It's actually really really fast. When you're going to access, when you want to pull up a user, and then immediately pull up all of its album. It's really fast axis to this this way right? you know, but I probably speculate that the reasons why relational, the relational approach eventually replaced these. And, and what can and will happen here as well, is that performance is not the number one priority, it's minimizing the amount of developer headaches. Okay. So, it remains to be seen, but I, but I think that this, this incremental walk step towards a big new scalable relational database is, is, is underway. Now again, that doesn't mean that I'm saying, use all the old databases. They really were designed for, sort of, a different work load and they really don't. There is really no evidence on these scale on some of these, some of these levels, but that doesn't mean that you're sort of throw out all the, you know, everything we learn, okay. But that's more me editorializing. So fine, how this work is there's a universe master at the very, very top and this is just a singleton. There's only one of these for a deployment and they sort of imagine there is only one or two of these for deployments anywhere so they have sort of one for test, one for production/test, and one for production. And that's it. So all, so, so, many different applications will use the same deployment of, of spanner. And so, this is mostly just status about status information about the zones. It doesn't need a, it doesn't interact with clients at all. Then there's a placement driver that's responsible for moving these directory sets of records. Around for load balancing purposes, and this happened on a scale of, of every few minutes, alright. And then within a zone there's a, a zone master that assigns data to spanservers. And there's a location proxy that sort of knows where everything is. And routes requests to the appropriate spanserver. And the spanservers themselves serve data, and so in here it's starting to look a little more like BigTable. You know, a zone is essentially a, an individual BigTable deployment. Okay? So inside of a spanserver this is where the, the big difference here is this is where they're going to try to support fully consistent transactions. So, across these, you know within a group of these replicas. They can, they support 2 phase commit. This only is needed when a transaction actually accesses data that's, that's you know in that's across the, is not, you're not constrained in one particular replica. Okay. Other than that, it just skips over this, this logic, and it doesn't cost anything. Okay. And then one step down, below this, across so this is, sorry, I guess I'm using the wrong terminology. So it should basically come in as across groups, and when all of them, the transaction and all these contained in one single group then you drop down a level and you run the Paxos algorithm that I didn't talk about in detail but I mentioned exists. In order to sort out the reads and writes for, in order to handle the write. Okay? And the only other piece I'll mention here is that this term colossus is new. It's the successor to Google File System. And Google File System is the- original turn for the op, you know, the open sourcing limitation of HTFS which underlies Map Produce and Hadoop. Sorry. GFS is to Map Produce as HTFS is to Hadoop. So when I'm throwing these acronyms at you, that's how to keep it straight. Okay, so that's all I want to say about spanner in particular. Let's just take a step back and look at all the different systems that Google has for a second. You know, map reduce was a paper in 2004 that had a ton of impact, BigTable had a ton of impact, then there's Megastore, there's this tens thing that we didn't talk about, but it's a SQL system on top of mat produce, much like hive, if your familiar with that or if you've heard me mention it, and then spanner very recently. And so you can or sort of organize things into a timeline this way, just to kind of get a sense of this. And because of these systems have had so much influence I want you to be aware of what they are and sort of how they fit together so it doesn't just sound like a big jumble of terms. So MapReduce was, you know, the, one of the earliest ones. It wasn't quite the earliest. There was actually another one called Sawzall, that really didn't get a ton of traction, but it was a nice paper. and then BigTable came a couple years later, and I drew a dotted line there representing that there's sort of compatible design to go together. One was map produced for analytics, BigTable is for the, sort of micro operations. And then both of these, a few years later have an open source limitation in Hadoop and H-base respectively. Fast forward a few more years, and you got a mega store in spanner coming very quickly one right after the other. And this hard, this heavy blue line represents you know it's pretty clear that he influence is fairly direct. In fact, I would suspect that there's a lot of code being borrowed and then megastore makes plenty of references to BigTable and spanner makes references to both megastore and, and BigTable. And they, the papers have many, many of the same co-authors. Okay, and then MapReduce depends directly on, I'm sorry, excuse me, Tenzing depends directly on MapReduce. It provides a sequel layer on top of MapReduce. And then there's some other systems here. One is called Dremel which was originally for very fast aggregate queries but really just aggregate queries but of extremely low latency. So this is you know, in the analytics camp cause your doing these sort of aggregate questions as opposed to a sort of micro updates. but it was extremely low latentcy unlike map produce it was more of a batch system and so this is this the a great fit and its a very nice system. And in fact, since they then can do joins not just aggregates and more importantly this was exposed as a service that you can just use directly over the web, even in your browser, called Big Query. And that's a, that's a, that's an important word to watch. It's one of the few systems that is do, available as a service through, that let you, as a cloud service. but scales a very, very large data, and sports analytics, okay. And, then, another one that we'll, we'll talk about yeah. But we'll come back to is Pregel. And this adds the one secret ingredient that I, is sort of near and dear to my heart which is iteration. And what I mean by that is when you, when you run MapReduce jobs and do analytics you're sort of taking step one. And then step two. And then step three. And you stop. But for many kinds of tasks, especially in data science, many of these analytics tasks, these machine learning tasks. You have to do something again and again and again and again until some kind of convergence condition is reached. And Pregel and one of our systems and a few other systems, are the ones that tried to, you know, notices this limitation of map reduce and extend it. So people were doing this with map reduce, but they would do it sort of in fairly ad hoc ways. Okay. And so we'll come back to that and talk about it but, you know, analytics, low latency micro-updates. So those two big classes of systems. And then analytics with iteration is perhaps a third class of system that we, that we'll talk about.