Module, in terms of the Big Data Architecture: Fundamental lecture here at the JPL-Caltech Virtual Summer School on Big Data Analytics. I'm Chris Mattmann. Thanks for sticking with me here to the end. So, we've done a sort of a hitchhiker's guide really fast paced history of the sort of research and software architecture, software engineering, how some of that might apply to big data. We've covered things like components and connectors, core architectural elements, styles, patterns, reference architecture. We've covered ways of capturing these sort of core elements in terms of architectural models, why you may want to visualize those. We've talked about the difference between code, as implemented and what can happen when code drifts or erodes from the actual architecture and design and what we can do to kind of deal with that, architecture recovery. And so here, in the final module, going to do an actual case study in the realm of grid, grid computing. To show you some of the value of software architecture and, and how you can actually apply it Within the context of big data. And your guy's applications and then maybe wrap up with some conclusions. So this, this lecture actually covers a specific example that's going to bring together soft architecture connection and Big data. So, you guys may be familiar or you might not be with the domain of Greek computing before clouds and big data and things like that, there was Greek computing. And the goal, as sort of defined by the fathers and, you know, forefathers and mothers of grid computing. Was to provide an infrastructure and an architecture for, for basically seamlessly bringing together resources for data and computer across organizations. To create virtual organizations that was the goal of grid computing systems circa, you know, around 2001, 2002 which was sort of the. Initial hot heyday that cloud computing era if you will of grid computing. Okay. So we were studying those in my research group around the time. We were studying as we found out there were two fundamental types of grid computing systems corresponding to the two fundamental types of resources that existed at the day. There were data grids which were really focused on federating data resources across institutions, searching them in sort of a nice and easy way. And then, there were computational grids, which were focused on how do we run jobs at your institution and my institution, and share sort of identities. And only let these science groups do these jobs, and let these other ones use these resources, and so on and so forth. So there are two predominant types of grid computing systems, data grids and computational grids from there. So, we were actually writing a paper on trying to understand the architecture of grid computing, trying to moosh them together, because we said, look, there were these sort of canonical papers on grid computing. There was the grid's anatomy, the anatomy of the grid, which was written by Foster and Kesselman and Alf. And then there's this paper called the Physiology of the Grid. The anatomy, there's more, the architecture, the physiology was more how do we implement this thing. Also its written by Foster, and Kesselman and Steven Tuck and a number of people. And so we were writing a paper on our grid technology at the time that we felt, kind of consistently met this sort of requirement of sharing data. We're sharing compute resources and things like that and we were writing in the context of mobile computing. We're building this project called Glide which was a technology for, it's like a mobile grid for sharing data and compute on wireless PDAs of the time before cell phones were super poplar. We submitted this paper to workshop. We got back a review. The workshop basically said the grid technology you're studying is nothing more than a simple object oriented framework. And so, we scratched our head and we said you know was the reviewer right about this? I mean, a simple object oriented framework it really seems to fit at least the definition of competing in terms of these anatomy and the physiology paper. You know, and the kind of current literature at the time. So we're wondering is, you're right. How do they know that? Okay. And we were also kind of annoyed, you know, for the review that we got. This Alleged Object Oriented Framework was the 2003 Runner-up NASA Software of the Year, okay? So, we basically started doing a little bit more research and looking around at the time which you guys might do here. You know, maybe as PhD students or post docs or as petitioners in the realm of big data. And we started to look at the research literature in terms of architectures and grid computing, and we basically found out that was, and still kind of persists to this day. That little was sort of known in terms of the architecture of the as implemented grid computing technologies of the day. There was a study by Finkelstein. And all of that was published in the Journal of Great Computing in which they were trying to look at things like requirements, and so forth. But, they didn't really know much about the architecture of it, the components, the connectors, the stuff that we're talking about here, in the summer school. Okay? Little was known about that. And so, you know, there's this big risk if you, you know, know a little bit about the requirements, but you're not exactly sure of how that maps. You know, and how that went through the design process, and how that went through the architectural process to the eventual implementation. It's a big risk of architectural drift or erosion, as we were just talking about, you know, in the second module of this course. Right? It seems like grid technologies all generally claim to have the same capabilities, but. They seem to be implemented in vastly different ways, have different sort of requirements that their satisfying and so on and so forth. Maybe evidenced by the reviewer that was reviewing our paper at the time. So we develop an approach that's base on what's your learning here in terms of software architecture. And related to big data and so forth, because grid computing errors, arguably within that realm. We developed an approach to studying, initially five, and eventually, almost 20, grid computing technologies and their implementations. And we developed an architectural approach to studying, these technologies, and I'll, and I'll talk about. So we had to select software codes. We decided to select, great computing technologies that were object oriented simply because some of the techniques we use to recover the the architecture were real amenable to that. We wanted them to be open source because we want to be able to look at the source code, the source code being the most, like I said, up to date version of the software and most sort of canonical representation of your system sort of as it evolves. And, we also wanted to use, kind of off the shelf, you know, systems, that seem to have a big user base, So, these were the initial five that we studied. We studied Globus, which was the de facto grid technology at the time, OODT, DSpace, GLIDE, our own technology because we wanted to put it to the test for building, and something called JCGrid. And that was where the initial five technologies that we studied. In 2005, 2006, having since expanded the study to include over, like I said, 20 grid technologies circa 2009, 2010, including Hadoob, things like Wings, Pegasus and so on and so forth. So more modern technologies. So our approach was to initially do sort of a detailed literature review, and to try and look at the four seminal grid papers at the time. Just this data grid paper, by by Ann Chervenak, as well as three papers from Foster and Kesselman on the grid's anatomy, its physiology, and a checklist of what it means to be a grid, software technology. So looking at these allowed, these literatures allowed us to distill a set of reference requirements. Where have you guys heard that before, right. We talked about reference architectures and domain-specific software architectures. We came up with a set of reference architectures for grid, reference requirements for grid, the domain of grid computing. And so for each acquirement, requirement we derived from studying the literature we also defined a layer in the classic five-layer grid architecture, which I'll talk about here in the next slide. We mapped the requirement to which layer it actually had to do with. So if the requirement had to do with sharing resources across multiple organizations, it likely went in the collective layer, which is responsible for that. If it had to do with application or building sort of applications on top of the grid computing infrastructure. We mapped that requirement to the application layer, and so on and so forth. And then what we did is, for each of the technologies, initially five, eventually 20, like I said, we applied one of the sort of more automated clustering static analysis dynamic behavior architecture recovery techniques. We actually applied one called Focus. Just developed by Malevich and, and, and Yachabach, to these, these open source code bases to arrive at sort of a partial architecture recovery model. And we use these reference requirements as well as the layers that they map to, to help us guide and shoehorn which one, of the layers, the components that we recovered, then the connectors and this partial recovered model actually went into, okay. And then we studied that. So the approach is sort of depicted graphically here. There's a five layer grid architecture. Real quickly, the layers correspond to fabric. These are, like, the data disks, the storage devices, the actual physical processors and things like that across a grid. A connectivity is all the networking layer and things like that that make these available via network devices that make these resources available in a network way. Resources are specific individual resources services in a grid environment specific to a single institution. Collective is a layer on top. And by the way, you're seeing this in the upper left hand corner of this diagram. Which is sort of the five layered great architecture as defined in these papers. Collective is multiple resource, sort of services across multiple potential institutions. It's what allows sort of this virtual organization to deal with resources across multiple institutions, find them, find compute, find storage and so on and so forth. And then, that application, or applications, that are grid enabled that are built on top of this sort of underlying infrastructure and then interact with them. So, so that architecture was found in the physiology and the anatomy paper. We studied that. We pulled that right out of there. That's in the upper left. The, reference requirements were distilled by studying that, those papers and the data grid paper amongst others, these sort of four seminal, papers. And then for each grid technology source code, there on the right, we took the source code, any documentation, and anything else we found and we ran it through this sort of focus process to produce this initial sort of recovered architectual model. We sort of compared those reference requirements to see if it could give us any information about the components about the connectors. We shoehorned them into the upper left there, into the actual sort of perscribed. Great architecture to see how well the implementations of these grid technologies conform to what they said the grid architecture should be. And we use this as sort of a metric to determine, like, okay, how well does our technology do? Because, remember the original empadis for this was, someone told us our technology wasn't a grid technology. It wasn't a grid software system. So we wanted to develop sort of a tried and true research approach, a scientific approach for verifying this. Okay, so, we we basically, I basically talked, sort of, through this, visually, on the last diagram. Again, we clustered our components according to sort of the reference architecture and shoehorned them in there from that recovered architectural model, which we sort of gleaned according to that focused approach that we had. And then we used the reference requirements to sort of give us further information about what layer in the grid architecture or the components or the connectors that sort of go into. And so these are the reference requirements, the list of around 17, that we gleaned from that basic grid literature. And, the sort of impacted layer in that grid architecture, that sort of gave us some more insight as to well, you know, if it's something that's dealing with single sign-on it likely goes in the connectivity layer, and so on and so forth. So, this is a sort of sequence of steps that we took. We did these sequence of steps for all of the grid technologies that we were looking at. This happens to be, in the upper left, a static class diagram of the object Apache OODT software, object oriented data technology. And Starting from the upper left and following a, b, c, going down that column on the left. And up again to d on the right. Down all the way to f. There. Down at the bottom. That's sort of the steps of focus. Focus is a clustering technique. It's a graph sort of clustering technique. That sort of looks at coupling and cohesion, the relationships between these sort of classes, Remember they're object oriented codes that we're talking about classes and relationships between classes like generalization, inheritance, and things like that. So, focus is a graph clustering technique that allows us to sort of measure the coupling and cohesion. Between these to try and derive what classes are related and to suggest they might be a component. What classes are interacting with one another to suggest they might be a connecter, and so on and so forth. So we ran sort of ODT and all the other technologies up to 20, like I said, through this process. Ran them through Mr. Wizard here. And what we got out the other end is this sort of recovered architectural model, when none of the clustering techniques for focus, when it sort of synthesizes, when it sort of arrives at a point where they can't really be clustered anymore. That's basically this recovered architectural model. It's a partial recovered architectural model. It doesn't have all of the detail, the styles, the configurations,. And all of that, but it has some basic information about components and connectors and things that we can work with and which we did work with, according to the process that we'd talked about before. I took, we took this recovered architectural model, and then we shoehorned it into the grid reference architecture, those five layers. Okay, again, with those reference requirements that gave us some insight, which we distilled from the literature, to figure out which layer it went into. As well as, just with looking at things like documentation and any other information that we can find. And what you find is that, you know, does this sort of look perfect? well. Let's see, the canonical layered architectural style has the following constraints. First, any layer in the, layered architectural style should only be communicating with it's most, adjacent most layer. Communication should flow top to bottom, looking at this in a vertical way. The top-most components being sort of client consuming components and, or layers, if you will. And the bottom most layers, or components being service providing layers, to all of the above, sort of layers and components above it. Okay, so things on the bottom should be service providers to things on the top and layers on the top. Right? And again, any two layers, communication should only happen between any two layers, not from example the top layer all the way down to the bottom layer. Okay, so do you see anything wrong with this picture? It might be hard to see unless you sort of look at this, so let me help you here. There all sorts of things that are wrong. All sorts of violations. You see, components here in the application layer crossing a two layer boundary to get down to the connectivity layer, okay? You see upcalls, you see components in the actual fabric layer making an upcall to a component in the resource layer, okay? You see things like components that we couldn't determine the right layer that it went into, all right? We just didn't have enough information, or we couldn't tell. What about this one? This is Globus's recovered architectural model. Okay. Woah two layer boundary and up call. Three of them right there. What else? Five components that we couldn't determine which went in the right layer. And Globus is a comparitively larger system you know than but it still exhibits the same if not more sort of basic characteristics for that. And we found this in a number of the software systems that we're looking at. Up call, call, up call. Here. So, so basically, there are numerous violations of the reference architecture. There are things like component upcalls, which indicates sort of, definitely a violation of the layer architectural style, that potentially indicates architectural drift or erosion, components or code in the actual code that originally weren't designed to talk to one another but that are talking to one another there. Crossing two or more layer boundaries. This typically indicates developer sloppiness. You know, you had a library that you intended for other, sort of elements to sort of call. Or the most adjacent elements to that to have some sort of coupling to it. But you find later on down the road, shoehorning and bolting on components. That you just decided to just make call this library even though there no where related. In the software system at all or nowhere related in the actual architecture as well. All right? And then there was just, you know well it could be a refactorization problem. It could be a number of things. But And then there are some components that you just, at the end of the day you look at your coding. Why is this here? It's not being called by anything. You know, nothing. It has, it doesn't really have a purpose. So you know, look at all the types of information that we could glean from performing this type of process. Okay, what we'd glean from that is that you know, grid technologies tend to be sort of a domain specific software architecture for the realm of great computing. Right? They exhibit a core set of reference requirements. They have sort of, kind of similar components in style even though they're violations. They have a large variation in terms in their size and number of components. And cardinality, and things like that. They also have optional requirements, right. And these were identified through, you know, not all grids satisfy these sort of collective level requirements. Not all of them satisfy even these layer specific requirements for that, okay. So, you know, it seems like single sign on is optional. It seems like, you know, data grids are optional, okay. And also a distinction between data grids and computational grids, right? In, in terms of the identified requirements for that, okay? So, these were sort of the results of our study, just, just taking a sort of an analysis of you know, the actual code and its mapping to software architecture. And these actually led us to develop and publish. Basically we're working with the Journal of Good Computing to publish a new architecture for great computings that actually more carefully and accurately represents the architecture based on all of these purported grid technologies, these grid codes. And so on and so forth. And this was simply by looking at the architectural styles the components and things like that as derived from the code. And needing things like architectural recovery and processes like we've talked about here. So applying and figuring out how to think things through and how to sort of discern the differences between these large scale software systems and you have some big data too. Identifying which requirements are optional is really important. They may tell you when you build your next system, you don't need to include components or software that supports that. And save you money, time, resources and a number of other things. Okay, so just thinking about and having this sort of basic understanding of software, architecture, components, connectors, the basic elements, styles and patterns, how they apply, what they suggest, what. You know what should be true, what shouldn't be true about them has a real big impact on software systems. 'Kay? And your big data systems. There's a lot of related work. I'm not going to go over this. I'll leave it here for the read, for you guys to sort of peruse. You know, the idea is grid systems, software systems typically have a hard time. Following the architectural style and that's just because there's a lot of overlap between grid layers and software systems in general are hard to understand. Unless from an architectural perspective especially when they get large you're dealing with data systems, frameworks, middle wares, libraries. The only way to truly properly understand them is to approach them from the realm of software architecture. Right, which is a summarization if you will, of the principle design decisions about your software system. There's lots of work kind of going on and, and, and areas for people to expand this. I encourage you guys to take a look. I encourage you to take a look at some of these pointers, again to my software architecture class at USC, to my homepage. And then all of data behind this particular study of grid middlewares is available on the following two links and websites, including. The data for the as submitted paper to Journal of Grid Computing. So I encourage you to take a look at that too. Encourage you to contact me if you have any questions, I'm Chris Mattmann. And thank you, this ends the first full module here on, on big data architecture fundamentals. Thanks.