Architecture fundamentals here, in the JPL-Caltech Virtual Summer School. Big Data Analytics. Hi, I'm Chris Mattmann. So in the first part of big data architecture we covered, covered two main things. We covered sort of an introduction to software architecture. Its principles, some of its research, some families of understanding of what we should be thinking about in terms of self architecture. And we sort of wrapped up thinking about things like styles, patterns, and reference architectures. Today or in this next sort of lecture here, we going to be covering architectural modeling and visualization. These are higher level, sort of architectural operations, that operate based on, capturing information about components and connectors and configurations, and then figuring out how to visualize them, and interact with them. And we'll talk about, another, thing that you see quite often and, which is another real good applicability of software architecture. When you have code and you have software architecture just making sure or understanding how consistent they are between one another, and identifying when they aren't and what those impacts are and what types of things you can do to sort of mitigate that. So, let's talk a little bit about architectural modeling basically as I stated in, in the sort of first part of this series on big data architecture. The Medvidovic, Taylor and Dashofy definition of software architecture as the principle design decisions, in a software system. So, architectural modeling, if you think about it within that context, is really the capture, right, or the recording, of those, principle design decisions. You may build a UML model. And what are you doing, when you capture those components or you capture the interactions between them? You're capturing the principle design decisions. You're creating an architectural model, of a Big Data system, okay? So, so architectural models defined, it captures and records the principle design decisions about a software system, right? So, there are several notations that have been developed to capture architectural models and instances of architectural models. Probably the most popular that you guys are all familiar with is UML. If you take a undergraduate course in software engineering, you've been exposed to UML. A lot of times if you take graduate courses you're exposed to UML. And there are deri, derivatives of it that exist nowadays, SCML, which is the systems engineering modeling language, a number of other modeling languages end in ML for modeling language as so forth. UML was the unified one, Sort of, was pioneered by, Grady Booch and the number of other folks who were thinking about how TO model software back in the day. CML is by far kind of the most prevalent architectural modeling notation you might be familiar with. It's also by far not the best and it has it's issues and it's problems and we'll, potentially talk about some of those, depending on how much time we have. But it's worth pointing out that UML was also emergent out of sort of a heyday of architectural modeling research back in the day, especially during the 90s and the 2000s. This is sort of the golden age of software architectural research. Right, this is the days when Roy Fielding was at East Irvine, and Jim Whitehead were there and a number of people were, you know, in a high powered software engineering, that this was when there was tons of research by David Garland and through Mary Shots, see, to me, this is when lots of things were going on at CU Boulder with Alexander Wolf. And Andre Van der Vulk and lots of, sort of work going on in architectural modelling, back in the day. So out of that came a series of different architectural description language, which are really the predecessor to common architectural modeling tools. ADLs, or Architecture Description Languages, okay? And we need ways of understanding even though many of these aren't still in existence to today, elements and principles from them have made their way into things like UML. And they've made their way into other modeling languages and notations. And it's important for you to think about ways of evaluating architectural modeling notations. So if you go out and think that you need to develop your own, or even if you're trying to pick, you know, amongst UML versus simply doing PowerPoint or, you know, recording an architectural model in text, you know. All of these are viable means for capturing those principal design decisions. You need ways of evaluating why, when, how, where, why you should, you know, do it or not. And the ways that we typically evaluate them are the ability or the capability of an architectural modelling notation to capture things like, the structure, the components and the connectors. The behavior of the architecture and to model that, like it's interactions and for example, what states components transfer between what type of data they have, what data is shared between those states, the workflow and things like that. So, this is how we evaluate architectural modelling and it's effectively why we need architectural modelling notations, and architecture description languages. So, these ADLs which again kind of came out of the 90's and, and the 2000's, and have since sort of emerged into architectural modeling notations. Some of them back in the day in terms of software architectural research you can kind of break them down between what they were focusing on. There are things like Darwin which was really focused on modeling components and their structures. And when you modeled all the components and their structures, you did it in such a way that you could use a pie calculus to analyze things simplistically like reachability of one component to another. Is this, are these two components going to be able to communicate? It's really nice to have that in something like Darwin, because you didn't even need to build your software system, per say, or implement it, but from an architectural model, you could determine whether or not two components had reach ability and could communicate with one another, simply by capturing one of these architectural models. Ra-peed took that sort of to the next level and really started to focus on interaction. Having a capability, capture not just structure in terms of components and connectors, but then, interaction and having an understanding of the effects of interaction, which led, we'll talk about visualization later, to things that effect visualizations. Wright focused on components, connectors, structure and behavior 'kay? Then we had some architecture description languages that were specific to a particular domain, like Koala which was used very heavily to model product lines. Okay, so large numbers of product lines for televisions, or for television software systems and controllers that were modeled by things like Koala. Okay, avionics, in the avionics domain, you had things come out like, the avionics architecture description, like, okay, or the AADL for them. All the way to the sort of the kind of common modern predecessor in, in even today, still now used extensible architectural description languages, or and there are two sort of three kind of main examples of that. ACME, which came out of David Garlan and Mary Shaw's group. Okay, and it supported components, component types, connectors, connector types. An arbitrary properties about those components and connector types. Because what they quickly realized with early architecture description languages, is you were limited to the information that you were capturing in them. Them, If you only knew about components, but you didn't know about connectors, you couldn't do downstream analyses of your architectural models that had to do with interaction. You could only do limited analyses that had to do with structure, or computation. Okay? So, what people, and then in avionics, they were modeling a bunch of avionics properties. In terms of ADL and other things. Well that might not have done as great of job as modeling properties relevant to say data movement in those Avionic systems, but it might have done a great job at modeling properties, properties related to architectural reliability. Because that was really important in the domain. So what people started to realize looking across all the architecture description language is, they are being developed in sort of this hay day of architecture research. Is that they really wanted to kind of pick the best of all of them. And so then came this sort of notion of extensible, architecture description languages, okay. And ACME was sort of the first example of that. ADML was ACME, that canonical model, represented as XML. Okay, and there was an ACME studio that came along with that, which was this sort of eclipse looking studio on which you can author and create these architectural model instances that were ADML instances, that had these ability to capture, for example in a component, its memory footprint. And when have to maybe specify its bandwidth requirements. To specify on a connector what type of interaction protocol it has. So, arbitrary properties specified per component, per connector, per canonical architectural element, and and really recorded in this architectural modeling instance. xADL or xADL is, very similar to AC, to ADML. And it came out of UC Irvine, XADL did, and the screen shot here is actually a screen shot of Arch Studio 4, which is, the environment in which you author architectural modeling instances that are guided by xADL or, xADL. 'Kay? So now we're moving from architectural models and its be, it being important to sort of capture these properties so that, eventually we can do downstream analyses. And I don't cover architectural analysis here in this in this summer school 'because we simply don't have time. But I encourage you to take a look at some of my architectural analysis slides on my website my USC website that I linked earlier. The whole goal of capturing these architectural modeling instances is really ultimately to have, the ability to understand a lot of properties and elements in our software system long before we implement the code. Well one thing that can really help us with sort of just interacting and looking at architectural modeling. Are architectural model notations, driven by these architectural description languages and these architectural modeling languages. One of the things that can really help us is architecture visualization. Okay? And visualization is more than simply putting a pretty picture of an architecture up on a screen, like for example the one that I showed you with Tika earlier. But it's, it's not just depicting the architectural model, but it's interacting with it okay, so this is what separates visualizations from, architectural visualization from simple static diagrams or simple drawings. Okay? And there are sort of four kind of common approaches to architectural visualization. There's textual, yes, some people capture architectural moral instances and then visualize them in text files, in the text editor. Right? We write a bunch of text, you know, you publish a paper on an architecture and big data, or whatever, and we've got text, which has a lot of information about the components, the connectors, those principal design decisions about the software system, so text is a completely valid way to visualize, an architectural model. And you know what, there's a textual visualization even for things like UML. Right? It's called XMI, [LAUGH] it's called interacting, the X, the XML MetaData interchange language. It's called interacting with the EML model, [INAUDIBLE] in largely a textual representation. Okay? There's graphical, yes people make PowerPoints. And there's architectural model, principle design decisions in there. Okay? Hybrid which is sort of what happens when you meld you bring together text and sort of these informal graphical notations. And then there are things called effect visualizations, and you saw these, you see these when you build early on, when you build models and things like Rapide, but what they are is. Basically building an architectural model, and then running it through a simulator. And we see these in big data systems sometimes. We take an architectural model that represents a detailed specification of the system, maybe some interactions, and then we simulate different data or operational profiles into the system, to then see what types of maybe the degradation of memory over time or things like that. So, sort of effect visualizations are visualizations, what if scenarios visualizations on your architectural model that you capture for that, a lot of times when you are thinking about visualization there are multiple views of the same architectural model or server architectural even, these are two views of the Tika library right. You'll notice a different number, cardinality of components and connectors from the left diagram. From the one of the right okay and we typically call the differences even though there are multiple views and architectural visualization. Two views of the same thing or the fact that architectural models typically have potentially multiple visualizations, and potentially multiple visualizations represent different architectural views. A component view, a view not being architectural components to the physical hardware, deployment view, for example, and things like that. So, again, just to sort of review, architectural model captures, some or all of the principle design decisions about the software are conjecture, right? So, you might have an architectural model that simply focused on the behavior of your software system. You might have an architectural model that simply focused on the structure, and so it may subset or capture just the subset of principal design decisions that have to do with that. Okay, an architectural view is a further filter. Okay? A further filter of that model that you want to visualize. That is, depict and interact with. Okay? It's some subset or collection or filter, further filter on the model. So, maybe we've got the entire architectural model and we want to show the deployment view so it's some filter of the principal design decisions that have to do with that. Now the deployment view is actually the deployment viewpoint, because the viewpoint is the name of that filter, the name of that sort of subset of the principal design decisions that we're talking about. So, the way I like to think about is we've got an architectural model, it may or may not be a subset of all of the principal design decisions of the software system. A view is a further subset of that model, and the viewpoint is the name of that view, okay. So we're going to move just real quickly into thinking about, some, higher level processes for this in terms of, architectural recovery. Now that we've covered architectural modeling, architectural modeling notations, architectural visualizations. We're going to shift gears a little bit, and talk about architectural recovery. What you find is that no matter what you do, to record architectural models, and do your best to think about the principle design decisions, to visualize them, before you implement the code. Is that when someone goes and implements the code, that code may be different then ones you originally theorized or, or prescribed if you will. We called this process architect, architectural drift. Well we call it two different things, we call it either drift or erosion. Drift is basically when the code, the as implemented architecture, varies from the prescribed, or the, you know sort of thought of, or software design or based on architecture modeling notations or, or whatever. If that code varies from that, but that it varies in such a way that it doesn't violate any of the core assumptions of the architecture. We call that drift, and it's subtle difference between that an erosion. Erosion is pretty much the same, that same case except the variations in the code from the actual prescribed architectural model, cause such a difference that it actually does violate like a core requirement or it violates sort of a core fundamental principal in the software system. Either way, there are processes for dealing with architectural drift and erosion, and that process is called architectural recovery. Basically architectural recovery is the the act of recovering the architectural or architectural model, and instance of an architectural model from the as implemented architecture or from the code. This is the sort of best way to insure we actually know what the true architecture of the system is, because the code is the most living canonical representation of your software system period. Right it's what you are constantly are working on. You may even have great architectural design and all of these documents, but eventually you get to the code and that's really what the invest is going to go into in terms of maintaining. So the best way is to have some process for automatically sort of, or as automatically as possible, recovering an architectural model from that. And that's the process of architecture recovery, there are various processes for doing this that have been developed in the architecture research literature over the years. Including ones by Kazman et al, Jakobac and Medvidovic et al and various other researchers and things like that. Typically this involves looking at the code, doing static analysis on it, determining what pieces of code talk to one another, looking at dynamic statement analysis, what states are these classes, or are these components in. Component connector analysis which are groupings typically, of kind of common classes and file and object oriented code are functions and procedures, and procedural oriented code. And then eventually potentially mapping it to an architectural style and things like this. So, we going to wrap up in this module with architectural recovery and so, in the next module here in terms of big data architectural fundamentals I'm going to actually illustrate in the context of big data. Architecture recovery, some of these core architectural elements, components and connectors and discuss it in the, in the context of a real world example. So, thanks.