Hey everybody, welcome to the JPL-Caltec Virtual Summer School on Big Data Analytics, here in September, from the 2nd to the 12th in 2014. I am Chris Mattmann, I'm going to be, lecturing, about, a really interesting topic which is understanding soft architecture within the context of big data or, or big data architecture and it's fundamentals. So, the next set of slides is broken up into a series of, three talks, so I'll, stop sort of at each, at the end of each, set of the talks, giving you time to sort of review the materials and what's going on. So let's, let's go ahead and pop into it. We're going to cover a variety of topics in this lecture, we'll start out just giving you a basic introduction to the field of software architecture and software engineering research. As it pertains to big data. I'll talk to you guys about architectural styles, the difference between styles. I'll talk to you about patterns, and reference architectures are a little bit of an advanced topic. And then we'll get into some other areas within software architectural in terms of architectural modeling. visualization, how do we represent, things like components and connectors as models, how do we visualize them. We'll get into some other processes in terms of drift and recovery. What happens when your code actually drifts away from your, sort of intended architecture and how do you reconcile that. And then the latter part, the last sort of third of the lecture, is going to cover a case study on understanding software architecture within the context of great computing. So, how can you actually apply this. So this first lecture here is going to go ahead and cover an introduction to software architecture and then styles, patterns in reference architectures. So who's talking to you right now? Let me give you a little bit of background sort of on myself. My name is Chris Mattman, I sort of wear three hats, I'm the Chief Architect of the Instrument and Science Data Systems section at the jet propulsion laboratory at California Institute of Technology. I have a team of about 20 data scientist at JPL. And what they're doing is working across various projects from DARPA and defense investments, building out open source solutions for Earth science remote sensing missions. Uh,we work within the context of astronomy projects, specifically the square kilometer array and its 700 terabytes of data per second that its going to generate when it's online. So a wide variety of things, and those data scientists, the software that they're building, has sort of a dual role. Every piece of software that my team writes every line of code at JPL that my team develops is contributed to the Apache Software Foundation. Apache is a 501c3 non-profit organization. It's also home to the world's most deployed and used web server, the Apache Web Server. But also, relevant within the context of, of here in our big data summer school, to things like Hadoop. Which is the defacto big data technology, in terms of processing, right now. It's also home to emerging big data technologies, like Spark and Shark, and Which are, basically, these sort of in memory data analytics. And so, you know, you should be interested that, those types of things if you're interested in big data. So those software that we contribute at Apache where I'm on the board of directors, is also what I use to each to other students. I also professor at the University of Southern California in the Computer Science department there. I teach two classes, one in search engines and information retrieval, big surprise [LAUGH]. And then, which I'll be lecturing, to you guys later on, during, this, summer school here, at Cal Tech. And then, another class in, another big surprise, software architecture, which is what I'm going to cover today. Software Designs, so they're both graduate classes, some of them are available remotely, as well. I encourage you to reach out to me and I'll mention that sort of during the remainder of my lectures here for you guys today. So some quick notes on this lecture, its the talk is definitely optimized for breadth and not depth, and that's for a reason, I'm going to try and go through, I'm going to try and give you the history and the applicability of software architecture, the big data which. Software architecture in terms of the software engineering research community has a pretty storied history over the past 20 years. It has a golden age, a heyday, things like that. There's no way in a 45 minute lecture cut into three that I'm going to be able to cover the entire sort of depth of some of these topics. So I'm going to point you in areas. To things, give you sort of a breadth of that. And, you know where I don't cover something, I'll be sure, in the middle of the lecture and the topic, to kind of point and say, hey, you might want to look here for further research. And I have a slide, sort of at the end, that has some, pointers to future research at the end of the third lecture here. You might also want to take a look at the syllabus from my USC software architecture class CSCI 578, the link there is right there on the slide for you guys so have a look at that. And you can get slides and other things to download to get some of the depth in some of these topics that I'm going to cover. You know, so in other words, if you really are interested in architectural modeling. In terms of what I'm going to teach you guys about today and you know, you don't get enough, if you're really interested in architecture description language and things. You want to look that up. Take a look at the syllabus here that's like here my USC syllabus. It outta help you. And then I just encourage you I'm really open, send me and email, my email address is right here on the slides. Or you know, contact me on Twitter. You can probably find me on Facebook, Google+, LinkedIn [LAUGH], not a places that you can't find me. In an IRC chat room somewhere. So, you know, see if you can find me and I'll do my best to help, if you can kind of connect with me here after the class. And we really welcome you here to Cal Tech for the Cal Tech, JPL Cal Tech Big Data Virtual Summer School. Okay, so let's get started with some actual material. Software architecture and, and why do you care. So, you care about it because the way that you design your system, even before you make technology choices, has a lot of applicability to the way that your software and your big data system is going to sort of come out the other end. If you've ever heard of the analogy, garbage in, garbage out. And has a lot more to do with what stack you pick, what big data technology you're deploying. Then in terms of the way that you're actually thinking about the design of your system, there are a lot of analogies to building architecture, you know. When we construct, a building like the one that I'm presenting in to, here to you guys, and we don't start out by, you know simply laying, copper wire. And, you know, laying the plumbing and things like that, without having sort of a detailed plan. And so it's, just, the same is true in software. It's really important. All software architectures have some some software architecture whether or not it's implicit in the design, even if you start coding, you know, believe it or not there's a software architecture in your mind when you're doing that, or explicitly described. Okay, so there's a lot of analogies, you know, here, and there's some other sort of tangential analogies and differences, between building architecture and software architecture. One of them that's, that's really obvious, and really important, too, is that buildings and building materials are, are malleable. You can touch them, they're tangible, you can. Can change them. You know the impact, for example, if you decide, that a room should be of a certain, you know, height, an, and depth, and width. And you know the impact very quickly if your measurements, or if your thoughts about that, in your design were wrong. From an engineering and construction perspective. For example, a door won't fit, you know, where you intended it to go. In software, that's not exactly the same scenario. You don't really have that understanding. Software is, is not tangible. You can't touch it. It's not clear to us sometimes the impact of our decisions. Even at the design level but at the implementation level as, as well. But even more so, the design level. Before you start coding, if you're just simply thinking about. What's the impact of, if I pick a, request reply protocol here in my big data architecture, and what type of downstream impact will it have on data dissemination? Versus if I picked up peer-to-peer type of, you know, data communication protocol and things like that. So, so the, the impact to some of these decisions isn't as, sort of, obvious, so. Even though software architectures are related to building architectures, building architectures in tangential or I'm sorry, tangible materials give us a lot more sort of testability and coverage and thoughts than software. Which is why you know, this isn't a science exactly yet. And, and, a lot of it is based on the ability of software that we tested, have adequate test coverage. And to have an adequate architectural model, that effectively represents what you're actually implementing. Okay? And so we'll talk about that during the remainder of the lecture. Software architecture has sort of enjoyed at least 20 years of sort of fundamental research. Some of the early research on software architecture was done. By Dwayne Perry and Alexander Wolf in 1992. And they had a sort of seminal paper called the Foundations for the Study of Software Architecture. That they published in Software Engineering Notes. Which is an ACM publication. Or SEN. And basically what Perry and Wolf will tell you is that software architecture really boils down to thinking about the design. And thinking about the design of software and software as elements. Form, and rational. And what they're talking about there are the elements of your software system in the form of components, the units of computation, the form, the way that you put those components together, and the rational, for, the rational for doing so. Why did you pick this number of components to decompose and break. Your software system down into. Why did you pick a component representing, this particular protocol or data conversion? Or why did you, do this or why did you do that? So, the rationale for, describing. Another very, sort of kind of the next step in some of our architectural understanding, from a research perspective, came from David Shaw. I, I'm sorry, from Mary Shaw and David Garlan. In 1996 and they wrote a book called Software Architecture Perspectives On An Emerging Discipline and that book for many years up until recently was sort of the canonical software architecture text that was taught at the university and graduate level. And the contributions out of that book you know, all state basically you know, in my opinion are basically. The identification, the first identification; so they build upon Perry and Wolf's definition of Software Architecture, form elements and rationale. Except they sort of decompose the the elements, into more than just components. More than just the units of computation in your system. They basically argued that you know, basically the way components in a software system. Interact, or something called the connectors, which are models of the interactions between software components, they argued in their book that connectors should have a first class status in software architecture. That the way that the interactions occur is really important, for them. So we'll discuss some of that here, sort of in the reminder of the slides today. And then finally, if you're thinking about fundamental software architecture, research, Philippe Krutchen, who was one of the folks behind, the rational unified process, was heavily involved in rational rows and other things. He published a paper in 1995 called the 4+1 model view of software architecture. That's another paper that's in which he's basically trying to think how do we model software systems, effectively what, what types of things should it have in terms of their elements. And his definitions are very consistent with Perry and Wolf, and with Shaw and Garlan. Where Krutchen sort of added something is he stated, you know, there's a lot of. The, the, there's a lot of aesthetics to software architecture. Software architecture, what, what might be a good architecture to me, you know, where I maximizing some particular properties, extensibility, composibility, reliability, may not necessarily be a good architecture to you. right? You may care more about. Its capabilities in terms of data conversion or its flexibility in terms of the interaction protocol of things like that. So the same two software architectures though may they maybe similar in cardinality and size and eventually when we get down to code or implementation. Maybe different to other people based on their sort of user defined quality and their perspective on that. So for Krutchen, sort of introduced this notion of aesthetics, that, that's really, is sort of important there. Sort of another, emerging and, and building upon philosophy for software architecture came out of Medvidovic and Taylor. And where they, where, where they sort of built on is they're building on sort of Shaw and Garlan and the notion of, Perry and Wolf, the notion of components and connectors. They built on Krutchen having this sort of aesthetic capability. Medvidovic and Taylor basically suggested software architecture boils down to the following canonical things and that's going to be the definition we use here in the class. And that's sort of the thing I want to convey to you because it's consistent with the other. Definitions, and it's pretty much kind of what we're doing today. Software architecture consists of the components, which are the units of computation in the system. Right? The components may compute on things. They may maintain internal state. We may have components for data processing. We may have, components that basically are responsible for conversion of things, for sending, information. Connectors, which are effectively the way that components interact, request replied protocols, event-based, publish-subscribe interactions, implicit invocation, these types of things. How are the components interacting amongst one another, and how do other connectors. Interact too, sometimes you have connectors which interact with other connectors. Pub sub is a great example of that or request replier federated protocols. So we've got components and connectors and then we have configurations which are arrangements of components and connectors. In other words your architectural typology and any rules that sort of guide that composition. Okay, this component may be connected to, this connector, only two components may not be able to be connected to one another without a connector in between them. These two components can never communicate, because there's no path through the typology of the architecture for them to communicate, so on and so forth. So the core definition sort of building upon this of software architecture that we're, we're leveraging in this course and that's sort of defined in kind of this newer book on software architecture that's sort of rapidly merging as the standard textbook. Is this book from Taylor, Medvidovic and Dashofy on, on Software Architecture. And basically in their book, they're discussing architecture as the following in terms of the definition, and it sort of builds on this. They're the principal design decisions about a software system, right? So when you're thinking about your big data software system, what are the principal definitions that you're thinking about there for the design? And what makes a decision principal? Well, building upon sort of what we're learning about in, in the prior slides, a decision that makes it principle might be it affects one of the core architectural elements. It's related to some component some computation in the system. It's related to the interaction, maybe. Maybe it has to do with the implementation, but it's so important because it represents some sort of load bearing wall or a bottleneck in that implementation. Maybe it has to do with system evolution. Maybe it's a component we're going to want to switch out later, or it's a decision or principal decision that is going to have some impact on the way that we sort of further evolve the system after we initially develop its code. And then there are others. Okay, so these are various examples of sort of what makes a decision principal, and as you can tell from these various examples, it may be dependent on different stakeholders in the software system. So. Right, your manager may have a different, perspective, of what represents a principle architectural decision than, say, your buddy, tester, developer. Okay. What are some, sort of examples, you know, of software architectures? Here is an example. This is a, very popular, content detection analysis toolkit, called Apache Tika. Tika is been referred to as the digital babel fish, I'll be talking to you a little bit about Tika in one of my further lectures here in the summer school. Effectively what you can think of it is as a small library that handles identification of file types, extraction of text and metadata from those file types, and then identification of language, and potential translation of those language. Tika here if, is thinking about it in the context of software architecture. represents a number of components. I know that they are components because the nice person that made this diagram for you guys, me, [LAUGH] put in a legend for you that states when I see a box. Here in the diagram that represents a software component. And what are those arrows between the software components there. Well there's different types of variables. One has a dot or some type of dash lines on it. One is like a solid black line what do those mean. Control flow is the sort of dashed arrow. Control flow meaning the transfer of control from one component to the other. Data flow is a black line, that's how data flows between the components, okay. I see something here, called a, a, registry or a repository, seems to be represented by that sort of familiar data base E shape. And so on and so forth, right? So, so regularly software architecture is described in this way. You'll notice I haven't stated what programming language Tika is implemented in, I haven't stated. You know, what, what sort of physical host this is deployed on, is it put on your Mac Book, is it, you know, are you putting this on windows, is it a virtual machine? All that sort of discuss the sort of architecture at, is at the level of these components, the connectors, which represent in these diagrams those arrows, control and data flow interaction. Amongst the components. Right? And, and potentially some other meaningful or important annotations, here that are present on the diagram. Right? So, so this is a common way to represent software architecture. Now thinking about this, if we built a whole bunch of these. Sort of content detection and analysis tool kits we may find that our architectures after awhile may look the same. They may use familiar components you may use some type of language identification component you may reuse or use some similar text extraction parser or some type of component for that. So over time, what we find typically in domains, and you find this in the domain of big data too, is that an architectural style or sort of emerges. Common sets of things like components and connectors and we'll talk about that. Architectural patterns, which are kind of common arrangements of, of these things with a little bit of the information left unspecified. They start to emerge. A commonly used, maybe, vocabulary. For talking about architectures within this domain start to emerge sort of as well. So architectural styles boil down to commonly used sort of components and connectors and potential their types or their classes. That you see, in a family of software systems, built in a domain. Normally, across many years, okay? It takes a while for software architectural styles to emerge. Some examples are things like peer-to-peer, right? In a peer-to-peer software architectural style, and, and in software systems that implement the peer to peer style, we know that there's one component type. There is a peer. Okay, and we may instantiate many of those things but effectively we know that the type of component is a peer at maintains some internal state, potentially, we know that peers communicate directly with other peers, okay? So the connections and, and typically in an arbitrated way. Through some initial interaction in which they locate all the other peers on the network, and then those peers addresses or some root set of them, which forward along communication to other peers and so on and so forth in this sort of organically. So we know a lot about a system simply by stating it's a peer to peer system, right? We know about its components, its interactions, and things like that. Client server rest, layered, these are all other examples. Okay, of software architectural styles. Think about it, I have a question for you guys to ponder here; while you're learning here in our summer school. What types of architectural styles do you regularly see emerging in big data systems, Okay? Are they combinations of these styles, like do you ever see a system that's strictly peer to peer or strictly client server, or do you see ones that are more hybrid? Okay, and that's sort of an open question, there's no correct answer to that, just sort of think about that while we're talking. So, when I think about architectural styles, I like to think of styles as things like ingredients. If you're thinking about things like ingredients in recipes, and food, and stuff like that. And I do a lot, right, I'm a hungry guy, you can see by my stomach. I, I tend to, [LAUGH] tend to think about food a lot. So, so architectural styles think of them as sort of the ingredients in your recipe, right. Like, like what, how much, what, what salt or I need pepper or I need vegetables or I need, you know, carrots or this or that. Architectural styles are more, sort of ingredients, okay? So you might have heard, if you think about software architecture too, things like patterns, code patterns, if you've read the code patterns book by Gamma, and, and things like that. patterns, you might have heard in terms of things like model view controller or. Three tiered pattern or sense compute control. Architectural patterns are typically lower level than architectural styles and what I mean by lower levels is that styles are mostly like what I said think ingredients right so patterns take us closer to the actual implementation of software. Patterns may tell us, yes, there are these particular styles of components or connectors in our software system. And they may tell us basically that we arrange them in a particular way and they communicate in this way. Leaving sort of, outside of the box for that some information to be specified. We arrange in, for example, a three tiered pattern. We have three components. And three component types, we have a sort of UI or presentation tier. We have a logic tier. Sort of that the UI and presentation tier communicates with and then we have some back end data tier. Or, or, or component or you know, whatever you want to call it. So we know a lot about a system too, we know even more than a style by talking about patterns. Okay for that and so I like to think of patterns more like recipes right. We may take the ingredients like that we have from styles, salt and pepper and things like that and vegetables, and put them together in to a bolognese sauce. Okay leaving for you to specify in that, if you want to make a variation on the amount of salt or the amount of pepper or whether or not you like carrots in your bolognese or not or whether you strictly like things like celery or things like that. so it takes us closer to sort of the implementation of pattern does, but it doesn't prescribe or provide everything okay. And finally we're going to wrap up this first part of the module just talking about Reference Architectures which were a term that was originally coined in the architectural domain by Will Tracz in a paper in which he was talking about architectures and domain specific software architectures. And ACM Software Engineering Notes way back when in 1995. And Tracz basically said that reference architectures have sort of three common things to them. They've got kind of common architectural styles and patterns, found in a particular domain like grid computing or big data or avionics. They've got reference requirements that drove the sort of derivation of those components, those styles, those patterns, and things like that. Then they typically have a domain model or common vocabulary for talking about systems and data and components within that domain. Okay, so, so domain specific reference architectures are typically or domain specific software architectures are typically reference architectures for a particular domain, like for avionics. And so forth. So we use these sort of interchangeably for that. You know, examples within big data, you might think about [INAUDIBLE] being a reference architecture, okay, for a particular domain. Common [INAUDIBLE] of components and connectors and things like that. You might think about systems like Globus, which was very popular. You know, many years ago, but still has some popularity and remains that way today. Especially with things like Globus online and so on and so forth. So, that represents the end of the first, sort of ten minutes. or, sorry, not ten minutes. The first, part or module here, in our Big Data fundamentals. So we'll move on, in the next module, talking about some advanced. Architectural concepts in terms of architectural modeling. Architectural visualization and, and so on and so forth.