Hi there, my name is Amy Braverman. I'm a statistician at the Jet Propulsion Laboratory, and welcome to the JPL Caltech Virtual Summer School on Big Data Analytics. And the modules here are on inference and uncertainty. As I said, I'm a statistician, and so I'm very concerned with these two topics, and I want to tell you a little bit about. How I think they're relevant for big data analytics. So before we begin, I'd like to just say a few things the goal of the modules is to present statistical inference and it's role data analytics in the simplest terms possible. There's obviously no way in this short amount of time. That we would be doing the equivalent of a full semester course in even a subset of the topics that actually belong to inference and uncertainty. So the way I approached this was to look back at a number of textbooks that I'm fond of and to ask myself what I thought the key. Items were that I would want to impart to you during this, set of lectures. And what that means is I tried to pick the absolute fewest number of topics because I don't want to overwhelm anyone and to present in a very heuristic, and intuitive way with the minimum amount of math. That it is possible to even discuss this stuff with. I do assume that everybody has a basic knowledge of at least a little bit of calculus like knows what a derivative is and you know what a limit is and I am going to, with apologies to my mathematically oriented colleagues, I'm going to present this information. In a very conversational way, and without stating all the possible conditions one might really want to check if one was to approach this as a math problem. These lectures are not intended to be either comprehensive or thorough. They are intended to be something of a survey of what I regard as the key. Topics from, within this subject matter. And that said I have three broad topic areas. I have review of basic probability, basic concepts of inference, and an introduction to two popular nonparametric procedures for performing inference. That you may or may not be familiar with. Some of the things that I discuss in the basic probability lecture are therefore completeness. We will use some of them as we proceed into the basic concepts of inference. And some of the basic material in the probability section will not appear again, but it would be hard to leave it out in any discussion of basic probability. So with that said, I'd like to move on to the introduction where I'd like to try to give you a broad, broad picture of what I'm trying to get at here. And I'm going to also, I define a few terms for you that will be good to know. And that we will probably end up using again later. Okay. So this is my basic cartoon of what goes in statistical inference. We have an unknown population. Sometimes we might want to call that a process, particularly. If we are looking at scientific, scientific disciplines where we're talking about modeling physical, physical, processes or, phenomena, and we understand that process to be generating, Observations. And some of those observations we actually get to see through the process of sampling. So the way I've drawn this, I've drawn a little probability distribution, a cartoon probability distribution, for the unknown population or process. Because we're going to use probability models to model. How that process or population behaves or looks like. And the, my little sampling cartoon there is supposed to look like a screen that lets certain of those items that are produced by the process through into the sample that you actually get to see, or that we actually get to see. And typically what we do is we. Estimate something or draw a conclusion from the sample that we actually have. And then we try to make a statement about some feature of the unknown process or population that we care about based on what we learned by interrogating the sample. And that's why written inference with quantified uncertainty down there at the bottom. So we say that sampling supplies us with realizations from the probability model that describes the population. And what we have to recognize here, is that in the one sample that we actually got, we could have gotten another sample. And we might have drawn a different conclusion. From that other sample and it is the that possibility, the possibility of getting more than one different conclusion that leads to there being uncertainty in whatever it is that we infer. Okay, so here's my little cartoon picture of what's going on there. This is what we would really like to be able to do. We would like to be able to collect many samples from our population, compute on each one of them, and then look at how different the results are, and I've put that over here on the right with this little histogram with a density curve plotted on top of it, and we call that the sampling distribution of the statistic that we computed, and that will give us some idea. Of how different things could have been. Now, in, in the real world, we often don't get to have more than one sample. We only get the one. So we're going to have come up with some ways of understanding what that sampling distribution looks like in order to make our inference. I will say since this is a summer school on big data analytics that having our one sample be extremely large actually gives us a way to obtain additional samples by sampling from the one sample we have. And we'll come back to that at the end of this little section and talk about that a little bit. Because that maybe a key for leveraging the information in a massive data center or bid data, as we're calling it. So we want to infer the prop, the characteristics of the true probability model that describe our population. And we want to infer it from. Let's go back to the old fashion way, from just one sample that we have. And we need to find a way to understand not what the computed value is but where it sits in that notional sampling distribution. And there are generally two schools of thought on this. Statisticians tend to fall into either the frequentists, sometimes also called, classical statistician. Camp, or the Bayesian Camp. And many of you may have heard those terms before. And we will come back to that. I'm pretty much going to stick to the frequentest point of view in these lectures, because, like I said, I'm trying to be as uncluttered as possible. But I will say something about Bayesian inference later on. So, now what I'd like to do is define, what I think are a few key terms that you, may or may not have heard before. An observational unit is an object about which we want to know something. So if I was interested in, my. Was interested in all the students at Caltech, let's say. That an observational student would be one student at Caltech. A variable is a quantity we measure on an observational unit. So, let's say it's the height of all the students at Caltech that I would like to know about. So I would measure height on the observational units. The population is the entire collection of observational units, that would be all students at Caltech, and we should at this point distinguish between finite populations and infinite populations. The example I just used, finite populations. Would be all the students who are here at Caltech today or all the U.S. citizens that are alive today. Because in principle we could actually figure out how many of them there. And there would be a countable number of them. But there's also a concept of an infinite population, where we might talk about all citizens of the United States that ever were or ever will be. Or all Caltech students there ever were or ever will be. I'm not sure that's actually infinite, but, you get my meaning. And the concept of an infinite population is actually especially useful when we think about physical mechanisms, things like surface temperature on the Earth. Surface temperature where you might ask. You might ask what's the observational unit. Is it a point location? Or is it my block or my census tracked or some other larger aggregation of space, that defines the observational unit. And there are many ways to think about that. And if we really wanted to get down to it, and I know many of my physics friends like to think this way. Space would be continuous. And we would be thinking about the process that describes temperature at any arbitrary location in a, let's say a two dimensional space. So that notion of an infinite population and the notion of using a probability distribution. To describe that infinite population actually comes fairly close to the kind of thing that we often want to do in science. okay. So a parameter is a quantity computed from all the units in the population. All right. A sample is a subset of the population and we may obtain it purposefully or. Coincidentally and a statistic is a quantity computed from that sample. And note that the statistic itself will be something that's random because the sample is random. The sample might have been different therefore the calculated value of the statistic might have been different. I have this little side bar here about an observational study versus and experiment because I think that's an important distinction. Back in the old days when statistics was young and one of the original applications was to understand the impact of different treatments on, let's say, how well your corn grows. People did what were called experiments. They would divide up their plot in a particular way and they would add fertilizer and water in different combinations, and then they would measure the height of the corn. And try to determine which is the best combination of fertilizer and water to get the best corn crop. That's an experiment because the observational units were actually being changed and affected by that experiment. So I'm calling that si, circumstance an experiment. In the kind of stuff that we do at JPL, we have a some what different situation. We, let's say we fly a satellite that looks down and measures or observes some characteristics of the Earth's surface. The satellite flies at a prescribed orbit that is not under our control. It may have been under control of the people who designed the mission but. At the point where we enter the picture it's not under our control. And typically that satellite flies in an orbit that has some regularity. So, that's what we call an observational study. We do not have the option of actually experimenting on the experimental units or on the observational units that we want to understand. We simply observe what's there. So that's an important, an important distinction to keep in mind. Let me now say something about exploratory versus confirmatory analysis because I think this is where, I'd like to distinguish between this module and some of the other modules that you might have heard, particularly from my machine learning colleagues, you may have heard the terminology exploratory vs confirmatory analysis. Exploratory data analysis sometimes called EDA is a popular term, and that pertains to taking the data that you have, as I've shown it here, the data from the sample. And exploring it as the name says. Maybe making plots, looking at it, computing some summary statistics, or perhaps even doing something extremely sophisticated like using a machine learning algorithm such as the one Dave Thompson might have talked about to you. But in any case you are applying that algorithm to the data you have in front of you. And that can be a daunting tak because those data sets can be large and there's a lot of clever things that can be done there but in the end at some point I think that our objective will be to draw a conclusion about the population from which those data were obtained and that's the job of confirmatory analysis or inference. Where we have to understand the uncertainty of the result that we obtain from the one sample that we have. So let me go here, EDA illuminates structures, patterns, relationships, and so forth in the sample. EDA is often necessary to formulate hypotheses about the unknown population. One of our major problems in big data is that we often don't actually know what's in the data, we don't really know much about it at all, and in order to form a hypothesis that we might want to test later, or to know what's important, we have to go through the process of understanding what's in at least the sample in front of us. In confirmatory data analysis, we use the tools of statistical inference to make a definitive probabilistic statements about the population, based on what we learned from the sample. And I'll throw out the notion that there are two tools of statistical inference. That turns out that these are really flip sides of the same thing hypothesis testing. And estimation and I'm not going to talk about hypothesis testing in these lectures, just about estimation. So, let's get back to massive datasets or big data and I allude at the beginning of this module, sometimes the sample that we have is too big to treat like our friends from a hundred years ago that would like to treat a sample. Sometimes it's even too big to get on to our computer. And so there comes a question about how to interrogate that sample for EDA purposes and then how to make inferences from it. There are in general two strategies for that. One is bigger, better, faster algorithms and machine learning provides a lot of those. And the other strategy is to make the data smaller, where you might sample again, from the big data sample that you have. Or you might choose to apply some algorithm that reduces it in some way to a more manageable form. And, I will the machine learning algorithms to the machine learning experts, and I'm not really going to talk about data reduction, although, we're allude to it again later. So finally, let's talk about the question of whether massive data sets are populations or samples. They're kind of both actually. They are samples in the sense that they are a random selection of some kind from a larger population or process but they're also like populations in that we have to do something else, we have to sample from them again, we have to. We know that we can interrogate them completely and understand them completely in the way we might have treated the smaller samples. So, here we have an option and I, as I said I alluded to this earlier. To perhaps sample again from our big data sample and now we could sample many times and we could actually make that histogram on the right. By computing the statistic of interest over those samples-the secondary samples. And, wouldn't it be nice if we could, some how, compute the thing we really wanted to know on the big data sample, and then look at the sampling distribution that we have on the right. And from that relationship, somehow infer what the relationship between the computed value from the big data sample was and the true population. And people are working on methods for doing things like that and thinking about that. I will not go that far today, but it's out there in the literature. There's something developed up at Berkley called the bag of little boot straps which is starting to get in this direction. And you may want to go have a look at that. So with that I think we'll move on to the first module on probability which is we will just do a brief review of probability theory.