Hi, and welcome to module 7.1 of Digital Signal Processing. Today we will talk about Stochastic signal processing. Stochastic means random. And so we'll start with a simple example of a random signal. We will describe its power spectral density, which is the spectral representation of a random signal. We will see what happens when we filter a stochastic signal. And then we will talk about the concept of noise. So why stochastic signal processing? So far in the class we have assumed that every discrete time signal we used, could be described exhaustively. Either via a close form representation like so, or, and this is really the advantage of discrete time signals, simply by enumerating its non zero values. However interesting signals are not known in advance. For instance what I'm going to say next contains information that this not known in advance. But I do know that what I'm going to say next, is going to be a speech signal. So I could try to describe stochastic signals in terms of a probabilistic model. The good news is that we can do signal processing with stochastic signals. Using the same tools that we have developed so far. In this module, we will not try to treat the subject of stochastic signal processing, either exhaustively or very rigorously. But we will try to give you enough intuition and the mathematical tools to deal with ubiquitous random signals, such as noise. So let's start with a simple example of a stochastic signal. Suppose we generate a discrete time signal by tossing a coin for each sample and setting the sample's value to +1 if the outcomes head. And -1 if the outcomes tail. So because of the mechanics of the coin toss, each sample is independent from all others, and each sample has a 50% probability of being +1, and a 50% probability of being -1. So this is our stochastic signal generator, and every time we turn on the generator, every time we repeat this experiment of coin tossing, we get what we call a different realization of the signal. We can plot it, and so for instance the first run of cost gives us this series of +1 and -1. And then if we repeat the experiment again, we toss the coin another 32 times. We got a different realization, and we can repeat the experiment as many times as we want and the outcome will likely be different every time we run the machine. Although we cannot describe in advance the values of the signal, we know the mechanics behind the generation of the signal. And the question is can we analyze this signal for instance can we get a spectral representation of a random signal? So, let's try by taking the DFT of a finite set of random samples. So, if we do that, we just run the machine for 32 samples and then we take a DFT. But every time we'll repeat the experiment, we get a different plot. And the values don't really seem to follow any pattern. So maybe, we tell ourselves, maybe we don't have enough data to discover a pattern in the DFT, so maybe we should take longer realizations. Instead of 32 points, maybe we should take more. And so let's try and take 64 points. And still we don't see any pattern. So let's try and take a 128 points. And again, the spectrum doesn't seem to show any trait that we can readily understand. So we need a new strategy. When faced with random data an intuitive response is to take averages. So we average out the values in order to eliminate the fluctuations. In probability theory the average is computed across a realization not along the time access. But, across different repetitions of the experiment. So, for the coin toss signal, the expectation of each sample is minus one times the probability that the n-th toss is tail plus one times the probability that the n-th toss is head. But we know that these probabilities are one half each, so this sum ends up being zero. Since the DFT is a linear operator averaging the DFT values will not work either. If we try and do that we quickly realize that the average of each DFT sample will be zero as well. However, the signal does move from minus one to plus one. So its energy or it's power must be non zero. If you remember the definition of energy and power of the signal from module 2.1. We can readily see that the energy is infinite. Because the limit for N goes to infinity of the sum. From minus n to n of the values of the sequence squared It's simply two n plus one. Because each value squared will be equal to one. And so, as n going to infinity. this will diverge. However, the signal does have finite power over any interval. We take the energy over a minus n to n interval, and we normalize that by the length of the interval. And we find out that the power is actually 1, regardless of the interval's length. Let's try the following strategy. Let's try to average the DFT's square magnitude, normalized. So we pick an interval N. We pick a number of iterations, so a number of times that we will repeat the experiment. We run the signal generator M times, and we obtain M N-point realizations. We compute the DFT of each realizations, and we average their square magnitude, divided by the length of the interval. So if we do that, well, of course, the first DFT will be this random pattern that we have seen before. But as we increase the number of realizations, we see that the points seem to converge to something. And indeed, by the time M hits 5000, we see that the average of the squared managed of the DFT seems to converge to the constant one. So we have defined a quantity here, P[k] which is the expected value of the squared magnitude of the kth bin of the of the DFT over capital N points. Divided by capital N, and it looks very much as if P[k], this expectation, is equal to one for all k's. So, if the square magnitude of the DFT tends to the energy distribution in frequency of the signal. Then the normalized square magnitude of the DFT. Tends to the power distribution, or the power density, in frequency. So what we just have derived is a new frequency representation for signals that have infinite energy but finite power, and it's called the power spectral density. Let's try to develop some intuition about the power spectral density of the coin toss signal. The fact that it is constant. Means that the power is equally distributed over all frequencies. In other words, we cannot predict if the signal will move slowly or super-fast. We cannot predict that because each sample is independent of each other. So we could actually have a realization where just by luck. We have a constant signal because all coin tosses are heads. Or we could have a realization in which at each coin toss we have a different outcome and so the signal will oscillate at the maximum digital frequency. The power spectrum of presentation. Embodies this behavior by distributing the probability of power over the entire frequency axis. Lets try now to filter a random process. And to run this experiment, we take our coin toss signal once again. And we take what is probably the simplest filter we can think of just a 2-point Moving Average. So the output of the filter will the average of two neighboring points in the input random signal. And the question we want to find an answer to is, what is the power spectral density of the output? So we complete this numerically, we run the experiment as before. We generate a random signal. This time we filter it. And then we take the average over several realizations of the DFT squared of the filtered output. So, we run the experiment like before, which was an interval M, we generate the random signal. This time we filter it. And then we take the average over several realizations of the DFT squared. For M equal to one, we don't really see a pattern. But as we see that the power spectral density converges to a very precise shape. Now, if you remember the frequency response of the 2-point Moving Average, that is just H of e to the j omega equal to one plus e to the j omega divided by two. And this shape here Is nothing but the square magnitude of this frequency response evaluated at the DFT points, i.e. At a two pi over N K. So indeed, if we plot this we see that the match is almost perfect. The power spectral density of the output seems to be the power spectral density of the input, i.e. The constant one, multiplied, by the square magnitude of the frequency response of the filter this time computed on the DFT grid. We can generalize these results beyond the numerical experiments and move it to the world of infinite support stochastic processes. The details are in the book but here we will summarize the key points. So, stochastic process is characterized in frequency by its power spectral density and it can be shown. That the power spectral density is the DTFT of the autocorrelation of the process. Where each sample of the autocorrelation is obtained by taking the expectation of the product of the stochastic signal times a delayed copy of itself. For a filtered stochastic process, the general result is that the power spectral density of the output Is equal to the power spectral density of the input, times the frequency response in magnitude squared. The good news is that this result guarantees that the filters that we designed in the deterministic case can still be used with stochastic signals. A low pass will remain a low pass and a high pass will remain a high pass. We do however lose the concept of phase, and this is understandable, since we don't really have any advanced information of the shape on the stochastic signal. That will depend on the particular realization. All we know is where the power is distributed in frequency. Now that we have some stochastic tools in place, let's attack the concept of noise. So, noise is everywhere. It appears in the form of thermal noise and circuits. It could be the sum of various extraneous interferences in communication systems. Or the quantization on numerical errors that complex digital system produced and that we cannot predict in advance. Because of a lack of knowledge about the sources of the noise, we will model the noise as a stochastic signal. And the most important type of noise is white noise. With the term white we indicate a stochastic process where all the samples are uncorrelated. If the samples are uncorrelated, the auto-correlation of the process will be zero everywhere except at zero, where it will take the value of the variant. As a consequence, the power spectral density is the constant sigma squared, where sigma is the variance of the stochastic signal. Graphically the power spectral density of the white signal couldn't be any simpler. We have seen the example of this in the coin toss experiment. The power spectral density of white noise is independent of the probability distribution function for the single sample. Distribution however will be important to estimate the bounce of the noise signal in the time domain. Very often we use a Gaussian distribution to model the underlying probability distribution function for the samples. The reason for that is that the Gaussian distribution is the model of choice when we want to represent the effective, many unknown, super impose sources. As is the case for noise. In this case, the noises were called additive white Gaussian noise, or AWGN for short.