Hi, welcome to module 9.3 of digital signal processing. We're still talking about digital communication systems. In the previous module we addressed the bandwidth constraint, and in this module we will tackle the power constraint. So first we will introduce the concept of noise and probability of error in a communication system. We will look at signalling alphabet and the related power. And finally we'll introduce QAM signalling. So, we have seen the transmitter sends a sequence of symbols, a of n, created by the mapper. Now we take the receiver into account. We don't yet know how, but it's safe to assume that the receiver in the end, will obtain an estimation. Hat a of n, of the original transmitted symbol sequence. It's an estimation.because even if there is no distortion introduced by the channel, even if nothing bad happens. There will always be a certain amount of noise that will corrupt the original sequence. When noise is very large, our estimate for the transmitted symbol will be off and will incur a decoding error. Now this probability of error will depend on the power of the noise with respect to the power of the signal and will also depend on the decoding strategies that we put in place, how smart we are in circumvent and defects of the noise. One way we can maximize the probabilty of correctly guessing the transmitted symbol is by using suitable alphabets. And so we'll see in more detail what that means. Remember the scheme for the transmitter. We have a bitstream coming in and then we have the scrambler, and then the mapper. And here we have a sequence of symbols a of n. These symbols will have to be sent over the channel and to do so we up sample and we interpolate and then we transmit. Now, how do we go from bitstreams to samples in more detail. In other words, how does the mapper work? The mapper will split the incoming bitstreams into chunks. And will assign a symbol, a of n, from a finite alphabet to each chunk. The alphabet we will decide later what it is composed of. To undo the mapping operation and recover the bitstream. The receiver will perform a slicing operation. So, the receiver will receive a value of hat a of n where hat indicates the fact that noise has leaked into the value of the signal. And the receiver will decide which symbol from the alphabet which is known to the receiver as well is closest to the received symbol, and from there, it will be extremely easy to piece back the original bitstream as an example, let's look at simple two level signalling. This generates signals of the kind we has seen in the examples so far alternating between two levels. The way the mapper works is by splitting the incoming bitstream into single bits. And the output symbol sequence uses an alphabet composed of two symbols G and minus G, and associates G to bit value one, and minus G to a bit of value zero. At the receiver the slicer, looks at the sign of the incoming symbol sequence which has been corrupted by noise and decides that the nth bit will be one, if the sign of the nth symbol is positive and zero otherwise. Let's look at an example, let's assume g equal to 1, so the two level signal will alternate between plus 1 and minus 1. And suppose we have an input bit sequence that gives rise to this signal here. After transmission and after decoding at the receiver. The resulting symbol sequence will look like this, where each symbol has been corrupted by a varying amount of noise. If we now slice this sequence by thresholding, as shown before, we recover a symbol sequence like this. Where we have indicated in red the errors incurred by the slicer because of the noise. So if you want to analyze in more detail what the probability of error is. We have to make some hypothesis on the signals involved in this toy experiment. Assume that each received symbol can be modelled as the original symbol plus a noise sample. Assume also that the bits in the bits stream are equiprobable. So zero and one appear with probability 50% each. Assume that the noise and the signal are independent and assume that the noise is additive white Gaussian noise with zero mean and known variance, sigma zero. With this hypothesis the probability of error can be written out as follows. First of all, we split the probability of errors into two conditional probabilities conditioned by whether the n-th bit is equal to one or the n-th bit is equal to zero. In the first case when the n-th bit is equal to one, remember the produce symbol will be equal to G. So the probability of error is equal to the probability for the noise sample to be less than minus G. Because only in this case, the sum of the sample plus the noise, will be negative. Similarly, when the amplitude is equal to zero, we have a negative sample. And the only way for that to change sine is if the noise sample is greater than G. Since the probability of each occurrence is one half, because of the symmetry of the Gaussian distribution function, this is equal to the probability for the noise sample to be larger than G. And we can compute this as the integral from G to infinity of the probability distribution function for the Gaussian distribution with the known variance here. So what we have here is the tail probability of a Gaussian with standard deviation sigma zero. So, this is the integral that we're trying to compute. Usually the tail probability of a unit variance Gaussian is indicated by the notation Q of G. The Q function of G. Now because our variance is actually sigma zero squared, we have to normalize the Q function argument by the standard deviation. And we find that the probability of error is equal to the Q function of G over sigma 0. Since this integral is not computable exactly, if we want to calculate the Q function, we have to resort to numerical packages, or tabulated versions of this function. And usually, what you would find in numerical packages in, is a derived function called the error function, which is related to the Q function by this formula here. But the point that is really important at the end of this derivation is that the probability of error is equal to some function of the ratio between the amplitude of the signal, and the standard deviation of the noise. Now we can carry this analysis further by considering the transmitted power. We have a bi-level signal, and each level occurs with one half probability. So the variance of the signal, which corresponds to the power is equal to G squared times the probability of the n-th bit being equal to 1, plus G squared times the probability of the n-th bit being equal to 0, which is equal to G squared. And so if we rewrite the probability of error, we can see that it is equal to the Q function of the ratio between the standard deviation of the signal and the standard deviation of the noise. But this is really equivalent to saying that the probability of error is equal to the Q function of the square root of the signal to noise ratio of the transmitted signal. If we plot this as a function of the signal noise to ratio in dBs and I remind here that dBs here mean that we compute 10 times the log in base 10 of the power of the signal divided by the power of the noise. And since we are in a log, log scale. We can see that the probability of error decays exponentially with the signal to noise ratio. This exponentially decay is quite a norm in communication systems and while the absolutely rate of decay might change in terms of the linear constants involved in the curve, the trend will stay the same even for more complex signal and schemes. So the lesson that we learned from the simple example is that in order to reduce the probability of error, we should increase G, the amplitude of the signal. But of course, increasing G also increases the power of the transmitted signal, and we know that we cannot go above the channel's power constraint. And so that's how the power constraint limits the reliability of transmission. The bi-level signal is keen, is very instructive, but it's also very limited in the sense that we're sending just one bit per output symbol. So to increase the throughput, to increase the number of bits per second that we send over a channel, we can use multilevel signaling. There are very many ways to do so, we'll just look at a few. But the fundamental idea is that we take now, larger chunks of bits, and therefore we have alpha bits that have a higher cardinality. So more values in the alpha bit means more bits per symbol and therefore a higher data rate. But, not to give the ending away. We will see that the power of the signal will also be dependent on the size of the alphabet and so in order not exceed in the probability of error, given the channels power of constraint, we will not be able to grow the alphabet indefinitely. But we can be smart in a way we build this alphabet. And so we will look in some examples. The first example is PAM, Pulse-amplitude Modulation. We split the incoming bitstream into chunks of M bits. So that the each chunk corresponds to an integer between 0 and 2 to the m minus 1. We can call this sequence of integers, k of n and this sequence is mapped onto a sequence of symbols a of n like so There's a gain factor G, like always. And then we use 2 to the M minus 1, odd integers around 0. So for instance, if M is equal to 2, we have 0, 1, 2, and 3 as potential items for k of n. And a of n will be either lets assume G is equal to 1 will be either minus 3 or minus or 1 or 3. We will see why we used the odd integers in just a second. At the receiver, the slicer will work by simply associating to the received symbol, the closest odd integer, always taking the gain into account. So graphically again, PAM for M equal to 2 and G equal to 1 will look like this. Here are the odd integers. The distance between two transmitted points, or transmitted symbols, is 2G, right here G is equal to 1, but it would be in general, 2 times the gain. And using odd integers creates a zero-mean sequence. If we assume that each symbol is equally probable, which is likely, given that we've used a scrambler in the transmitter, then the resulting mean is zero. The analysis of the probability of error for PAM is very similar to what we carried out for bi-level signaling. As a matter of fact, bi-level signaling is simply PAM with m equal to 1. The end result is very similar and its an exponential decaying of the ratio between the power of the signal and the power of the noise. The reason why we don't analyze this further is because we have an improvement in store. And the improvement is aimed at increasing the throughput, increasing the numbers of bit per symbol that we can send without necessarily increasing the probability of error. So here's a wild idea, let's use complex numbers and build a complex valued transmission system. This requires certain suspension of disbelief for the time being, but believe me, it will work in the end. The name for this complex valued mapping scheme is QAM, which is an acronym for quadrature amplitude modulation. And it works like so. The mapper takes the incoming bitstream and splits it into chunks of M bits, with M even. And then it uses half of the bits to define a PAM sequence which we call a of r of n. And the remaining m over two bits to define an another independent PAN sequence a i of n. The final symbol sequence is a sequence of complex numbers where the real part is the first PAM sequence and the imaginary part is a second PAM sequence. And of course in front we have again factor G. So the transmission of alphabet A is given by points in the complex plane with odd valued coordinates around the origins. And the receiver just lies through works by finding the symbol in the alphabet which is closest in Euclidean distance to the received symbol. Let's look at this graphically. This is the set of points for QAM transmission with M equal to two, which corresponds to two bi-level PAM signals on the real axis and on the imaginary axis. So that results into four points. If we increase the number of bits per symbol, we set M equal to four. That corresponds to two PAM signals with two bits each. Which makes for a constellation. This is how these arrangement of points in the complex plane are called. A constellation of four by four points at the odd valued coordinates in the complex plane. If we increase M to 8, then we have a 256 point constellation with 16 points per side. Lets look at what happens when a symbol is received and how we derive an expression from the probability of the error. If this is the nominal constellation the transmitter will choose one of these values for transmission, say this one. And this value will be corrupted by noise in the transmission and the receiving process. And will appear somewhere in the complex plane, not necessarily exactly on the point it originates from. The way the slicer operates is by defining decision regions around each point in the constellation. So suppose for this point here, the transmitted point, the decision region is a square of side, 2G, centered around the transmitted point. So what happens, is that when we receive symbols, they will not fall on the original point, but as long as they fall within the decision region, they will be decoded correctly. So for instance, here. We will decode this correctly. Here we will decode this correctly, same here. But this point for instance falls outside of the decision region and therefore it will be associated to a different constellation point. Thereby, causing an error. To quantify the probability of error, we assume as per usual that each received symbol is the sum of the transmitted symbol plus a noise sample. Inter of n. And we further assume that this noise is a complex value Gaussian noise of equal variance in the complex and real components. We're working on a completely digital system that operates with complex valued quantities. So we're making a new model for the noise. And we will see later how to translate the physical real noise into a complex variable. With this assumptions. The probability of error is equal to the probability that the real part of the noise is larger than G in magnitude, plus the probability that the imaginary part of the noise is larger than G in magnitude. We assume that real and imaginary components of the noise are independent, and that's why we can split the probability like so. Now, if you remember the shape of the decision region, this condition is equivalent to saying that the noise is pushing the real part of the point outside of the decision region in either direction and same for the majority part. Now if we develop this, this is to equal 1 minus the probability the real part of the noise is less than G and the imaginary part of the noise is less than G. This is the complimentary condition to what we just wrote above. And so this is equal to 1 minus the integral over the decision region, d, of the complex valued probability density function for the noise. In order to compute this integral, we're going to approximate the shape of the decision region with the inbound circle. So instead of using the square. We're going to a circle centered around the transmission point. When the conciliation is very dense, this approximation is quite accurate. With this approximation, we can compute the integral exactly for a Gaussian distribution. And if we assume that the variance of the noise is sigma 0 squared over 2 in each component, real and imaginary. It turns out that the probability of error is equal to e to the minus g square over sigma 0 square. Now to obtain a probability of error as a function of the signal to noise ratio, we have to compute the power of the transmitted signal. So if all symbols are equiprobable and independent. It turns out that the variance of the signal is G square times 1 over 2 to the power of M, which is a probability of each symbol times the sum of over all symbols in the alphabet of the magnitude of the symbol squared. Now it's a little bit tedious but we can solve it exactly for M and it turns out that the power of the of the transmitted signal is G square, 2 3rds, 2 the M minus 1. Now if we plug this in to the formula for the probability of error, that we've seen before, we get that the result is an exponential function where the argument is minus 3. That multiplies 2 to the minus m plus 1, that multiplies the signal's noise ratio. We can plot this probability of error in a log log scale, like we did before. And we can parameterize the curve as a function of the number of points in the constellation. So here you have the curve for a four point constellation. Here's the curve for 16 points. And here's the curve for 64 points. Now, you can see that for a given signal to noise ratio, the probability of error increases with the number of points. Why is that? Well, if the signal to noise remains the same, and we assume that the noise is always at the same level. Then it means that the power of the signal remains constant as well. In that case, if the number of points increases G has to become smaller in order to accommodate a larger number of points for the same power. But if G becomes smaller then the decision reaches become smaller. The separation between points becomes smaller and the decision process becomes more vulnerable to noise. So in the end, here is the final recipe to design a QAM transmitter. First you pick a probability of error that you can live with. In general 10 to the minus 6 is an acceptable probability of error at the symbol level. Then you find out the signal to noise ratio that is imposed by the channel's power constraint. Once you have that, you can find the size of your constellation by finding m which based on the previous equations is the log in base 2 of 1 minus 3 over 2 times the signal to noise ratio divided by the natural logarithm of the probability of error. Of course you will have to round this to a suitable integer value and potentially to an even power of 2 in order to have the square constellation. The final data rate of your system will be m the number of bits per symbol, times w which if you remember is the board rate of system. And corresponds to the bandwidth allowed for by the channel. So, we know how to fit the bandwidth constraint by upsampling. With QAM, we know how many bits per symbol we can use given the power constraint. And so we know the theoretical throughput of the transmitter, for a given reliability figure. However, the question remains, how are we going to send complex value symbols over a physical channel. It's time therefore to stop the suspension of disbelief and look at techniques to do complex signaling over a real value channel.