Hi and welcome to module 5.11. We have come a long way since the bare bone definition of a filter and it is now time to put to use everything that we've learned before and build a real time signal processing system on your pc. So here we will review the technicalities involved in implementing real time processing. On a general architecture and then we will review the code that will allow you to implement real time guitar effects on your PC. I hope you're going to have fun with this one. Hi and welcome to module 5.11 of digital signal processing. In this module we'll talk about real time signal processing, with a particular focus on general purpose architectures such as your PC. We will first examine how data transfer take place between your PC and peripherals such as the sound card. We will talk about multiple buffering which is the technique to ensure smooth data communication between the CPU and the external devices. We will look at the general implementation framework from the point of view of the API. And then we will look at how to implement some guitar effects for real time[INAUDIBLE]. When we talk about real time signal processing, we indicate a situation where, the maximum time that we can spend to process one sample. Is indicated by a system clock of period t s. We really haven't talked about sample and interpolation. We will do that in the next module, but you already know that to process a real world signal, we will have to record that in the form of samples. And we will record a sample every t s seconds. Then we will process this sample in a causal filter, because we are in a real-time situation. And the maximum amount of time that we can spend on the processing of one sample is again TS seconds. And then, every TS seconds, we will have to play out the output sample. So everything needs to happen in at most TS seconds. Let's look in more detail at the output process. We have a series of samples, X of M. We have a sound card, which is teh device that will bridge our digital world to the analog world of teh loud speaker. And the sound card operates with the system clock of t s. Now let's assume that our sequence, x of n, is generated by some algorithm. By a central processing unit. On a dedicated audio device, we could use the same system block to drive both the production of sample and the output process. So the system will be rather simple. Now on a PC on the other hand, the CPU works on a clock that is completely Independent of the sound card clock. As a matter of fact, there is a clock that is much, much faster then the clock you would find in your sound card. So this situation is the following, the CPU will generate samples, add its own pace, and put these samples. In memory. The sound card,on the other hand, will fetch the samples from memory at its own pace, which is dictated by its clock, and then output them to the loudspeaker. The transfer of data from the CPU to memory happens by a standard memory bus transaction. The transfer of data from memory to the soundcard takes place in the form of a DMA transfer. Dma stands for direct memory access. In other words, the sound card does not need to invoke the CPU in order to access the data in memory. And so this transfer and this transfer. They can happen asynchronously. The problem of course, when you have to asynchronous process is how you have them communicate. In particular, how does the sound card tell the CPU that more data is needed. This problem is solved by having the CPU issue a ...an IRQ, an interrupter request, to the CPU, and when the CPU receives this signal, it will put another batch of data in memory for the sound card to access. Now, if we were to interrupt the CPU every time we need a new sample, there would be too much overhead. So the idea is that the sound card will consume a chunk of samples already in memory, called a buffer. The sound card will notify when the buffer is completely used up and the CPU will fill a new buffer for the sound card to use. The key point here, since we're talking about real-time signal processing, is that the time it takes for the CPU to fill up the buffer needs to be strictly less than the time it takes for the sound card to consume that buffer. Of course the buffering operation introduces a delay, but this delay is a price So by processing the data in batches. We insulate data generation process against concurrent demands on the CPU. Let's look in more detail how this works. This is an example of double buffering. The buffer is this strip here, which is a set of contiguous location and RAM. Say L locations for L samples. The CPU will write newly computed samples into this buffer and the sound card will fetch the samples from this buffer for playing. There are two pointers into this buffer. There is a write pointer, used by the cpu to identifiy the next memory location into which write a new samples. And there's a read pointer used by the sound card, to identify the location from which the next sample has to be fetched. The pointers are initialized as such. We assume that the buffer has been filled with zeros in the beginning, and we position the read pointer at the beginning of the buffer, and the write pointer at the midpoint of the buffer. And now the dance begins. The sound card is started and the sound card will start fetching samples from the read buffer and move the read pointer in this direction. At the same time the CPU will start writing samples into the buffer and it will move the write pointer in this direction. Remember the key is that the CPU write samples faster than the sound card reads them. So as time progresses, you can see the read pointer advances, but the write pointer advances faster. And, after a few miliseconds, the CPU will have filled the second half the buffer and it will stop. At the same time, the sound card will keep reading and when it reads the entirety of the first half of the buffer It will signal to the CPU that this part of the buffer is now depleted before moving on to reading samples in the 2nd half of the buffer. When the CPU receives the IRQ, it knows it will have to start filling the first half of the buffer. This is a circular buffer, as we have seen in the previous module. So the pointer will be rolled around, and the CPU will start right in the samples here. So the reading process continues at the same pace as before, and now the filling process continues on the first half of the buffer, at a faster pace than the reading process. So again, the CPU will have filled the first half of the buffer. While the soundcard is still reading the second half. When the second half is depleted, a new RQ will signal to the CPU that the second half of the buffer needs to build again. And so on, and so forth. Double buffering uses a delay. That is equal to the system clock times half the length of the buffer. If the CPU is not fast enough, the read pointer of the sound card will trail over memory cells that contain invalid samples, old samples. And we will have a situation of Underflow. In practical application it's coming to you as multiple bufferings instead of simple double buffering. In this case the buffer is divided into sub-buffers, and now your cue will be emitted every time the read pointer trails one of the boundaries. The advantages that the CPU will be called more often, so the load on the CPU will be better distributed in time while still affording a reasonable under full protection. So, so far we looked at the output process of the sound card. What about the input? The input is really symmetrical. Instead of having the sound card fetch sample from memory, we will have the sound card put samples in memory and notify the cpu that a new buffer is available for processing. The cpu will work at its own pace, it will share the RAM with the sound card. The sound card will Both place input sample in memory. And fetch out the samples to play. And will signal to the cpu when either of these buffers are empty. Very quickly, the roles are reversed if this is the buffer in RAM. We now have that the write pointer belongs to the sound card, and the read pointer belongs to the CPU. We assume as per usual that the initial buffer is filled with zero. >> And we start the system. The CPU will start getting samples from the buffer, and the sound current will start writing samples to the buffer. The CPU will be faster and so it will deplete the buffer before the sound card has finished filling the other half of the buffer. When the sound card has finished filling the first half, it will trigger an IRQ and the CPU will know that this buffer is now ready for consumption. So, it will roll the pointer and start fetching the data and process it. And so on, and so forth. We now need to put it all together. So we will have an input buffer. And an output buffer. We will use a multiple buffer strategy, say, with 3 sub buffers. We will select the buffer sizes and the number of sub buffers to be the same both for input and output. And we will use the input IRQ to drive the processing. This is the input buffer and so the sound card will place samples into this buffer as they com from say, the microphone. This is the output buffer, and so the sound card will fetch samples from this buffer and play them through the loudspeaker. Assumed that the buffer is a 0 field, since the clock for the sound card is the same for both input and output, the input buffer will be filled at the same rate as the output buffer is emptied. After a few milliseconds the first chunk in the buffer will be filled. And conversely, the output buffer will be emptied. An IRQ will be launched, at which point the CPU will intervene, and start processing the input samples, say filtering them, and placing them. In the corresponding chunk of the output buffer for later use. Again, the CPU will have to perform this operation faster than the time it takes for the sound card to fill a chunk of the buffer. So here, for instance, when the sound card. Has filled say, a third of the next buffer, the CPU is done with the processing, It will mark the input subbuffer as used up, and the output subbuffer as ready for playing. The process continues. The sound card is still filling up the second chunk of the buffer. And when it's done, it will launch naricu/g, and the CPU will start processing this input chunk into this output chunk. The processing requires less time than the sound card needs to fill the third part of the, The buffer. And when the third part of the buffer is filled, an IRQ will be triggered, pointers will be rolled around, and the sound card will start outputting the first chunk of the buffer that was processed earlier. While the corresponding chunk of the input buffer will be filled with new input samples. And the process repeats. Total delay of the system is equivalent to t s times the number of samples in the buffer. Because of how we position the read and write pointers. We usually start the output process first, because it doesn't matter if the input trails a little bit and buffers can be collapsed. In other words, the computation can be done in place. Ok. So, how do we do this in practice? We get our BC. And 1 strategy is the low-level strategy. We study the sound card data sheet, each chip is different, we write code to program the sound car directly, this will require us to write some Particular values to some particular IO ports, we need to write an interrupt handler, and we need to write the code that handles the data. It's certainly interesting, but very time consuming, and not very portable. The high level strategy is to choose a good. Api that allows us to attack the functionality of the sound card without worrying about the details of it implementation and just write a callback function that handles the data when we are notified that a new Buffer is available. Let's look at an example of a typical callback prototype. We will use C in our next examples. And this callback will take simply three mandatory arguments. A pointer to the input buffer, a pointer to the output buffer and the length of said buffers. So for example, we could write a call back that simply. A cast to convert the generic pointers to the buffers and to the right data type, in this case, we choose to use floats. And then, a simple loop that, for each sample in the input buffer, calls a processing function and stores the result. [unknown] Into the output buffer. So now the question is, what do we put into the processing function? So let's expand that and let's define a generic processing gateway. If we're going to use filters, we will need to put in place buffers to store past samples of the input and of the output. So, we do this here on this left part of the panel. As we said before, we decide to use floats. So we defy a buffer of a certain length and we choose a power of 2 for convenience so that the circular buffer can be rolled around with a simple masking operation. So we define two buffers of equal length and two indices into the buffer. The processing function will take the current sample. It will store the sample into the input memory. It will call a function that we call effect because we're interested implenting guitar effects that will compute the current output. It will store the output in the memory for the output samples. And it will update the indices into the buffers with the usual circular strategy. So, we increment the pointer by 1. And then we roll around the pointer by binary masking. And the processing function will return the current output sample. That, as you remember, will be put in the output buffer of the sound card. So let's look at the simple effects that we can generate with this programming paradigm. An echo is a situation where you have a sound that bounces back and forth between 2 reflecting surfaces. And a simple simulation is given by the following impulse response. You have the original sound and then you have a first reflection. After N samples, and the 2nd reflection after 2N samples. So, you have equally spaced replicas of the signal, which we scale by different weights in order to simulate a decay in time. So, we have the original signal scaled by A, A is usually 1, a 1st reflection scaled by B And the 2nd reflection[UNKNOWN] by C. And then, we normalize the sum of the reflection by the sum of the coefficients in order to have unit gain. To implement this in our framework, all we need is a straight forward translation of the transfer function into C code. So here, we define the 3 coefficients that we used for the 3 Replicas of the signal, the normalizing coefficient, and the delay between replicas, which is here expressed as a function, of course, of the sample of frequency. The output value is simply the wave sum of three delayed replicas of. The input. Acoustically a simple echo sounds like this. First you will hear the original sound. And then you will hear the same sound processed by the simple echo. [music] The simple echo has two drawbacks with respect to a natural echo. The first one is that it has only a finite number of repetition, whereas in a natural situation it would have a theoretically infinite number of back-and-forth reflections. The second drawback is that each repetition is just... (End of transcription.) An exact replica of the original sound simply scaled in amplitude. Whereas the reflection process in a natural echo would introduce a low pass characteristic in each reflection. To obtain a better echo we can go back to our old friend the carpal strong algorithm. If you remember the carpal strong works by building equally spaced replicas of the input signal. To produce the output. The fact that the input signal is final support prevents overlap between the output copies, but if we input a standard sequence to the feedback loop, we will have overlap and[UNKNOWN] Time. So the feedback loop already takes care of the fact you want an infinite number of repetitions in the echo. In order to introduce the low pass characteristic of a natural echo, we simply introduce a low pass filter. In the feedback path. And the output at this point, can be expressed as the usual decay factor times the convolution of the output, and the input's response of the low pass delayed by capital M plus the input. If we choose a simple low pass such as the leaky integrator, we can actually write the constant coefficient difference equation, which turns out to be like that. And if we compute the impulse response numerically, we get something that Each repetition of the period of the echo gets smoother and smoother, which is really what you obtain if you apply a low-pass filter repeatedly. The implementation in our processing framework is, once again, very straight forward. We just need to convert the constant coefficient difference equation to see statements. We define the decay factor, the lambda factor for the leaky integrator. We define The normalizing factor that will preserve unit gain, and again, the echo delay as a function of the sampling frequency. And then we return the output as a linear combination of input and output samples. If we now play the same clip as before through this natural echo. It will hopefully sound better and more realistic. So here is a comparison between the two. The next linear time and variant effect that we will consider is reverb. Reverb is really the superimposition of very, very many echolike reflections that come from different directions and therefore have different delays... (End of transcription.) And different type of attenuation is associated to them. Because of the complexity of the impulse response associated to reverb, precise simulations are very costly in terms of computational power. But if we want to just approximate the reverb as would be generated by a very small environment with highly reflective surfaces You can get a way, by using a simple allpass filter as shown here. Remember the allpass filter is a filter whose magnitude response is unit over the entire minus pi to pi frequency axis. By appropriately choosing the parameters alpha and capital N in the allpass, we can obtain a phase response with the staircase like characteristics shown here. This phase response will introduce different delays to different frequency regions. Thereby simulating a multi-faceted reflective surface. The implementation of the all pass filter is again extremely straightforward. We only have two parameters to choose. And then we just implement the cutsom coefficient difference equation in c. We can now try to apply the reverb to a little guitar snippet. We'll play the original and then the processed signal. Let's now look at some very common guitar effects that do not belong to the class of linear time invariant filters. This is to give you an idea that there is a whole world out there beyond linear processing that is very, very interesting. The 1st type of effect that we will consider is the distortion. Simple distortion can be achieved by truncating the scaled version. Of the input signal and bringing the signal back to the original amplitude. Tremelo is an effect achieved by multiplying the input signal by a slowly varying sinusoidal component. This will change the volume periodically in time. Flanger is obtained by summing a signal to a delayed version of itself and by varying the delay in time. Usually according to a sinusoidal pattern as well. When you sum a signal to a delayed version of itself, you obtain some spectral nulls and frequency. By changing the delay, the spectrum nulls will change their position in time, and this will give rise to a variety of effects, from say, a robotic sound to an artificial chorus. From the point of view of implementation, the fact that these effects are nonlinear doesn't really complexify the associated code. And this is really one of the strong points of digital signal processing. We can easily move beyond the standard filtering paradigm without incurring much complexity. So distortion simply takes the input signal, operates a hard limit, and rescales the output. And it sounds like this. [music] Tremolo requires you to define the speed of the modulated sinusoid and then the output is obtained simply by multiplying the input times an appropriately scaled sinusoid. Tremolo sounds like this. Finally, with flanger, we need to define the maximum delay, the frequency at which this delay will oscillate around its mean value, and then we compute the instantaneous delay. By converting to an integer, the product of the maximum delay times the oscillating component. And with some, the two delayed versions of the input. [music].