In this brief video we're going to talk about alignment. A very low level but very important detail of machine architecture that every compiler writer needs to be aware of. First, let's review a few properties of Contemporary machines. Currently, most modern machines are either 32 or 64 bit meaning you have the 32 or 64 bits in a word and the word is actually subdivided into smaller units. We would say that there are eight bits in a bye and then four or eight byes in word depending whether it's a 32 or 64 bit machine. And other important property is that machines can be either byte or word addressable. Meaning that in the native language of the machine in machine code it may be possible to either name only entire words or it may be possible to reference memory at the granule area of individual bytes. They say that data is word aligned if it begins at a word boundary. So if we think about. Data in memory or the organization in the memory and is laid out into bytes. And let's say. That this is a 32-bit machines so that four bytes make a word and one word begins here and the next word begins here and if data is allocated on a word boundary, say, it needs more bytes then that would be a word a line a piece of data. If a piece of data begins in the middle of the word, so let's say for example that begins here, and we have some data that's allocated here, this data is not [inaudible] doesn't begin on a word boundary And [inaudible]. Property or the important issue is that most machines have some alignment restrictions. So these restrictions come in one of two forms. So, on some machines, if the data is not properly aligned, meaning that you tried to reference data that isn't aligned the way the machines requires, then the machine may just fail to execute that instruction. Your program may hang or even the machine may hang and it's, but, the important thing is that program will not execute correctly. So there's a, it's incorrect to not have the data aligned properly. Now, there are other machines that well, actually al low you to put the data anywhere you like but at a significantly cause And maybe that accessing data that is aligned in word boundaries is cheaper than accessing that's on non-word boundaries And these performance penalties Are often dramatic so it can easily be ten times lower to access missile line data than to access data that has the alignment favored by that particular machine. So let's take a look at an example where data alignment issue tend to come up. One of the most common situations where we have to worry about the alignment is in the allocation of strings. So let's say we have this string, the string Hello and then we want to put it in memory. So let me draw our memory as a linear sequence of bytes so I'll mark out some bytes here. And let's assume this is a 32-bit machine so let me make the word boundaries a little bit heavier boundaries. So, one, two, three, four. Okay. So, there are the, the word boundaries And now let's say there were we are trying to have aligned data, a word aligned data and so allocate this string beginning in the word boundary. So, the each character will go on the first byte when e, then l, then l, then o. And now, we may have terminating null depending on how strings are implemented. And let's assume that we do. And this is fine placement of the strings extremely begins in the word boundary and. That assess by presumably any alignment restrictions of the machine and now the question is where does the next data item go? So we could begin the next data item right in the next available byte and that would be good if we are very concerned about not wasting memory. But, I noticed that, that data item will then be were aligned. We may either run into correctness or performance problems if the machine has restrictions on the alignment. So, the simple solution here is to simply skip to the next word boundary and allocate the next data item whenever it is on the next word beginning at the next word boundary. And what happens to this two bytes here, well these bytes are just junks. T hey're not used at all, they never reference by the program. It doesn't matter what they're value is because the program should never refer to them. It's just unused memory. And note that if we didn't have the terminating zero then there would be the terminating, no character then and then would be three unused bytes after the [inaudible]. So to summarize this is the general strategy for dealing with alignment when you have alignment restrictions. Data begins on the boundaries, typically word boundaries that are required and if the particular data that you're allocating has a none integral length. Meaning that it doesn't end directly on the next required boundary and you just skip over whenever bytes are in between to get the data, the next data that's going to be allocated on the correct boundary.