So far we implicitly assumed that our data is linearly separable. What this means is, we kind of assumed that it's always possible to find the line that will perfectly separate positive and negative training examples. However most of real data sets, they are noisy which means that there is no such sep, separating hyperplane. So the question is what happens to our, for, formulation when we have data that cannot be nicely separated. For example, imagine my data set here. Where I have these data points that are kind of are on the wrong side of the, of the bound, the decision boundary. And there is actually no linear decision boundary here that would allow me. To, to draw a line and put all the pluses on one side, and all the minuses on the other side. So, when we are dealing with such, such data. Where finding nice linear separator is impossible and this is basically this happens in every data part, data set. What we have to do is, we have to change our formulation a bit and introduce a penalty. So, the idea will be the following. What we want to do now, is we still want to maximize the margin. This is the first part of our objective function. But what we want to do is we want to have some parameter, and we will call this parameter C, and plus the number of mistakes. Right? So, what we are doing right now is basically saying, we want to find w that has good margin, while also makes a small number of mistakes. Right? So, the idea, in a sense now give, minimizing w gives us the, gives us the line that has high margin. While the second part of the optimization problem, the, the one on the right basically wants to control for the number of mistakes. Right? And the idea here is that we will have the value of C and set it. And the goal is to find the separating hyperplane to find the line, that both has good margin and makes a small number of mistakes. And now, of course, the question is, how do we penalize mistakes? Because not all mistakes are of same severity. And the idea is that not mistakes are equally bad. Which means we will be using margin, in order to penalize them. So how do we use margin to penalize mistakes? We introduce this notion of slack variables. And the way we think of slack variables are basically these additional constraints, or these additional penalties. That we get for misclassifying a data point. So the idea is that we have our separating plane. And then the, the value of slack variable, or the value of penalty will simply be wha, what is the distance from the other side of the sep, or the margin. To the, to the data point itself. Right? So in this case, psi the value of psi is this much. For example for the mis-specification of this data point plus, the value of psi is all the way from the other side right? This is kind of the, how much we are mis-classifying. Because it would require us kind of to move that data point plus to the other side of the margin. If you, if we would want to make it be classified correctly. So what this means is, now we are arriving to our new optimization problem, right? We still say, okay, what is our goal? Our goal is to find w and b, and we want to also find the values of the slack variables psi. Such that the norm of w is small, which means the margin is large. Plus the, the sum of the span of this psi is as small as possible. While what, what we also require, we also require the confidence in our classification is at least 1. And if it's if it's not 1, then we have to subtract the value of psi. Right, so this is basically whenever our correct, example is correct, correctly classified, the our confidence will be greater than 1. And we can, in that case, will be able to set the value of psi to 0. Otherwise, if that is not the case, we will have to set the value of psi to some nonzero value. Which means we will occur some penalty in the optimization problem. And the idea here is basically that if we take our data point Xi and it is on the wrong side of the classification margin. Then we incur some penalty for misclassifying it. And this data, this optimization that we set it so far, this is called the SVM with soft constraints. Why soft constraints? Because now we can also allow for misclassifications. So, one more thing that would be good to get some intuition about is what is the role of this parameter C? We call this parameter the slack penalty. Why do we call it the slack penalty? Is because it controls between the size of the cost of the margin. How much are we wishing to make the margin big, big? And how much, are we penalizing our misclassification mistakes? So the way we can think of C is the following. If we set C to be infinite, right, what this basically means is that we only want to find w that separates the data. So for example, in our case, if I have a data set here and I would set C to be very big, then this is the decision boundary we would find, right. It's a decision boundary, that nicely separates the data. For example, if you would set C equal 0, right? Which would basically mean that we don't really care about misclassifications, but we just want to make our, our W to be as, as short as small as possible. Then, basically, what, what this would do, it would ignore the data, and the whole decision morally would just be something that goes through the coordinate origin. So, it could be this [INAUDIBLE] line that I show here. But, however, if we choose a good value of C, then we are nicely trading off between our line nicely separating the data, so having large margin, while also not making too many mistakes. And for a good or appropriate value of C, this is the line we would like to find, right. We still have a relatively nice separation between pluses and minuses, while making one small mistake. So, having discussed the value of the slack penalty in the formulation of the support vector machine, here is now what we call the support vector machine optimization problem in it, in its natural form. So the way we can think about it is the following. Our goal is to solve the following optimization problem, where we want to find b and w, such that the. 1/2 square of the, of the, of the square of the normal w, plus the slack penalty times our misclassification costs. The whole thing is minimized. What is, what is this doing, the way we are thinking about this, we are thinking of the first part of the optimization problem as maximizing the margin. Right, we want. The length of w to be as small as possible. And we think of C as a slack penalty. Which is something is something we have to kind of set by hand and it tells us how much are we trading off between fitting the data and making the margin large. And then the the last part is we call it empirical loss. Right. Because this is saying how well are we fitting the data. Right. So the left part of the equation is trying to maximize the margin. Find a good separator and the second part is to, trying to say let's try to fit the data as well as possible and the cost of how well are we fitting the data is called the loss. On how we can now think about machine learning is that basically machine learning is trying to trade off between finding a, a good separation between the two classes while also miminzing the loss. And in particular the loss that we have written here goes under the name of the hinge loss. So we can think of support vector machines to be using or minimizing the hinge loss. The reason why we call it the hinge loss is the following. What we ould really like to do is, is, the idea is that if we have our classification. And on the y axis, we plot the penalty. The idea would be that if we misclassify, we obtain a penalty of one, right, if misclassification means that we predicted one class, and the, the true class was, was, of the other sign. So the. Product of the two signs is negative. While if we made the correct classification, we would like to obtain 0, meaning no penalty. So an ideal 0/1 loss would be, you obtain penalty of 1 if misclassify, and obtain penalty of 0 if we classify correctly. What is the penalty that support vector machine is using, is called the hinge loss. The reason we, we call it the hinge loss is, because there is this hinge at one. Which basically means, if we are classifying the point correctly, and the point is away from the margin, it's basically away from the decision boundary for a least value of 1, then we obtain the class the the cost for the loss of 0. However if the mid, if the point is inside the margin or inside the classification boundary so can still be classified correctly but is too close to the boundary. Or is actually on the wrong side of the boundary then we are incurring the penalty and its penalty is proportional to how far away is our point from the from this decision boundary. So this is called a Hinge Loss, and support vector machine is exactly optimising this hinge loss in the, in the lost part of the term.