So, so far we discussed SVD and we will conclude our discussion of singular value decomposition by looking at an example of its usage, and then talk a bit more broadly about the method. So, here is the idea. Let's think about the following problem. Imagine we want to identify all the users who like our Matrix movie. So, the idea is we have this matri matrix now of values of users to movies. And now, we would like to identify all the users who liked the movie Matrix. And what we would like to learn from this task in a sense is that given that we saw that there are people who like SciFi movies, so maybe we would like to kind of find a person who didn't even see the movie Matrix yet, but we may want be able to say yes, but given that they like these other SciFi movies. The mal, they may also like the movie Matrix. And the question is, how can we do this using SVD? The answer to this question is that basically we want to map, map our query point into the concept space. We are given our data users to movies matrix A, right? We did, we do the SVD of it, so here's the SVD. And what we want to do now, is we are given our query point. Our query point q is simply we say, let's find all the users that like the movie Matrix. So, let's create, in some sense, this query, this query user, this artificial user that likes the. A movie matrix, and the idea is, we want to find other users who are close to this given user in the, in the concept space. So, what we'll do we have, we have our movies space, we have our data point q here, our query point. And we want to project it into our concept space. The way, the way we do this is that basically we simply do the inner product of our query point with each concept vector in vector v. Because vector V is movies to concepts vector right, so if we go if we go do this, y is taking the inner product a good idea, because of example q times the first singular vector will simply take, take our position of the point q. And tell and will tell it, tell us its location along the, the, the axis of the first singular vector. The second singular vector is orthogonal to the first one, so here it is the V2. And when I multiply q times V2. We will basically now get the projection of the data point and its position on the second singular vector. So, that's basically what will happen. So, if we, if you do our, our projection. So, we take our vector q multiplied with matrixVv we do the thing, and here is what we obtained. So, we obtained that this, for example, this particular user, now we are in this concept space, in this two, two-dimensional concept space, where the first column of V is SciFi, and the second column of V was romance. And once we do the inner product, we basically see that a lot, that, our query point in some sense corresponds heavily to the SciFi concept, and very low, has a very low coordinate value along the romance concept. So, this is now how we mo, do the query and kind of map it into the concept space. So, now for example, imagine I have some user, some user d. That we don't know what they think or they haven't told us anything about what they think about the movie matrix. But they didn't tell us they really like movies, Alien they haven't said anything. So, if we take this user and again multiply them by our vector V and basically move them to the concept space. Here are the coordinates of the position of that user in the concept space. So, what is, what is a good thing that happens? So, for example, if I now compare the positions of the original user and the query in the original space, and, the, the, the locations in the concert space, I find the following. So, the similarity between q and d in our original space is zero, right, in the sense that for Alien and Serenity, our query doesn't. doesn't want them. For our movie matrix, the user didn't tell us anything, so there is no similarity between these two vectors. But if I go back to my concept space, here I see that both of these data points, or both of these users, q and d, they actually share high values on the SciFi concept. And they share low values on the romance concept. So, in some sense, I may be able to identify or I may be able to put together that q and d, are actually close together in our space even though in the raw data representation they don't share any coordinates together. So, in some sense even though q and d have zero ratings in common, we are able to identify that they are similar, because SVD was a SVD was able to identify that kind of people who liked Alien and Serenity also liked the Matrix movie. So, this basically is how we can make use of Singular Value Decomposition. What I want to do now, now very briefly is to relate Singular Value Decomposition to another. Type of, decomposition of a matrix, that is called Eigen value decomposition. so, first, I will you what is the relationship between singular value decomposition and Eigen-decomposition, and then we will con, conclude. So, what we know so far is that SVD is, given a matrix, we represent it as a product of three matrices, using my retranspose. What is Eigen value decomposition? Eigen value decomposition is kind of, more constrained. It says, given sub matrix A, I want to represent it as a matrix x times the ma, ma, matrix capital lambda times x transposed again. So, here I only have kind of product of two matrixes, if you like, lambda and x. So, for Eigen-decomposition to, to even exist what we have to do is first A has to be symmetric which means that the values above the diagonal have to be the same as the values below the diagonal. While, for example, in SVD we did not have this constraint. And then ce, both in svd and Eigen value decomposition, all the matrices are columnar phenomena, which means the columns are orthonormal to each other and have unit length. And in both cases, you sigma and lambda are diagonal matrices. So, now the question is, what is the corres, corresponding, correspondence between singular value decomposition and Eigen-decomposition? So, let's consider this simple case. Let's consider what is A times A transposed? So, we know that we can take the matrix A and perform singular value decomposition of it. So, let's do that. So, what is A times A transposed is the singular val, value decomposition of A times the singular value decomposition of A transpose. Which is the same as singular value decomposition of A and then trans, trans, transposing that. So, now let's start thinking what do we what do we get next? All right, what we get next is that because the, the multiplication is cumulative we can kind of reword reword the terms. Right? So, for example we can take the original expression and now just reword the terms of, or the order of multiplication. What we notice now is that we get, we get a multiplication of V transposed times V. And given that our matrices are orthonormal, this means that V, a matrix multiplied with itself, gives us an identity matrix. An identity matrix is simply a matrix that has zeros, off the diagonal and it has values of 1 on the diagonal. So, what this means is that it's basically an identity matrix, right. So, what this means is that we can take. AA transpose and transform is down if you take each SVD and see what happens, it turns out to become U times sigma, sigma transpose U transpose, okay? similarly, the same thing happens or a similar thing happens if I ask what is the singular value decomposition of A of matrix a times a transpose. I do, I do the same, the same trick as before, and here now, I obtain that this equals V times sigma times sigma transpose times V transpose. One thing that I have to remember is that sigma is a diagonal matrix. So, in some sense sigma times sigma transposed is nothing else than another diagonal matrix that has the squares of the values on the diagonal, right? So, what do we learn from this? Is the following. So, if I take the A times A transposed, which is a symmetric matrix. I do a singular value decomposition of it. What I end up with, is an expression like, like this which basically means that, in this case, u is a set of I can think of u as a set of item vectors. Right? So, as a part of the item decomposition. And I can think of Sigma times Sigma transposed as a set of item values. So, what this basically means is that if I have a matrix I can do singular value decomposition of it and from singular value decomposition of it I can do the Eigen value decomposition. Where the relationship between Eigen values and the singular values is that, singular values squared are the Eigen values of the corresponding matrix. So, that's the co, that's the connection between the Eigen value and the singular value decomposition. So, to to finish talking about singular value decomposition, here is, here is kind of the overview. So, what is good about singular value decomposition is that it gives, it gives us the optimal low rank approximation right, in terms of the Frobenius norm. So, it means that if I allow myself to take my data and represent it using a small number of dimensions, then SVD will be able to identify the best possible number of dimensions. That basically give us the best possible prejection of the data into some small dimensional space in such a way that if we go from this small dimensional space back to the original high dimensional space, the sum of the squares of the reconstructions error, errors, will be as small as possible. So, that's great. What is problematic with SVD are two things. The first one is the interpretation problem. What this means is that the singular vectors specifies some, a linear combination of input columns or rows. Which means that many times singular vectors are very hard to interpret. When I say singular vectors are hard to interpret in our cases of movies to users matrix, I was able to interpret the first singular vector to correspond to the SciFi movies, and the second one to the Romans movies. Many times, that is hard to do, and the second big drawback of singular valid decomposition is what is called lack of sparsity. What this means is that, the input matrix A, is often very sparse which means it, it is full of zeros, and only has a few non zero elements in it. But when you, when we do singular value, the composition,. The matrices U and V transpose, they will be dense. What I mean by that is all, basically all the values in this matrix will be non zero. So, many times, even though U and V have a very small number of columns or rows, so in some sense in, in terms of the row column or row, row size, they are much smaller than the matrix say in terms of the data size, maybe bigger because a has very few non zero elements. Or values and then matrices u and v have a large number of non zero elements so what we will do next is we will look at the method that is much easier to compute much faster to compute then the singular value decomposition. And also main maintains the sparsity of rows and columns of U and V and this is what we are going to look at next.