Continued from Question 11.
In this question, we will use the update rule we just derived to implement a neuron that will learn from of two dimensional data that is given in the following file:
c10p1.mat (This file is provided as part of the
exercises from the Dayan and Abbott textbook recommended for the course).
c10p1.mat contains 100
(x,y) data points.
Move c10p1.mat to your Matlab or Octave directory and use the following command to load the data (note that you must include the '-ascii' option for the file to load correctly):
load('-ascii', 'c10p1.mat')
You may plot the data points contained in c10p1.mat using the following command:
scatter(c10p1(:,1), c10p1(:,2))
The equivalent python pickle files are:
c10p1.pickle (Python 2.7)c10p1.pickle (Python 3.4)
and can be loaded in the usual way:
import pickle
with open('c10p1.pickle', 'rb') as f:
data = pickle.load(f)
Assume our neuron receives as input the two dimensional data provided in c10p1, but with the mean of the data subtracted from each data point (the mean of all
x values should be subtracted from every x value and the mean of all
y values should be subtracted from every
y value). You should perform this zero-mean centering step and then display the points again to verify that the data cloud is now centered around
(0,0).
Implement the update rule derived in the previous question in Matlab or Octave. Let
η=1,
α=1, and
Δt=0.01. Start with a random vector as
w0. In each update iteration, feed in a data point
u=(x,y) from c10p1. If you've reached the last data point in c10p1, go back to the first one and repeat.
Typically, you would keep updating
w until the change in
w, given by norm(w(t+1) - w(t)), is negligibile (i.e., below an arbitrary small positive threshold), indicating that
w has converged. However, since you are implementing this as an online learning algorithm, you may prematurely detect convergence using this method. Instead, you may just run the algorithm for 100,000 iterations.
Run your code multiple times. You should find that
w converges to two very different vectors. Why does this happen?
Hint: Consider the eigenvectors of the correlation matrix of the mean-centered data. (The correlation matrix of a data matrix
X, where rows indicate separate samples, is
XTX/N, where
N is the number of samples. You can calculate its eigenvalues using eig().) If the data is mean-centered, the correlation matrix will be the same as the covariance matrix.