Quiz 5: Networks and Learning Help Center

Learn more

Warning: The hard deadline has passed. You can attempt it, but you will not get credit for it. You are welcome to try it as a learning exercise.

This week's quiz will test your knowledge of networks and learning based on lectures from Week 6 and Week 7. Some of the questions require you to use Matlab or Octave.

Note: For this quiz, we have written several variations of some questions, and these questions can change slightly if you attempt the quiz multiple times. If you re-attempt the quiz, please make sure to adjust your responses and code accordingly.

Good luck!

Question 1

Let's design some feedforward networks that can do some basic operations on their inputs. This could mean lowering their intensity, looking for strong changes, or one of many other possibilities. One nice way to build intuition for this sort of processing is to think of these networks as operating on images. Even though our networks will operate over only 5 pixels of image data, we can still build the same basic operations that we would for a regular image. For all of these questions, we will start with the following image as input:



Suppose we processed the image and it looked like this:



Which of the following weight matrices W would give us a feedforward network most closely approximating this image processing operation?

Question 2

In lecture 6.2, we encountered a process of conceptual abstraction, taking us from modeling individual neurons to modeling whole networks. By the middle of the lecture, we had abstracted away many of the interesting time dynamics of feedforward neural networks and arrived at a simple equation:

vss=F(Wu)

We have to make a number of assumptions to get to this equation. Necessarily we lose some interesting information when we make these assumptions. Which of the following details do you think we have lost between the beginning of lecture 6.2 and the point where we first see this equation (roughly 10 min in).

Question 3

Suppose that we had a linear recurrent network of 5 input nodes and 5 output nodes. Let us say that our network's weight matrix W is:

W=⎡⎣⎢⎢⎢⎢0.60.10.10.10.10.10.60.10.10.10.10.10.60.10.10.10.10.10.60.10.10.10.10.10.6⎤⎦⎥⎥⎥⎥

Suppose that we have a static input vector u:

u=⎡⎣⎢⎢⎢⎢0.60.50.60.20.1⎤⎦⎥⎥⎥⎥

Finally, suppose that we have a recurrent weight matrix M:

M=⎡⎣⎢⎢⎢⎢−0.500.50.500−0.500.50.50.50−0.50.00.50.50.50−0.5000.50.50−0.5⎤⎦⎥⎥⎥⎥

Which of the following is the steady state output vss of the network?
(Hint: See the lecture on recurrent networks from Week 6, and consider writing some Octave or Matlab code to handle the eigenvectors/values (you may use the "eig" function))

Question 4

Suppose we have a linear feedforward network with two input nodes and one output node. Let's say that we are learning our weight vector w and that we are using the Hebb rule.

Suppose the input correlation matrix Q is:

Q=[0.20.10.10.3]

If we allow learning to go on for a long period of time, which of these could be a final weight vector w?
(Hint: See lecture 7.1 and use Matlab's "eig" command)

Question 5

What does the weight vector we found in Question 4 tell us?

Question 6

The "learning rate" as mentioned in lecture 7.2 is used in many different contexts to denote the sensitivity of a parameter estimate to new data during online learning. Check all of the following which are true:

Question 7

In lectures 7.2 and 7.3, we saw multiple algorithms which use a two-step process for parameter estimation. EM is one such algorithm, consisting of the E and M steps. In each step, we iteratively update our estimates of parameters. Why do we need to alternate between the two steps? What justifies this approach?
(Hint: You may want to google "EM algorithm")

Question 8

The next three questions utilize the following code to model an integrate-and-fire neuron receiving input spikes through an alpha synapse: alpha_neuron.m (Python: alpha_neuron.py).

The parameter “tpeak” controls when the alpha function peaks after an input spike occurs (and hence how long the effects of an input spike linger on in the postsynaptic neuron). “tpeak” for excitatory synapses in the brain may vary from 0.5 ms (AMPA or non-NMDA) to 40 ms (NMDA synapse).

Vary the value of tpeak from 0.5 ms to 10 ms in steps of 0.5 ms and observe how this influences the output of the neuron for the fixed input spike train used in this code. Plot the output spike count as a function of tpeak for the given input spike train.

Which of the following answers best describes the relationship between the value of tpeak and the firing rate of the neuron?

Question 9

Continued from Question 8:

Which of the following explanations best explains how the value of tpeak influences the firing rate of the neuron?

Question 10

Continued from Question 8:

How would you turn this synapse into an inhibitory synapse?

Question 11

In the next five questions, we'll implement Oja's Hebb rule for a single neuron and explore some of its properties.

Recall that Oja's rule is:
τwdwdt=vu−αv2w, where v=u⋅w.

We will implement the discrete time version of Oja's rule in Matlab or Octave. To do this, first rewrite Oja's rule using discrete time rather than continuous time. Which of the following equations represents the discrete time version of Oja's rule? Let η=1τw.

Question 12

Continued from Question 11.

Now, in order to use Oja's rule, we need to translate the discrete time version into an update rule. Which of the following equations represents the discrete update equation for Oja's rule?

Question 13

Continued from Question 11.

In this question, we will use the update rule we just derived to implement a neuron that will learn from of two dimensional data that is given in the following file:
c10p1.mat
(This file is provided as part of the exercises from the Dayan and Abbott textbook recommended for the course).

c10p1.mat contains 100 (x,y) data points.

Move c10p1.mat to your Matlab or Octave directory and use the following command to load the data (note that you must include the '-ascii' option for the file to load correctly):
load('-ascii', 'c10p1.mat')

You may plot the data points contained in c10p1.mat using the following command:
scatter(c10p1(:,1), c10p1(:,2))

The equivalent python pickle files are:
c10p1.pickle (Python 2.7)
c10p1.pickle (Python 3.4)

and can be loaded in the usual way:
import pickle
with open('c10p1.pickle', 'rb') as f:
    data = pickle.load(f)
Assume our neuron receives as input the two dimensional data provided in c10p1, but with the mean of the data subtracted from each data point (the mean of all x values should be subtracted from every x value and the mean of all y values should be subtracted from every y value). You should perform this zero-mean centering step and then display the points again to verify that the data cloud is now centered around (0,0).

Implement the update rule derived in the previous question in Matlab or Octave. Let η=1, α=1, and Δt=0.01. Start with a random vector as w0. In each update iteration, feed in a data point u=(x,y) from c10p1. If you've reached the last data point in c10p1, go back to the first one and repeat.

Typically, you would keep updating w until the change in w, given by norm(w(t+1) - w(t)), is negligibile (i.e., below an arbitrary small positive threshold), indicating that w has converged. However, since you are implementing this as an online learning algorithm, you may prematurely detect convergence using this method. Instead, you may just run the algorithm for 100,000 iterations.

Run your code multiple times. You should find that w converges to two very different vectors. Why does this happen?

Hint: Consider the eigenvectors of the correlation matrix of the mean-centered data. (The correlation matrix of a data matrix X, where rows indicate separate samples, is XTX/N, where N is the number of samples. You can calculate its eigenvalues using eig().) If the data is mean-centered, the correlation matrix will be the same as the covariance matrix.

Question 14

Continued from Question 11.

What happens when the data is not zero-mean centered before the learning process?

In order to more fully explore the behavior of the Oja's rule when the data isn't mean centered, you should adjust the mean of the data a few times and observe the behavior of the learning rule. You can adjust the mean of the data by adding a constant to every x component of the data and a different constant to every y component of the data.

Question 15

Continued from Question 11.

What happens when the pure Hebb rule is used instead of Oja's rule? You can explore what happens by removing the subtractive term −αv2w in your code and running the code.
    
You cannot submit your work until you agree to the Honor Code. Thanks!