1
00:00:00,880 --> 00:00:06,055
Now, that we've learned the two concepts
of node degree and connected components,

2
00:00:06,055 --> 00:00:11,155
we can start to apply them immediately to
the dining-table's partner data set and

3
00:00:11,155 --> 00:00:15,696
see what insights we can get out of it.
First, let's look at degree.

4
00:00:15,696 --> 00:00:19,340
Now we have some choices.
This is a directed network, so we can look

5
00:00:19,340 --> 00:00:22,259
at InDegree, OutDegree, or just undirected
degree.

6
00:00:22,260 --> 00:00:27,969
Now, why doesn't OutDegree make sense?
Well, every girl is naming two other girls

7
00:00:27,969 --> 00:00:32,896
as their as her first and second choice.
Therefore, all the nodes are going to have

8
00:00:32,896 --> 00:00:34,568
the same degree.
Same OutDegree.

9
00:00:34,568 --> 00:00:38,602
So we don't want OutDegree.
Indegree, on the other hand, is a lot more

10
00:00:38,602 --> 00:00:41,619
interesting, because it's a reflection of
popularity.

11
00:00:41,620 --> 00:00:46,810
If lots of other girls want to sit with
you at the same table, well, that says

12
00:00:46,810 --> 00:00:50,917
something about how desirable you are in
this girl storm.

13
00:00:50,918 --> 00:00:59,642
So let's look at that.
The degree is calculated automatically, so

14
00:00:59,642 --> 00:01:06,345
you can just go straight to Ranking >
Nodes > InDegree.

15
00:01:06,345 --> 00:01:11,353
Now, we could color the nodes by InDegree,
but we're going to save color for the

16
00:01:11,353 --> 00:01:15,535
connected components.
So let's instead look at size, it's this

17
00:01:15,535 --> 00:01:19,739
little diamond shape here, so I'm going to
click on it.

18
00:01:19,740 --> 00:01:24,840
And we can chose the Min and Max size.
I'm going to go since these nodes are

19
00:01:24,840 --> 00:01:28,138
sized 5.
I'm going to set the Min size a little bit

20
00:01:28,138 --> 00:01:31,167
lower at 2 and the Max size maybe twice
that.

21
00:01:31,168 --> 00:01:36,226
There is this option called Spline.
And this just says, how does this size

22
00:01:36,226 --> 00:01:40,935
vary as this quantity of interest varies
in this case in degrees?

23
00:01:40,936 --> 00:01:47,642
Since the inequality in popularity isn't
that large and it's a relatively small

24
00:01:47,642 --> 00:01:53,955
network, this linear relationship is okay.
However, if for example, you were looking

25
00:01:53,955 --> 00:01:57,052
at a massive network a directed one such
as the web.

26
00:01:57,053 --> 00:02:01,170
And, well, you wouldn't be looking at the
whole web, but say a subset of the web.

27
00:02:01,170 --> 00:02:07,697
And there could be one page that has
100,000 in links, and many other pages

28
00:02:07,697 --> 00:02:13,891
that have one in-link or no in-links.
And you wouldn't want a node that's

29
00:02:13,891 --> 00:02:19,448
100,000 times bigger than other nodes.
It would just obscure your whole

30
00:02:19,448 --> 00:02:24,327
visualizations.
In that case, you may want to do something

31
00:02:24,327 --> 00:02:29,530
like a law of linear relationship,
something like this.

32
00:02:29,530 --> 00:02:32,600
But for our purposes here, we're just
going to go linear.

33
00:02:32,600 --> 00:02:38,417
So, let's apply that, and voila.
Instantly, we can see that some girls are

34
00:02:38,417 --> 00:02:43,032
more popular than others.
Marion and Eva look like they're most

35
00:02:43,032 --> 00:02:47,503
popular and Hilda Doesn't seem to be doing
too poorly either.

36
00:02:47,504 --> 00:02:55,007
Now this is all visual.
And if we want to actually look up the

37
00:02:55,007 --> 00:02:58,030
numbers.
What is the InDegree of each node?

38
00:02:58,030 --> 00:03:03,336
If we simply go to Data Laboratory.
Unfortunately, it doesn't populate the

39
00:03:03,336 --> 00:03:07,134
degree automatically.
So we're going to go back to Overview.

40
00:03:07,135 --> 00:03:11,437
And we're going to ask it to calculate the
Average degree.

41
00:03:11,438 --> 00:03:17,885
Now, the Average degree is 2, which we
sort of knew, because we have 52 edges and

42
00:03:17,885 --> 00:03:21,937
2, 26 nodes, so you can divide 1 by the
other.

43
00:03:21,937 --> 00:03:27,280
So nothing surprising there.
It has given us the Degree Distribution.

44
00:03:27,280 --> 00:03:30,478
We're actually interested in the In-Degree
distribution.

45
00:03:30,479 --> 00:03:39,330
And here, you can see that, for example,
two girls have In-Degree 6, but I don't

46
00:03:39,330 --> 00:03:45,623
know, nine girls have In-Degree 2.
So that's interesting in and of itself.

47
00:03:45,624 --> 00:03:51,807
However, what is super nice about this is
that we can go back to Data Lab.

48
00:03:51,807 --> 00:03:58,845
And now, we actually have the In-Degree
and clicking on top here of this sorts

49
00:03:58,845 --> 00:04:03,140
lowest to highest.
Let's click again and we have the two

50
00:04:03,140 --> 00:04:07,046
girls who have In-Degree 6 who are Marion
and Eva.

51
00:04:07,047 --> 00:04:11,958
So that's a nice way of just getting at
the, numbers.

52
00:04:11,958 --> 00:04:16,997
And, also, if you're interested in the
Debris Distribution you can have a quick

53
00:04:16,997 --> 00:04:21,318
look like this.
Of course, you can export this spreadsheet

54
00:04:21,318 --> 00:04:25,304
and analyze it further as well.
So back to the Overview.

55
00:04:25,305 --> 00:04:30,511
What we have is a picture where some girls
are more popular than others, but we don't

56
00:04:30,511 --> 00:04:34,858
really know whether there certain cliques
or how information might flow.

57
00:04:34,858 --> 00:04:41,186
So, I'm going to start out by asking a
question, you now, if someone had some

58
00:04:41,186 --> 00:04:46,470
piece of gossip, one particular girl.
Where could that gossip go?

59
00:04:46,470 --> 00:04:53,446
Or if you had a set of girls, would they
all be able to hear gossip from everyone

60
00:04:53,446 --> 00:04:56,536
else?
And this is precisely the question of who

61
00:04:56,536 --> 00:05:01,124
is in the same strongly connected
component, because if you're in the same

62
00:05:01,124 --> 00:05:06,230
strongly connected component with someone
else, it means that anything that they

63
00:05:06,230 --> 00:05:09,177
find out can potentially make its way to
you.

64
00:05:09,178 --> 00:05:12,550
So let's find the strongly connected
components.

65
00:05:12,550 --> 00:05:16,400
I'm going to go back to the Statistics
tab.

66
00:05:16,400 --> 00:05:19,999
If you have Filters instead, just make
sure you click on Statistics.

67
00:05:20,000 --> 00:05:23,120
And I'm going to click on Connected
Components.

68
00:05:23,120 --> 00:05:26,458
I'm going to say Run.
And it's going to ask me if I want the

69
00:05:26,458 --> 00:05:31,926
strongly connected components.
Which I can get in a Directed graph, or

70
00:05:31,926 --> 00:05:33,940
not.
Yes, I do in fact want the strongly

71
00:05:33,940 --> 00:05:37,148
connected components.
Because, if you imagine the flow of

72
00:05:37,148 --> 00:05:41,698
information, if one girl says she would
like to sit with someone else, but that

73
00:05:41,698 --> 00:05:44,374
one, that person doesn't want to sit with
her.

74
00:05:44,374 --> 00:05:48,077
Well, they may also not be willing to
share information with her.

75
00:05:48,078 --> 00:05:53,047
So, we really want to take into account
the directionality of the edges.

76
00:05:53,048 --> 00:05:56,542
So I'm going to say, okay.
And, it's telling me there is one weakly

77
00:05:56,542 --> 00:06:01,441
connected component, which just means if
we treat the edges as undirected, everyone

78
00:06:01,441 --> 00:06:04,600
is connected directly or undirectly, which
is true.

79
00:06:04,600 --> 00:06:09,424
But we just have a single connected
network but it's telling me that there are

80
00:06:09,424 --> 00:06:14,392
actually 11 different strongly connected
components, meaning that information

81
00:06:14,392 --> 00:06:18,875
wouldn't circulate throughout the network
following directed edges.

82
00:06:18,876 --> 00:06:25,245
So, we're going to close this, and now,
let's color the nodes by the strongly

83
00:06:25,245 --> 00:06:30,456
connected components.
So, I'm going to go to Partition here, and

84
00:06:30,456 --> 00:06:36,252
I want to partition the Nodes, and I'm
going to click Refresh to just get this

85
00:06:36,252 --> 00:06:41,132
new partition.
Which should be the Strongly-Connected ID.

86
00:06:41,132 --> 00:06:47,618
And, it's giving me some colors some of
these seem a little bit similar, so let me

87
00:06:47,618 --> 00:06:51,538
just see.
I'm going to right-click and say Randomize

88
00:06:51,538 --> 00:06:54,930
colors and these look a little bit
different.

89
00:06:54,930 --> 00:07:01,784
So I'm going to Apply them.
And I think this paints a very, very clear

90
00:07:01,784 --> 00:07:08,877
picture about the dynamics in this in this
among this group of girls.

91
00:07:08,878 --> 00:07:15,802
So first of all we have the purple cluster
here, where all of these girls may be

92
00:07:15,802 --> 00:07:23,146
sharing information with each other and
they also include these two very popular

93
00:07:23,146 --> 00:07:28,502
girls, Marion and Eva.
You see this one girl here, Alice, who

94
00:07:28,502 --> 00:07:34,266
names two girls in this cluster, but her
color is actually different.

95
00:07:34,266 --> 00:07:40,360
She's in the strongly connected component
all by herself because, actually, no one

96
00:07:40,360 --> 00:07:44,032
selected Alice as someone they wanted to
sit with.

97
00:07:44,033 --> 00:07:50,670
We see, this little component here of
three nodes in blue, Betty, Hilda and

98
00:07:50,670 --> 00:07:55,112
Hazel.
All had, well, Hazel had selected Hilda

99
00:07:55,112 --> 00:08:00,960
and vice versa, and Betty had selected
Hilda and vice versa.

100
00:08:00,960 --> 00:08:08,154
So they form one strongly connected
component a similar dynamic with Jane,

101
00:08:08,154 --> 00:08:13,956
Mary, and Edna.
And here is an example with Helen, Jean,

102
00:08:13,956 --> 00:08:17,600
and Robin.
So here, you actually have a closed triad,

103
00:08:17,600 --> 00:08:21,880
but what has happened is that Helen and
Jean mutually chose each other.

104
00:08:21,880 --> 00:08:26,850
And then, Jean chose Robin as her second
choice.

105
00:08:26,851 --> 00:08:32,230
And then, Robin chose Helen.
So you can kind of circulate in here.

106
00:08:32,230 --> 00:08:37,456
And then, Ada and Cora are their own
strongly connected component, because they

107
00:08:37,456 --> 00:08:42,807
chose each other.
So this shows you a bit of the dynamic of,

108
00:08:42,807 --> 00:08:49,560
of [laugh] girls in, in, you know, earlier
stages of their lives.

109
00:08:49,560 --> 00:08:55,020
But also, there's some recent research
showing that most social networks, when

110
00:08:55,020 --> 00:09:00,558
you're naming friends tend to be fairly
asymmetrical like this, and in this case,

111
00:09:00,558 --> 00:09:04,075
because it's a relatively small network,
etcetera.

112
00:09:04,075 --> 00:09:07,552
Even just looking at the strongly
connected components without applying

113
00:09:07,552 --> 00:09:11,360
community detection algorithms, which is
something that we're going to be doing

114
00:09:11,360 --> 00:09:17,420
later on.
We can figure out a lot of the dynamic

115
00:09:17,420 --> 00:09:20,373
between the girls.
