1
23:59:59,500 --> 00:00:05,077
[MUSIC]. 

2
00:00:05,077 --> 00:00:08,269
Okay. 
So, the next system we're going to look 

3
00:00:08,269 --> 00:00:11,267
at is CouchDB. 
Which began in 2005, but is still has 

4
00:00:11,267 --> 00:00:15,613
undergone a lot of updates and is still 
pretty popular today. 

5
00:00:15,613 --> 00:00:19,261
And so, this is an example of one of 
these document-oriented stores that Rick 

6
00:00:19,261 --> 00:00:22,598
Cattell talked about. 
And so here, just a look at the features 

7
00:00:22,598 --> 00:00:26,683
we're talking about here. 
You got the scale, you got primary index. 

8
00:00:26,683 --> 00:00:30,920
Now, now we start to see the requirement 
for secondary indexes. 

9
00:00:30,920 --> 00:00:34,430
Meaning that you can look up, not just by 
the key, but by other kinds of values in 

10
00:00:34,430 --> 00:00:38,007
the document. 
And we'll see how that's, how that works, 

11
00:00:38,007 --> 00:00:41,534
okay? 
And then transactions, also are a little 

12
00:00:41,534 --> 00:00:45,480
bit better. 
We can what I mean by record here is that 

13
00:00:45,480 --> 00:00:51,173
you can change multiple values within a 
single document of particular. 

14
00:00:51,173 --> 00:00:55,583
Which is a set of key values pairs, as 
opposed to just a single key value pair, 

15
00:00:55,583 --> 00:00:59,205
okay? 
And then, there actually is some support 

16
00:00:59,205 --> 00:01:02,640
for analytics, and we'll see how this 
works, too. 

17
00:01:02,640 --> 00:01:06,759
This is all through the notion, this 
concept of views in CouchDB. 

18
00:01:06,759 --> 00:01:10,791
But, you can actually run little 
MapReduce scripts to compute derive new 

19
00:01:10,791 --> 00:01:14,553
values from existing document source, 
alright? 

20
00:01:14,553 --> 00:01:17,847
And then, the others we have here is 
views, which is somewhat unique. 

21
00:01:17,847 --> 00:01:20,531
You can see this column is fairly 
sparsely populated. 

22
00:01:20,531 --> 00:01:23,331
And I hit the concept of views pretty 
hard when we were talking about 

23
00:01:23,331 --> 00:01:27,205
relational databases. 
And argued that it was pretty fundamental 

24
00:01:27,205 --> 00:01:30,660
to not only the relational model, but a 
pretty important concept in general. 

25
00:01:30,660 --> 00:01:33,274
It gave you this notion of logical data 
independence. 

26
00:01:33,274 --> 00:01:35,790
And so whenever you see views, that's a 
good thing. 

27
00:01:35,790 --> 00:01:40,250
So, CouchDB has those well. 
Alright. 

28
00:01:40,250 --> 00:01:42,950
So, the data model here is 
document-oriented, as we said, where a 

29
00:01:42,950 --> 00:01:46,930
document is a set of key value pairs. 
And so here, in this application, you 

30
00:01:46,930 --> 00:01:49,938
know, these are perhaps blog post where 
they have a subject, the subject, you 

31
00:01:49,938 --> 00:01:53,772
know, the key of subject. 
And the value is some texturing, and they 

32
00:01:53,772 --> 00:01:57,116
have an author, and they have a date of 
which they are posted. 

33
00:01:57,116 --> 00:02:00,660
And they have a set of tags and this is 
what we said makes it a DOC model. 

34
00:02:00,660 --> 00:02:03,433
We said it could be sort of nested so 
this is okay to have a list of objects 

35
00:02:03,433 --> 00:02:06,356
here. 
And then they have a body which is 

36
00:02:06,356 --> 00:02:12,380
another sting fine, and you'll notice 
that this is actually Jason compliant. 

37
00:02:12,380 --> 00:02:15,674
We looked at Jason in the Twitter 
assignment, and so it's another 

38
00:02:15,674 --> 00:02:18,708
occurrence of this. 
And one of the reasons why I want to make 

39
00:02:18,708 --> 00:02:21,100
sure your working with Jason before, is 
it does get popular. 

40
00:02:21,100 --> 00:02:24,754
It's a CouchDB, all data's are presented 
in Jason, and all the request Macro 

41
00:02:24,754 --> 00:02:30,860
represented in Jason as well, okay? 
So, how do updates work in this context? 

42
00:02:30,860 --> 00:02:35,414
Well, as I mentioned, they you know, you 
can make, it is fully transactional 

43
00:02:35,414 --> 00:02:41,168
within, within a single document. 
So, full consistency within a document, 

44
00:02:41,168 --> 00:02:45,712
meaning that I can sort of grab hold of 
the document and logically sort of lock 

45
00:02:45,712 --> 00:02:49,820
it. 
And up, make whatever updates I want. 

46
00:02:49,820 --> 00:02:54,784
Now, it doesn't actually take a lock 
because it uses what's called Optimistic 

47
00:02:54,784 --> 00:02:58,684
Concurrency Control. 
Meaning, that it sort of assumes that, 

48
00:02:58,684 --> 00:03:01,280
that conflicts aren't going to happen 
optimistically. 

49
00:03:01,280 --> 00:03:03,920
And so now, if I, if I check out that 
document and try to make updates to it, 

50
00:03:03,920 --> 00:03:08,015
anywhere I want, all throughout it, and 
then try to commit that change. 

51
00:03:08,015 --> 00:03:11,151
Somewhere else is doing this same has 
done this as committed their own team to 

52
00:03:11,151 --> 00:03:13,710
this. 
In the meantime, my changes will fail, 

53
00:03:13,710 --> 00:03:16,563
but we assume that, that doesn't happen 
that often. 

54
00:03:16,563 --> 00:03:19,025
So, it's okay to just sort of fail in 
those, in those cases. 

55
00:03:19,025 --> 00:03:22,125
And really all you have to do in cases 
like this his check out the new changes 

56
00:03:22,125 --> 00:03:25,117
and make whatever changes you want. 
Fine. 

57
00:03:25,117 --> 00:03:28,839
But there is no multi-row transactions. 
What I mean by multi-row here is 

58
00:03:28,839 --> 00:03:31,767
multi-document. 
There's no, let's say for example I 

59
00:03:31,767 --> 00:03:34,697
update my status. 
I have to update my document that 

60
00:03:34,697 --> 00:03:39,119
describes my current state of the world. 
But may be I also want to update all of 

61
00:03:39,119 --> 00:03:42,681
my friends Walls, right? 
Their, their pages to, with, that 

62
00:03:42,681 --> 00:03:46,460
reflects my new status. 
And you can't guarantee that that happens 

63
00:03:46,460 --> 00:03:49,988
synchronously in CouchDB. 
But, you know, in this particular 

64
00:03:49,988 --> 00:03:53,180
application and in many others, maybe 
that's okay. 

65
00:03:53,180 --> 00:03:55,383
All right. 
You know, you can do it in two separate 

66
00:03:55,383 --> 00:03:57,870
steps. 
You update your status, and then you 

67
00:03:57,870 --> 00:04:00,517
update theirs. 
And there will be a period of time when 

68
00:04:00,517 --> 00:04:03,478
their, you know, view of your current 
status is out of date, but maybe that's 

69
00:04:03,478 --> 00:04:09,670
all right. 
Fine. 

70
00:04:09,670 --> 00:04:14,590
Alright, so this notion of views works 
like this. 

71
00:04:14,590 --> 00:04:17,956
A view specification, which is in itself, 
a CouchDB document, a set of key value 

72
00:04:17,956 --> 00:04:21,946
pairs. 
But it's sort of a special one looks like 

73
00:04:21,946 --> 00:04:25,014
this. 
There's some metadata information here, 

74
00:04:25,014 --> 00:04:28,905
and then there's this key views which is, 
which is a dictionary of things. 

75
00:04:28,905 --> 00:04:31,987
And so, this this specification has three 
views in it, one called all, one called 

76
00:04:31,987 --> 00:04:34,840
by lastname, and one called total 
purchases. 

77
00:04:34,840 --> 00:04:39,049
And each one of these views is going to 
be is implementing a set of key value 

78
00:04:39,049 --> 00:04:42,004
pairs. 
I'm just going to be implementing a 

79
00:04:42,004 --> 00:04:45,903
dictionary. 
And so the all view, okay? 

80
00:04:45,903 --> 00:04:49,910
So, fine. 
So, how are these implemented? 

81
00:04:49,910 --> 00:04:53,447
Well, you know, it's, speaking of 
recurring concepts. 

82
00:04:53,447 --> 00:04:58,177
Views are recurring from logical 
independence to relational databases. 

83
00:04:58,177 --> 00:05:03,135
and here, MapReduce is recurring even, 
outside the context of literally Hadoop, 

84
00:05:03,135 --> 00:05:05,452
okay? 
And so here, the map function, reduce 

85
00:05:05,452 --> 00:05:07,740
function, are actually written in 
JavaScript. 

86
00:05:07,740 --> 00:05:11,070
Again, everything in CouchDB is 
JavaScript, and they look like this. 

87
00:05:11,070 --> 00:05:14,590
So, the all view has just the map 
function, no reduce function. 

88
00:05:14,590 --> 00:05:18,275
And in JavaScript, you can write these 
anonymous functions like in, in this sort 

89
00:05:18,275 --> 00:05:21,616
of syntax. 
So, this function doesn't have a name, it 

90
00:05:21,616 --> 00:05:24,979
just says, hey, I've, I've got to 
function with no name, with a single 

91
00:05:24,979 --> 00:05:29,445
argument called doc. 
And the body looks like this, it says, if 

92
00:05:29,445 --> 00:05:33,280
doc.Type equals customer, then emit a key 
value pair, or the key is null and the 

93
00:05:33,280 --> 00:05:37,511
value is doc. 
So, we don't really care about the key in 

94
00:05:37,511 --> 00:05:42,305
that case, we just care about the doc. 
Alright, so this is all customers. 

95
00:05:42,305 --> 00:05:45,506
Fine. 
The by last name view, you can imagine, 

96
00:05:45,506 --> 00:05:47,860
the key is going to be last name in this 
case. 

97
00:05:47,860 --> 00:05:49,555
And so here, we have another anonymous 
function. 

98
00:05:49,555 --> 00:05:54,611
If doc.Type equals customer, then emit 
the last name, followed by the entire 

99
00:05:54,611 --> 00:05:59,206
document. 
And so now, this allows clients to search 

100
00:05:59,206 --> 00:06:06,289
efficiently by last name, and these views 
will, are in custody or computed. 

101
00:06:06,289 --> 00:06:11,833
You know, they materialize sort of 
eagerly and stored in these distributed 

102
00:06:11,833 --> 00:06:17,850
BG struct index structures just for 
various look-ups. 

103
00:06:17,850 --> 00:06:20,597
And, so this is how they implement those 
secondary indexes in that column in our, 

104
00:06:20,597 --> 00:06:23,850
in our table, okay? 
So now, we can look-up by last name as 

105
00:06:23,850 --> 00:06:27,225
well as by document ID. 
And so, you can even go a little further, 

106
00:06:27,225 --> 00:06:29,925
you don't have to just do simple 
key-value pairs, you can even do a little 

107
00:06:29,925 --> 00:06:33,535
bit of computation. 
So total purchases here, the map function 

108
00:06:33,535 --> 00:06:37,125
again, takes in a document and says, 
well, if doc.Type equals purchase. 

109
00:06:37,125 --> 00:06:40,500
So, I'm not talking about customers 
anymore, just purchases. 

110
00:06:40,500 --> 00:06:44,131
Then, emit the key of doc.Customer and 
the value of doc.Amount. 

111
00:06:44,131 --> 00:06:47,722
But then, we're also going to find a 
reduce function and the reduce function 

112
00:06:47,722 --> 00:06:52,575
adds up all the values of the amounts. 
And to here, we have given a customer, if 

113
00:06:52,575 --> 00:06:56,095
I have give you a customer, you can 
return it's the total amount of all it's 

114
00:06:56,095 --> 00:07:02,564
purchases, okay? 
And CouchDB maintains all these views as 

115
00:07:02,564 --> 00:07:10,490
things change with eventual consistency 
sort of semantics. 

116
00:07:10,490 --> 00:07:14,530
And so, we have a lot of things going on 
here with one concept. 

117
00:07:14,530 --> 00:07:18,660
We have secondary indexes, we have 
logical data independents, we have 

118
00:07:18,660 --> 00:07:23,520
MapReduce computation, okay? 
So, I had checked the box in the table 

119
00:07:23,520 --> 00:07:27,732
saying the CouchDB could do joins in 
analytics. 

120
00:07:27,732 --> 00:07:32,010
And it's not really quite true there 
somewhat limited, so let me show you an 

121
00:07:32,010 --> 00:07:38,126
example of how they do, sort of join. 
So, they have this concept of view 

122
00:07:38,126 --> 00:07:43,694
coalition, and what you can do here is 
write a map function in a view that looks 

123
00:07:43,694 --> 00:07:48,019
like this. 
And so, here we're tying to group 

124
00:07:48,019 --> 00:07:50,730
together all the comments associated with 
a post. 

125
00:07:50,730 --> 00:07:54,512
So, it's sort of a join between the 
comments table, if you will, and, and the 

126
00:07:54,512 --> 00:07:59,100
post, the blog posts table. 
And so, this map function is the same 

127
00:07:59,100 --> 00:08:02,920
kind of thing we did in the MapReduce 
assignment to do a join. 

128
00:08:02,920 --> 00:08:06,165
This is what you have to do when you do a 
join in MapReduce is you take, you 

129
00:08:06,165 --> 00:08:09,835
pretend like the whole collection of 
documents. 

130
00:08:09,835 --> 00:08:14,100
Regardless of type, regardless of source 
relation Is one big set of objects. 

131
00:08:14,100 --> 00:08:16,746
And in your map function, you sort out 
which one's which and make sure to hash 

132
00:08:16,746 --> 00:08:20,828
on the same key. 
So here, the ID of the document post and 

133
00:08:20,828 --> 00:08:26,150
then in this document .docpost refers to 
some post ID. 

134
00:08:26,150 --> 00:08:33,265
That's now post ID go to one. 
And all the comments associated with post 

135
00:08:33,265 --> 00:08:36,380
id equals one will go together. 
But they do this funny thing. 

136
00:08:36,380 --> 00:08:39,418
They say the key here in the key value 
pair is DOC ID identity followed by the 

137
00:08:39,418 --> 00:08:43,564
number zero in the one case, and the 
number one in the other case. 

138
00:08:43,564 --> 00:08:49,280
And what's happened here is that all, 
everything in Mac, in CouchDB is sorted. 

139
00:08:49,280 --> 00:08:52,880
And so, you got things sorted by document 
ID and then sorted by this exturbate. 

140
00:08:52,880 --> 00:08:56,418
So, now you have the post ID coming 
first, then all the comments coming 

141
00:08:56,418 --> 00:08:59,724
second. 
And now, which you could, you know, you 

142
00:08:59,724 --> 00:09:03,940
think you could just do a reduce function 
to actually compute the join that you 

143
00:09:03,940 --> 00:09:08,218
might want. 
But it turns out reduce functions don't 

144
00:09:08,218 --> 00:09:12,230
quite work that way, and there's some 
scalability issues. 

145
00:09:12,230 --> 00:09:14,910
So, what people recommend is to do this 
view collation trick where you still spit 

146
00:09:14,910 --> 00:09:18,760
out the data in sorted order, but then 
you can query it. 

147
00:09:18,760 --> 00:09:24,008
Say a client application could query it 
like this where they say, start key is 

148
00:09:24,008 --> 00:09:29,160
just a post ID without the extra prefix 
at all. 

149
00:09:29,160 --> 00:09:33,020
And the end key is post ID with the 
number two. 

150
00:09:33,020 --> 00:09:35,918
So, now this gets the full range of the 
post Id with, which has the key of zero, 

151
00:09:35,918 --> 00:09:39,270
followed by all the comments for the post 
ID in one. 

152
00:09:39,270 --> 00:09:42,516
So, you got everything in one big list. 
And now, you can process it in order to 

153
00:09:42,516 --> 00:09:47,510
show a shape say in the application. 
So, you essentially constructed the join 

154
00:09:47,510 --> 00:09:52,200
by cheating with the order, okay? 
So, fine. 

155
00:09:52,200 --> 00:09:57,520
So perhaps, you could argue that you have 
some joins and analytics, but it's a bit 

156
00:09:57,520 --> 00:09:59,973
glib to say so. 

