1
00:00:00,357 --> 00:00:06,452
[MUSIC]. 

2
00:00:06,452 --> 00:00:09,212
Okay, so I want to spend a little bit 
more time on the details of MapReduce 

3
00:00:09,212 --> 00:00:13,791
versus our relation database. 
beyond just the sort of how the pair, 

4
00:00:13,791 --> 00:00:17,354
query processing happens. 
We saw that parallel query processing is 

5
00:00:17,354 --> 00:00:19,786
largely the same. 
Some of the, many of the algorithms are 

6
00:00:19,786 --> 00:00:21,972
sort of shared between them. 
There's a ton of details here that I'm 

7
00:00:21,972 --> 00:00:25,462
not going to have time to go over. 
but the takeaway is that, that the basic 

8
00:00:25,462 --> 00:00:30,670
strategy for, for performing parallel 
processing is the same between them. 

9
00:00:30,670 --> 00:00:33,442
But there's other features that 
relational databases have and I've listed 

10
00:00:33,442 --> 00:00:36,104
some of them here. 
So we mentioned declarative query 

11
00:00:36,104 --> 00:00:39,048
languages, and we mentioned that those 
start to show up in Pig and especially 

12
00:00:39,048 --> 00:00:42,223
HIVE. 
Now there's a notion of a schema in a 

13
00:00:42,223 --> 00:00:44,630
relational database that we didn't talk 
too much about. 

14
00:00:44,630 --> 00:00:48,315
But this is you know, structure on your 
data that is enforced at the time of the 

15
00:00:48,315 --> 00:00:53,121
data being presented to the system. 
So when your data does, does not conform 

16
00:00:53,121 --> 00:00:56,970
to the schema, can be rejected 
automatically by the database. 

17
00:00:56,970 --> 00:00:59,480
This is a pretty good idea because it 
helps keep your data clean. 

18
00:00:59,480 --> 00:01:03,260
It's not real feasible in many contexts, 
because you know the data's fundamentally 

19
00:01:03,260 --> 00:01:06,446
dirty and so saying that you have to 
clean it up before you're allowed to 

20
00:01:06,446 --> 00:01:10,470
process it, just isn't going to fly. 
Right. 

21
00:01:10,470 --> 00:01:12,948
And so this is one of the reasons why 
MapReduce is attractive, is that it 

22
00:01:12,948 --> 00:01:15,804
doesn't require that you enforce a schema 
before you were allowed to work with the 

23
00:01:15,804 --> 00:01:18,630
data. 
However, you know, it doesn't mean that 

24
00:01:18,630 --> 00:01:21,890
schemas are a bad idea when they're, when 
they're available. 

25
00:01:21,890 --> 00:01:25,358
And in fact, really, you know, even with 
MapReduce, a schema is really there, it's 

26
00:01:25,358 --> 00:01:28,500
just that it's hidden inside the 
application. 

27
00:01:28,500 --> 00:01:29,851
Right? 
So when you read a record, you're 

28
00:01:29,851 --> 00:01:32,603
assuming that the first record, the first 
element in the record is going to be an 

29
00:01:32,603 --> 00:01:35,140
integer and the second record is going to 
be a date, and the third record is 

30
00:01:35,140 --> 00:01:39,190
going to be a string. 
So that schema is really present, it's 

31
00:01:39,190 --> 00:01:43,890
just present in your code as opposed to 
pushed down into the system itself. 

32
00:01:43,890 --> 00:01:47,270
And there's a lot of great empirical 
evidence over the years that suggest it's 

33
00:01:47,270 --> 00:01:51,469
better to push it down into the data 
itself, when where possible. 

34
00:01:51,469 --> 00:01:54,647
And in fact you're starting to see this. 
So HIVE and Pig again have some notion of 

35
00:01:54,647 --> 00:01:58,965
schema, as does DryadLINQ, as does some, 
some emerging system. 

36
00:01:58,965 --> 00:02:04,708
there's, there's a system called Hadapt, 
that won't talk about really at all. 

37
00:02:04,708 --> 00:02:09,264
But combines sort of Hadoop level query 
processing for parallelism and on the 

38
00:02:09,264 --> 00:02:14,370
individual nodes there's a relational 
database operating. 

39
00:02:14,370 --> 00:02:21,022
And one of the reasons among many is to 
have access to schema constraints. 

40
00:02:21,022 --> 00:02:26,355
fine, logical data independence. 
This actually, you don't see quite so 

41
00:02:26,355 --> 00:02:28,710
much, so this is the notion of views, 
right? 

42
00:02:28,710 --> 00:02:31,792
Does the system support views or not and 
you haven't seen quite as many instances 

43
00:02:31,792 --> 00:02:35,705
of Hadoop systems that support views but 
I predict they'll be coming. 

44
00:02:35,705 --> 00:02:39,555
Indexing is another one. 
So, we talked about how to make things 

45
00:02:39,555 --> 00:02:43,983
scalable that one way to do it is, to. 
Who've derived these indexes to support 

46
00:02:43,983 --> 00:02:48,810
sort of logarithmic time access to data. 
that's not available in vanilla 

47
00:02:48,810 --> 00:02:51,628
MapReduce. 
Every time you write a MapReduce job, 

48
00:02:51,628 --> 00:02:54,480
you're going to touch every single record 
on the input. 

49
00:02:54,480 --> 00:02:58,440
You're not going to be able to zoom right 
in to a particular record of interest. 

50
00:02:58,440 --> 00:03:05,690
That's wasteful and it was recognized to 
be wasteful and so one of the solutions. 

51
00:03:05,690 --> 00:03:07,970
You, you see people adding indexing 
features to Hadoop. 

52
00:03:07,970 --> 00:03:12,552
And a H base is an open source 
implementation of a, another proposal by 

53
00:03:12,552 --> 00:03:18,395
Google for a system called Big Table. 
That among other things provides, kind of 

54
00:03:18,395 --> 00:03:21,992
quick access to individual records. 
And H based is designed to be sort of 

55
00:03:21,992 --> 00:03:25,604
compatible with Hadoop. 
And so now you can design your system get 

56
00:03:25,604 --> 00:03:27,833
the best of both worlds. 
Okay. 

57
00:03:27,833 --> 00:03:31,668
See, you can't get some indexing along 
with your MapReduce style [UNKNOWN] in 

58
00:03:31,668 --> 00:03:34,506
your face. 
And once again, I'll mention Hadapt here 

59
00:03:34,506 --> 00:03:36,864
as well. 
One of the motivations for Hadapt to be 

60
00:03:36,864 --> 00:03:39,868
able to provide index young individual 
nodes. 

61
00:03:39,868 --> 00:03:42,090
Okay. 
Fine, so I'll skip caching/materialized 

62
00:03:42,090 --> 00:03:44,474
views. 
This is the same same as logical data 

63
00:03:44,474 --> 00:03:48,147
[INAUDIBLE] accept you can actually 
pre-generate the views as opposed to 

64
00:03:48,147 --> 00:03:52,711
evaluate them all at run time. 
But we're not going to [INAUDIBLE] too 

65
00:03:52,711 --> 00:03:55,306
much about that. 
And then transactions which I'll talk 

66
00:03:55,306 --> 00:03:58,240
about in a couple of segments in the 
context of, of no sequel. 

67
00:03:58,240 --> 00:04:03,950
But while, databases are, so databases 
are very good at transaction. 

68
00:04:03,950 --> 00:04:06,290
They were thrown out the window among, 
among other things in this kind of 

69
00:04:06,290 --> 00:04:10,196
context of MapReduce and their sequel. 
And they're starting to come back. 

70
00:04:10,196 --> 00:04:12,861
Okay. 
But remember you know, what MapReduce did 

71
00:04:12,861 --> 00:04:16,566
provide was very, very high scalability 
so this is you know, thousands and up, 

72
00:04:16,566 --> 00:04:20,928
thousands machines and up. 
And it also provided this notion of fault 

73
00:04:20,928 --> 00:04:23,256
tolerance. 
So relational databases didn't unders, 

74
00:04:23,256 --> 00:04:25,490
didn't really treat fault tolerance this 
way. 

75
00:04:25,490 --> 00:04:29,890
They were unbelievably good at, you know, 
recovery, right? 

76
00:04:29,890 --> 00:04:33,312
If you were, because of this notion of 
transactions, if you were sort of 

77
00:04:33,312 --> 00:04:37,254
operating on the database and everything 
went kaput. 

78
00:04:37,254 --> 00:04:41,680
given some time, it would figure 
everything out and recover. 

79
00:04:43,300 --> 00:04:46,550
and you will, you, you can be guaranteed 
to have lost no data, okay. 

80
00:04:46,550 --> 00:04:50,138
That's fine but that's not the same thing 
as saying during query processing while a 

81
00:04:50,138 --> 00:04:53,740
single query is running what if something 
goes wrong. 

82
00:04:53,740 --> 00:04:56,220
Do I always have to start back over from 
the beginning or not? 

83
00:04:56,220 --> 00:04:58,940
And the sort of the implict assumpti-, 
assumption with relational databases was 

84
00:04:58,940 --> 00:05:02,315
that your queries aren't taking long 
enough for that to really matter. 

85
00:05:02,315 --> 00:05:06,735
But in the era of big data, of massive 
data analytics of course you have queries 

86
00:05:06,735 --> 00:05:10,890
that are running from many many hours. 
Right? 

87
00:05:10,890 --> 00:05:14,130
And so having to restart this, in the 
course they're running on many many 

88
00:05:14,130 --> 00:05:17,299
machines where failures are bound to 
happen. 

89
00:05:17,299 --> 00:05:20,755
And so that context is something that 
MapReduce sort of really motivated and 

90
00:05:20,755 --> 00:05:24,005
now you're, you see modern parallel 
databases. 

91
00:05:24,005 --> 00:05:28,194
Capturing some of [UNKNOWN] tolerance in 
general. 

92
00:05:28,194 --> 00:05:30,273
Okay. 
So this is sort of a list of some kind of 

93
00:05:30,273 --> 00:05:33,649
partialistic contribution for relational 
databases. 

94
00:05:33,649 --> 00:05:37,253
And this is a partialistic contribution, 
maybe a completeness of contributions 

95
00:05:37,253 --> 00:05:39,842
from MapReduce. 
And my point is that you see a lot of 

96
00:05:39,842 --> 00:05:42,065
mixing and matching going on but the 
design space is being more fully 

97
00:05:42,065 --> 00:05:44,302
explored. 
It used to be sort of all about 

98
00:05:44,302 --> 00:05:46,993
relational databases with their choi-,, 
their choice in the giant space, and then 

99
00:05:46,993 --> 00:05:49,560
MapReduce kind of rebooted that a little 
bit. 

100
00:05:49,560 --> 00:05:53,078
And now you see kind of a more fluid mix, 
people sort of cherry picking features. 

101
00:05:53,078 --> 00:05:55,319
Okay fine. 
And then the last one I guess I didn't 

102
00:05:55,319 --> 00:05:58,903
talk about here is, what I think was 
really, really powerful about MapReduce 

103
00:05:58,903 --> 00:06:03,410
is it turned, you know, it turned, it 
turned the. 

104
00:06:03,410 --> 00:06:07,490
Army of JAVA programmers that are out 
there into distributive systems program. 

105
00:06:07,490 --> 00:06:10,882
A mere mortal JAVA programmer could all 
of a sudden be productive processing 

106
00:06:10,882 --> 00:06:14,380
hundreds of terabytes without necessarily 
having to learn anything about the 

107
00:06:14,380 --> 00:06:18,721
distributive systems. 
That was really, really powerful. 

108
00:06:18,721 --> 00:06:22,320
Right for, with the, the analog in 
databases was, I mean, you had to become 

109
00:06:22,320 --> 00:06:26,140
a database expert to be able to use these 
things. 

110
00:06:26,140 --> 00:06:28,270
Okay. 
And so, I think that that impact is hard 

111
00:06:28,270 --> 00:06:31,894
to overstate. 
Right, the ability of one person to get 

112
00:06:31,894 --> 00:06:36,286
work done that used to be, you hire a 
massive team and six months of work was 

113
00:06:36,286 --> 00:06:39,680
significant, alright. 

