1
00:00:07,339 --> 00:00:13,990
In data science, the term data leakage sometimes just referred to as leakage,

2
00:00:13,990 --> 00:00:16,094
describes the situation where the data you're

3
00:00:16,094 --> 00:00:18,359
using to train the machine learning algorithm

4
00:00:18,359 --> 00:00:21,210
happens to include unexpected extra information

5
00:00:21,210 --> 00:00:24,269
about the very thing you're trying to predict.

6
00:00:24,269 --> 00:00:26,460
Basically, leakage occurs any time

7
00:00:26,460 --> 00:00:29,184
that information is introduced about the target label or

8
00:00:29,184 --> 00:00:35,240
value during training that would not legitimately be available during actual use.

9
00:00:35,240 --> 00:00:38,950
Maybe the simplest example of data leakage would be if we

10
00:00:38,950 --> 00:00:43,225
included the true label of a data instance as a feature in the model.

11
00:00:43,225 --> 00:00:45,655
The model would learn the equivalent of,

12
00:00:45,655 --> 00:00:48,159
if this object is labeled as an apple,

13
00:00:48,159 --> 00:00:50,590
predict it's an apple.

14
00:00:50,590 --> 00:00:54,490
Another clear example of data leakage that we've seen before is having

15
00:00:54,490 --> 00:00:59,914
test data accidentally included in the training data which leads to over fitting.

16
00:00:59,914 --> 00:01:03,159
However, data leakage can happen for many other reasons too,

17
00:01:03,159 --> 00:01:07,224
often in ways that are quite subtle and hard to detect.

18
00:01:07,224 --> 00:01:09,625
When data leakage does occur,

19
00:01:09,625 --> 00:01:14,694
it typically causes results during your model development phase that are too optimistic,

20
00:01:14,694 --> 00:01:18,265
followed by the nasty surprise of disappointing results

21
00:01:18,265 --> 00:01:23,015
after the prediction model is actually deployed and evaluated on new data.

22
00:01:23,015 --> 00:01:28,030
In other words, leakage can cause your system to learn a sub optimal model that

23
00:01:28,030 --> 00:01:34,405
does much worse in actual deployment than a model developed in a leak free setting.

24
00:01:34,405 --> 00:01:38,754
So leakage can have dramatic implications in the real world ranging from

25
00:01:38,754 --> 00:01:40,390
the financial cost of making

26
00:01:40,390 --> 00:01:44,920
a bad monetary and engineering investment in something that doesn't actually work,

27
00:01:44,920 --> 00:01:47,319
to system failures that hurt customers

28
00:01:47,319 --> 00:01:51,969
perception of your system's quality or impact to the company's brand.

29
00:01:51,969 --> 00:01:56,409
For these reasons, data leakage is one of the most serious and widespread problems

30
00:01:56,409 --> 00:01:58,209
in data mining and machine learning

31
00:01:58,209 --> 00:02:01,120
and something that as a machine learning practitioner,

32
00:02:01,120 --> 00:02:03,944
you must always be on guard against.

33
00:02:03,944 --> 00:02:07,870
So now, we'll cover what data leakage is, why it matters,

34
00:02:07,870 --> 00:02:12,780
how it can be detected and how you might avoid it in your applications.

35
00:02:12,780 --> 00:02:18,370
As an aside, this term data leakage is also used in the field of data security to

36
00:02:18,370 --> 00:02:20,800
mean the unauthorized transfer of information

37
00:02:20,800 --> 00:02:24,460
outside of a secure facility like a data center.

38
00:02:24,460 --> 00:02:27,069
However, in some ways this security based meaning is

39
00:02:27,069 --> 00:02:29,905
actually somewhat appropriate for our machine learning setting,

40
00:02:29,905 --> 00:02:33,115
given the importance of keeping information about the prediction

41
00:02:33,115 --> 00:02:38,060
securely separated from the training and model development phase.

42
00:02:38,060 --> 00:02:43,895
Let's look at some more subtle examples of data leakage problems.

43
00:02:43,895 --> 00:02:46,770
One classic case happens when information about

44
00:02:46,770 --> 00:02:50,504
the future that would not legitimately be available in actual use,

45
00:02:50,504 --> 00:02:53,280
is included in the training data.

46
00:02:53,280 --> 00:02:57,060
Suppose you are developing a retail website and building classifier to

47
00:02:57,060 --> 00:03:02,719
predict whether the user is likely to stay and view another page or leave the site.

48
00:03:02,719 --> 00:03:05,235
If the classifier predicts they're about to leave,

49
00:03:05,235 --> 00:03:10,135
the website might pop up something that offers incentives to continue shopping.

50
00:03:10,135 --> 00:03:14,789
An example of a feature that contains leaked information would be

51
00:03:14,789 --> 00:03:16,949
the user's total session length or

52
00:03:16,949 --> 00:03:21,150
the total number of pages they viewed during their visit to the site.

53
00:03:21,150 --> 00:03:23,430
This total is often added as a new column during

54
00:03:23,430 --> 00:03:27,180
the post-processing phase of the visit log data, for example.

55
00:03:27,180 --> 00:03:30,569
This feature has information about the future namely,

56
00:03:30,569 --> 00:03:33,705
how many more visits the user is going to make.

57
00:03:33,705 --> 00:03:37,034
That's impossible to know in an actual deployment.

58
00:03:37,034 --> 00:03:41,300
A solution is to replace the total session length feature with

59
00:03:41,300 --> 00:03:43,639
a page visit in-session feature that only

60
00:03:43,639 --> 00:03:46,685
knows the total pages visited so far in the session,

61
00:03:46,685 --> 00:03:50,444
and not how many are remaining.

62
00:03:50,444 --> 00:03:53,550
The second example of leakage might involve trying to predict

63
00:03:53,550 --> 00:03:57,569
if a customer on a bank's website was likely to open an account.

64
00:03:57,569 --> 00:04:01,064
If the user's record contains an account number field,

65
00:04:01,064 --> 00:04:04,680
it might normally be empty for users still on the process of exploring

66
00:04:04,680 --> 00:04:10,360
the site but eventually it's filled in once the user does open an account.

67
00:04:10,360 --> 00:04:12,719
Clearly the user account field is not

68
00:04:12,719 --> 00:04:16,409
a legitimate feature that should be used in this case,

69
00:04:16,409 --> 00:04:22,004
because it may not be available at the time the user is still exploring the site.

70
00:04:22,004 --> 00:04:25,730
Another example of future information leaking in the past might be,

71
00:04:25,730 --> 00:04:30,389
if you are developing a diagnostic test to predict a particular medical condition.

72
00:04:30,389 --> 00:04:34,279
The existing patient data set might contain a binary variable

73
00:04:34,279 --> 00:04:38,870
that happens to mark whether or not the patient had surgery for that condition.

74
00:04:38,870 --> 00:04:44,350
Obviously, such a variable would be highly predictive of the medical condition.

75
00:04:44,350 --> 00:04:49,300
There are many other ways predictive information could leak into this feature set.

76
00:04:49,300 --> 00:04:50,889
There might be a certain combination of

77
00:04:50,889 --> 00:04:55,795
missing diagnosis codes that was very indicative of the medical condition.

78
00:04:55,795 --> 00:04:59,379
But again, these would not be legitimate to use since

79
00:04:59,379 --> 00:05:05,339
that information isn't available while a patient's condition is still being studied.

80
00:05:05,339 --> 00:05:07,574
Finally, another example in the same patient is that

81
00:05:07,574 --> 00:05:11,115
it might involve the form of the patient ID.

82
00:05:11,115 --> 00:05:15,740
The ID might be assigned depending on a particular diagnosis path.

83
00:05:15,740 --> 00:05:20,475
In other words, the ID could be different if it's the result of a visit to a specialist,

84
00:05:20,475 --> 00:05:24,764
where the initial doctor determined that the medical condition was likely.

85
00:05:24,764 --> 00:05:27,569
This last example is a great illustration of the fact that there are

86
00:05:27,569 --> 00:05:31,334
many different ways data leakage could occur in a training set and in fact,

87
00:05:31,334 --> 00:05:36,009
it's often the case that more than one leakage problem is present at once.

88
00:05:36,009 --> 00:05:38,339
Sometimes, fixing one leaking feature can

89
00:05:38,339 --> 00:05:42,394
reveal the existence of a second one, for example.

90
00:05:42,394 --> 00:05:47,290
As a guide, here are some additional examples of data leakage.

91
00:05:47,290 --> 00:05:50,209
We can divide leakage into two main types.

92
00:05:50,209 --> 00:05:51,839
Leakage and the training data;

93
00:05:51,839 --> 00:05:56,259
typically where test data or future data gets mixed into the training data,

94
00:05:56,259 --> 00:05:57,925
and leakage in features,

95
00:05:57,925 --> 00:06:00,100
where something highly informative about

96
00:06:00,100 --> 00:06:04,795
the true label somehow gets included as a feature.

97
00:06:04,795 --> 00:06:09,850
One very important cause of data leakage is performing some kind of pre-processing

98
00:06:09,850 --> 00:06:15,079
on the entire dataset whose results influence what is seen during training.

99
00:06:15,079 --> 00:06:19,689
This can include such scenarios as computing parameters for normalizing

100
00:06:19,689 --> 00:06:23,620
and rescaling or finding minimum and maximum feature values to

101
00:06:23,620 --> 00:06:27,550
detect and remove outliers and using the distribution of

102
00:06:27,550 --> 00:06:32,230
a variable across the entire dataset to estimate missing values in the training set,

103
00:06:32,230 --> 00:06:34,670
or perform feature selection.

104
00:06:34,670 --> 00:06:39,675
Another critical need for caution occurs when working with time series data,

105
00:06:39,675 --> 00:06:42,389
where records for future events are accidentally

106
00:06:42,389 --> 00:06:46,209
used to compute features for a particular prediction.

107
00:06:46,209 --> 00:06:48,060
The session length example that we saw,

108
00:06:48,060 --> 00:06:51,360
was one instance of this but more subtle effects can

109
00:06:51,360 --> 00:06:56,535
occur if there are errors in data gathering or missing value indicators.

110
00:06:56,535 --> 00:07:00,165
If a feature relates to collecting at least one record in a time span,

111
00:07:00,165 --> 00:07:03,944
the presence of an error may give away information about the future.

112
00:07:03,944 --> 00:07:08,615
In other words, that no further observations are to be expected.

113
00:07:08,615 --> 00:07:12,360
Leakage in features includes the case where we have

114
00:07:12,360 --> 00:07:15,569
a variable like diagnosis ID and a patient record that we

115
00:07:15,569 --> 00:07:18,990
remove but neglect to also remove other variables

116
00:07:18,990 --> 00:07:23,154
known as proxy variables that contain the same or similar information.

117
00:07:23,154 --> 00:07:27,250
The patient ID in the case where the ID number had clues about the nature of

118
00:07:27,250 --> 00:07:32,214
the patient's diagnosis due to the admission process, was an example of this.

119
00:07:32,214 --> 00:07:35,709
In some cases, data set records are intentionally

120
00:07:35,709 --> 00:07:38,319
randomized or certain fields anonymized

121
00:07:38,319 --> 00:07:40,930
that contain specific information about a user such as,

122
00:07:40,930 --> 00:07:43,920
their name, location and so on.

123
00:07:43,920 --> 00:07:45,865
Depending on the prediction task,

124
00:07:45,865 --> 00:07:49,329
undoing this anonymization can reveal user or

125
00:07:49,329 --> 00:07:56,550
other sensitive information that is not legitimately available in actual use.

126
00:07:56,550 --> 00:08:01,689
Finally, any of the above examples we've discussed here could be present in

127
00:08:01,689 --> 00:08:04,040
a third party dataset that gets joined to

128
00:08:04,040 --> 00:08:07,220
the training set as an additional source of features.

129
00:08:07,220 --> 00:08:09,170
So, always be aware of the features in

130
00:08:09,170 --> 00:08:14,305
such external data and their interpretation and origin.

131
00:08:14,305 --> 00:08:19,770
So how can you detect and avoid data leakage in your applications?

132
00:08:19,770 --> 00:08:22,064
Before building the model,

133
00:08:22,064 --> 00:08:26,360
exploratory data analysis can reveal surprises in the data.

134
00:08:26,360 --> 00:08:32,335
For example, look for features very highly correlated with the target label or value.

135
00:08:32,335 --> 00:08:35,194
An example of this, from the medical diagnostic example,

136
00:08:35,194 --> 00:08:37,830
might be the binary feature that indicated

137
00:08:37,830 --> 00:08:41,500
a patient had a particular surgical procedure for the condition.

138
00:08:41,500 --> 00:08:47,269
That might be extremely highly correlated with a particular diagnosis.

139
00:08:47,269 --> 00:08:48,544
After building the model,

140
00:08:48,544 --> 00:08:52,309
look for a surprising feature behavior in the fitted model such as

141
00:08:52,309 --> 00:08:58,184
extremely high feature weights or very high information games associated with variable.

142
00:08:58,184 --> 00:09:02,649
Next, look for overall surprising model performance.

143
00:09:02,649 --> 00:09:05,259
If your model evaluation results are substantially

144
00:09:05,259 --> 00:09:08,740
higher than the same or similar problems and similar datasets,

145
00:09:08,740 --> 00:09:15,715
then look closely at the instances or features that have most influence on the model.

146
00:09:15,715 --> 00:09:19,500
One more reliable check for leakage but also potentially expensive,

147
00:09:19,500 --> 00:09:21,769
is to do a limited real world deployment of

148
00:09:21,769 --> 00:09:24,500
the trained model to see if there's a big difference between

149
00:09:24,500 --> 00:09:26,330
the estimated performance suggested by

150
00:09:26,330 --> 00:09:30,960
the model's training and development results and the actual results.

151
00:09:30,960 --> 00:09:34,940
This check that the model is generalizing well to new data is useful,

152
00:09:34,940 --> 00:09:39,139
but may not give much immediate insight into if or where the leakage is

153
00:09:39,139 --> 00:09:41,600
happening or if any drop in performance

154
00:09:41,600 --> 00:09:45,080
is due to other reasons like classical over fitting.

155
00:09:45,080 --> 00:09:47,889
There are practices you can follow to help

156
00:09:47,889 --> 00:09:51,519
reduce the chance of data leakage in your application.

157
00:09:51,519 --> 00:09:53,710
One important rule is to make sure that you perform

158
00:09:53,710 --> 00:09:58,669
any data preparation within each cross-validation fold, separately.

159
00:09:58,669 --> 00:10:02,019
In other words, if you're scaling or normalizing features,

160
00:10:02,019 --> 00:10:06,549
any statistics or parameters that you estimate for this should only be based on

161
00:10:06,549 --> 00:10:11,945
the data available in the cross-validation split and not the entire data set.

162
00:10:11,945 --> 00:10:13,600
You should also make sure that you use

163
00:10:13,600 --> 00:10:17,919
these same parameters on the corresponding held out test fold.

164
00:10:17,919 --> 00:10:20,687
If you're working with time series data,

165
00:10:20,687 --> 00:10:22,355
keep track of the time stamp that's

166
00:10:22,355 --> 00:10:25,315
associated with processing a particular data instance,

167
00:10:25,315 --> 00:10:30,250
such as a user's click on a webpage and make sure any data used to compute features

168
00:10:30,250 --> 00:10:35,620
for this instance does not include records with a later time than the cutoff value.

169
00:10:35,620 --> 00:10:38,049
This will help ensure you're not including information from

170
00:10:38,049 --> 00:10:42,539
the future in your current feature calculations or training data.

171
00:10:42,539 --> 00:10:44,669
If you have enough data,

172
00:10:44,669 --> 00:10:47,159
consider splitting off a completely separate test

173
00:10:47,159 --> 00:10:49,929
set before you even start working with a new dataset,

174
00:10:49,929 --> 00:10:55,470
and then evaluating your final model and this test data only as a very last step.

175
00:10:55,470 --> 00:10:57,605
The goal here is similar to doing

176
00:10:57,605 --> 00:11:00,259
a real world deployment to check that

177
00:11:00,259 --> 00:11:04,394
your train model does generalize reasonably well to new data.

178
00:11:04,394 --> 00:11:07,235
If there's no significant drop in performance, great.

179
00:11:07,235 --> 00:11:10,745
But if there is, leakage maybe one contributing factor,

180
00:11:10,745 --> 00:11:14,955
along with the usual suspects like classical over fitting.

181
00:11:14,955 --> 00:11:17,784
For more real world examples,

182
00:11:17,784 --> 00:11:21,111
analysis and guidance about preventing data leakage,

183
00:11:21,111 --> 00:11:25,000
you can take a look at the optional readings provided in the lesson plan.