My name is Matthew Graham. And in this set of six modules, I will be talking about data and databases. This first module we'll be doing a general consideration of data and some theory about, behind databases. This is an overview of the module structure. The second one, we will be talking about data models. In the third module, we will be talking about relational databases, in particular, since they are the most common database technology used currently. In the fourth and fifth modules, I'll be talking about SQL which is the language you use to query databases. Module four will be the basics of SQL. Module five will be about some more advanced features of SQL. And then, in Module six we will consider alternative databases to relational databases to consider such thing as, as NoSQL, and NewSQL, and other trends that are current in big data analytics. Building on top of relational databases but alternative solutions to that. So, it behooves us, I think, to try and understand exactly what it is when we talk about data. What do we mean by what is data? We have some understanding that we understand what we mean by that, but really, what were talking about of the values of qualitative or quantitative variables belonging to a set of items. You have some object that you want to describe. You've measured, you've observed it in the scientific context and those things that you've measured have values which may be integers, they may be numbers, they may be strings, they may be more composite objects. And those are the things that we refer to when we talk about data. It is that that we use. An important aspect of data is metadata. This is also data. And arguably it's more important. Metadata is the set of data which tells you something about the global properties of the data set that you're dealing with. If I was taking data at a telescope, the metadata might be the, the instrumental settings or the the weather conditions or the quality of the night, that sort of thing. It's a, a, a set of data that is about the actual binary data, the imaging data, or the spectroscopy data that's being taken by the telescope but it is constant for all the images that I might take in a night. Similarly, if I was doing a biology experiment, then if I was doing a gene, let's say of many genes the actual. Gene strings that I'm getting from the gene of A would be very different from gene to gene but the actual experimental setup, maybe the same in all the cases and then I would have metadata arising from the instrumental settings on the experiment, about the data sets that I'm getting. Metadata is also data, however, and so all the things that we can do with data we can also do with metadata. There is this idea that, the, the purpose of doing data analytics or doing data science is to extract information out of data. And there's, there's this particular pyramid structure that you can see, where you have data on the bottom layer, and the idea is, is that the cognitive value of the what you're doing increases as you progress up the pyramid, so from data we extract information. We combine information together to gain knowledge and then ultimately we're after wisdom derived from the knowledge. It's possibly a somewhat philosophical or fanciful notion but certainly the types of techniques that you're hearing about in this course are relevant for getting information out of the data. An example of the power of metadata, over data is, in the left-hand side you can see images of comet home 17p. This is a comet that went, was observed in 2007, Many people posted images on Flickr that they had taken with cell phones and very amateur camera about the, of the, of the comet. And, it was thought that there must be some way that we could use this information to improve our knowledge. Of the, orbit of this particular comet. And although the the images themselves are all ones and zeros. Standard camera image. What is useful is not that data precisely. What is more useful is the metadata for each image. And saying when it was taken, where it was taken. Using some notion of which particular bit of the comet's orbit it was capturing latitude and longitude of where the camera was situated. That sort of thing. And from the metadata attached to each image. It has been possible to combine the images into a single hall, a single scientific data set. Properly calibrated, each of those individual images, properly registered onto a scientific basis, on to an underlying framework. And therefore, put together a much more complete dataset. And then apply analytics to that more complete dataset. And get a much more accurate, orbital solution for this particular object. And that demonstrates the power of metadata as a form of data in its own right. So, as I say, metadata can be very important. Particularly useful when you're dealing with these very large datasets as we are. To give some idea of the scale of the data that we're working with, that we're possibly considering as the large end of big data analytics. We talk happily today about gigabytes of data, and terabytes of data. This laptop has a Terabyte disk in it. But a lot of the data we're talking about, or thinking about, for the next five, 10 years, is in the Peta scale. And possibly even larger, going towards the So, a Petabyte of data seems a large amount. It's enough to store the DNA of the entire US population. If I took the DNA of every member of the 300 million population of the US, I could put that into a petabyte data set and then I could clone each of those individuals twice over also take that DNA, put it in. And that would define a petabyte of data. If I was listening to music on my MP3 player, a petabyte of of of songs would last me about 2,000 years of continuous playback. A petabyte of data is about 13 and a third years of high definition video, such as you're watching now. And in the social media world where there actually more used to these very large amounts of streaming data, a petabyte of data is about one tweet per person on the planet per day for seven weeks. So, obviously, in a year, Twitter is creating, if everyone on the planet was tweeting, that would create somewhere between seven and eight petabytes of data. So, it gives you an idea of what a petabyte of data is. It may seem a lot. But in a scientific context in, in astronomy, the large. Large Synoptic Sky Telescope LSST, will produce a petabyte of data, in 60 nights. Comes online, at the end of this decade. And SKA, the Square kilometer Array Radio Telescope will produce a petabyte of data in about 90 seconds. So, these are much more believable data sets, than. You know, 13 years of video. But it demonstrates that these scale, datasets, will become very commonplace in the next decade. And we therefore need ways of managing that data. And, and working with that data. In addition to, techniques for, for analyzing it. Programmatically, when we work with data, we, we think of turning those values of the variables connected to the set of objects into integers and floats or bytes. We may put them into numerical arrays. The string for to put them into character strings, so we have a set of primitives in all of our computer languages for dealing with the values of those data. And then those values we put into more complicated structures to contain notions of how we can structure that data into something that's meaningful at a programmatic level. So we define data structures such as arrays, or maps, lists, sets, collections. We will define queues for certain operations. On larger scales, the more complicated systems we have data structures such as trees and graphs. And the list goes on and on. We do this because it makes our lives easier to have a programmatic representation of what the structure of the data is. So that we can write good computer code against it to make the analysis more efficient. In, in that fashion. But that's programmatically. There is a notion that data should also possibly be structured before it comes into the programmer and into the computer program the programmatic arena. To make it easier as well for purposes of data management. It should be said, that there is no un, unique solution for working with data. You can represent the same chunk of data, through any one of these programmatic structures. there, they will transition or map, from one to the other, quite easily. It depends on the problem being addressed, what you're trying to do with it, and also personal preference. I may favor, an array over a list. Or I may be programming in Python, which has its own particular language preferences for the, for data structure. So there's no unique solution. Similarly when we come to talking about data storage and data access, data management there is again not necessarily a unique solution. It may depend on the problem being addressed and also your personal preference. But broadly I would argue that you can consider data to be regarded as structured. Or, or connected. By structured I, I mean something similar that you could put it into tabular form. A lot of data can be represented in that way. By connected data I mean slightly more loose structure, but you can still associate bits of data together. A good example is Wikipedia. For a lot of entries in Wikipedia there is a set of facts. And for a lot of those entries, those set of facts may be similar. Let us take all the descriptions of famous people. There may be for each of them where they were born, when they were born. Where they died, when they died, who they married, who their parents were. All sorts of things like that, You could imagine tabulating that information and then you, for each one, you have a tabular data structure that you can repeat. But there may be other facts that are associated with that person. Where they went to school, what their area of expertise was, which soccer team they played for. That may not be true for everyone, but you still want to structure those facts together in, in a meaningful way. And in a way that can be utilized. And in that case, you'd be talking about connected data and you would be working with data structure that way. There is a thing called the linked open data movement, which I will talk more about in my final module which works on that sort of basis and is useful for doing web discovery. So let's talk about more of the meat of what we're here to, to, focus on in these modules, which are databases, so I've talked about, having programmatic structures for working with data, and how, with data, we can also have, notions of that there is some sort of structure behind them. So what is a database? A database is a structured collection of data residing on a computer system at can be easily be accessed, managed, and updated. So we're taking that innate structure in our data and we're formalizing it in some way. That is a database model. And there are different database models, depending on how you want to work with your data. I'll talk about that in the next module. But the whole idea is that in the database, data is organized according to the database model. It's structured. And then you have a database management system. A DBMS. Which is a software layer, software package, designed to store and manage databases. So your data store is in a database, and you have the DBMS to work with the database to provide you with the, the software tools, the software interfaces to access the data, to manage the data, to update the data. And do those sort of things. This is a great cartoon from XKCD. It shows some of the dangers of, of databases. Hopefully by the time we finish this set of modules, the punchline will be more obvious to you. Why should I use a DBMS. Why should I bother to learn. Database technology, surely if I structured the database myself, or if I have a good idea of how data is structured on my hard drive, I don't need to bother with the, the necessities of the [INAUDIBLE] necessities of a database management system. Well, a DBMS buys you a number of things. It takes you away from the data in the sense that you do not need to worry about the data being somewhere, the data is independent, just sort of abstraction layer on top of the data. And allows you to consider it in a more abstract way. It is normally optimized for efficient access. I certainly could write a program which would go through a directory structure trying to find the bits of data that meet the type of search or query that I'm trying to do. But because a database management system is optimized for that sort of work it would be a more efficient way. It also allows concurrent access in the sense that there may be many people who want to access the data that I'm after because I'm using a DBMS that may offer a number of client's access to the data that I'm after. DBMS will buy you data integrity, security and safety. This is important. It'll, make sure that you don't do anything silly with your data, like deleting it accidentally. It will make sure it's secure, if you want it to be, so no one can change the values. That only possibly certain people can get access to your data. If it's proprietary data. It allows uniform data administration in the sense that I can put many different data sets into a D.A.B.M.S. but I'm going to use the same interface to manage that data. It means reduced application development time because if I know that my data is sitting in a DBMS such as MySQL then I only need to learn the MySQL interface once to be able to access any data that is sitting in a MySQL database. Might as well be DBMS system anyway. So I don't need to continually rewrite pieces of code factors stating different context, I only need to do it once through the DBMS interface. And finally, it gives me access to data analysis tools. There are a number of. Commercial applications which are build for working with particular DBMS's. Some are better than others, but it means that I can try those on my data. And it doesn't really matter what my data reads, it is just data, numerical non numerical,. And the data analysis tool, which worked for the particular type of numerical data that I can apply and I don't need to worry about building the necessary interfaces or necessarily actually encoding the algorithms either. It's already been done for me by someone, and I can use those tools. So there are a number of advantages, for using a DBMS over just. Putting your data on your hard drive and, and, and doing the analysis way with it. Scales of databases. There's a great quote from Jim Gray, the Turing Award winner for database work. From Microsoft Research and Tony Hey from 2006 where they said that databases and sweet spots for managing data form about 1 GB to about 100 TB. We've probably pushed the envelope at the, the high end on that with some new stuff which I'll be talking about in the sixth module in this series. But, essentially, You can have a very lightweight database management system, SQLite is such an example, it's part of the Mac OS operating system automatically, so you've got a database there, a relational database that you could start with. If you want to move to something slightly larger scale. MySQL, I've already mentioned, PostgreSQL the other alternative, both from open source both widely used for free if you need something slightly more enterprise scale,. SQLServer from Microsoft, or Oracle, IBM all have their own solutions as well. Perfectly usable. And then for moving to the Web-scale, there are things like Hive and Hadoop and then alternative non-relational technologies like SciDB, and Redis, key-value stores. MonetDB, which is column-oriented database. NuoDB which is a new SQL database. And I'll be talking about, those again in, in the sixth module. So, there is a, a range of technology solutions out there for you depending on. How much data you have and, and how much money you have in certain circumstances. That certainly that whatever your needs there is, a, a database solution out there for you. And I'll leave it there for the moment