Hello, my name is Matthew Graham, and this is the first of two modules on semantics. You may remember from the final module of the database set that we talked very briefly about database structures for sophisticated, representations of data. In the forms of things called RDF and also ontology-based databases, and this set of two models will expand on that to some degree. The idea behind semantics and the semantic web is that it's a set of technologies and methodologies for representing knowledge in a machine processable manner. It's envisaged and has been for at least the last decade or so, very much as the future of the Internet. If you regard the first version, the first generation of the Internet as a set of connected web pages or some sort of glorified online library. The next generation would be a, a connection of data collections much richer than just pres printed material. Which would allow you to do arbitrary queries, arbitrary analyses with them in a connected fashion. The idea is that there's this web of data, web of databases, a decentralized platform for distributed knowledge. So you just go into the semantic web to collate your information and knowledge and bring that back together. It's predicated on a set of logical pieces of meaning that can be mechanically manipulated. That means that there is a computational representation of not just data and information, but also knowledge, the main knowledge, ideas, concepts, how they relate to eachother. And in such a way that at a computer can process them and make decisions based on them. It encompasses vocabularies that can be used for making assertions about things. In a way that you can actually check for consistency between meaningful statements, conceptual statements. And also infer new pieces of information or new pieces of knowledge based on representations of that. And this leads to the notion of smart applications. Applications which actually have some idea of the main knowledge encoded into them and allow you to do sophisticated things with them. An example of that would be let's say we have a data collection of various bits of data related to zebra fish. We have images, we have the results of genomic array experiments, so we have bits of DNA. We have other types of experiments in imaging with that, we have this all in our large data collection. But we have arranged it in such a way that each piece of individual data is tagged with a semantic descriptor which tells us which part of the anatomy it relates to. We also have a conceptual scheme encoded into our smart application which describes how those individual anatomical descriptors relate to each other. Both in terms of an anatomical structure, but also maybe in terms of how the the zebra fish develops with time. So that this particular structure at this particular stage of development, when it's an embryo, becomes this particular structure when it's a juvenile. Which becomes this particular type of structure when it's mature. What this allows us to do for example, is that we could, for example, find all the data that's relevant to a particular anatomical structure, say the hindbrain. But because we have this conceptual description behind as well, we can actually infer on this, and do broader and narrower searches. So, we could identify the central nervous system, which the hindbrain is part of. We could also identify those individual anatomical structures which comprise the hindbrain and then use those additional search terms to, to widen our search and make a smarter search. Other types of referencing we could do on that would be also to do particular stages of development that are related to each other. Or metabolic processes or molecular information depending on different levels we have in our conceptual schemes. So these sort of smart apps are much more, we, we already have the data stored in a relational database but then we have an additional structure on top which is doing knowledge management. As it's called, and this is very much the regime of, of semantics and semantic technology. So the fundamental basis for this is that we regard knowledge as a, as a graph structure. Information or knowledge is, is both, is best expressed in this idea as a labelled, directed graph. This is the entity attribute value data model which you will remember we covered in one of the database modules. So you can refer back to that for a description, but the idea is that we have essentially triplets, which represent pieces or facts of information, facts of knowledge. So in this particular graph that we are showing here, this is a way of representing something which has a, a, a arbitrary name at the moment, _1. But it has a name, lanthanum, so identified as having something which has a name, lanthanum. There is a property called has_Atomic_Number, which has a value, 57. This thing has a property called has_Color, silvery white. Then we have another entity called _2. That has a name, samarium, but it also has a color, silvery white. So we have, a set of, of entities of subjects. We have a set of entities or objects, and we have a set of properties relating an object and a subject together. And in this way, we can, we can build up arbitrary graphs of, of knowledge. And the way we do this, the underlying technology is the resource description framework, RDF, which we covered briefly in database module number 6. This is a W3C standard for encoding knowledge. That's the World Wide Web Consortium the same body that endorses HTML and all the other standards for the web. So the idea with their, this RDF thing is that a fact is expressed as what's called a subject-predicate-object triple or statement. So in our example in the graph we just showed one of those triples, as they're known, would be _1, that's the subject. Predicate has atomic number, that's identifying some property of it or some statement about it. And then a value 57. The subjects, predicates and objects are given as names for those specific entities in the way that RDF works, each of those names is a URI very similar to a URL. Objects, the the, the, the, the thing that, you know, the subject-predicate, the, the object bit can also have text values, which are called literal values. And they can also be data typed if you want to, so you could say that this thing is actually an integer or this is a string, or it's an array. Something like that for programmatic ease. There are various different ways of representing RDF. Depending on how programmatically you want to do it. The top one here is n3 or turtle. Which is a fairly succinct freeform expression. You define a name space identifying, maybe, the domain regime that you're using to carry your information. And then we have, within the first angle bracket, is our subject. This is the thing that we're saying is lanthanum, La. Then we have a predicate, which in this case is pt:name, and then we have an object which is the lanthanum encased in quotes. A semicolon is used to say that we're continuing we're going to add another predicate object to the same subject. In this case we have atomic number 57, and then another predicate object, which is color silvery white, and then full stop to finish our statement. So that is the representation of the graph structure for lanthanum and n3/turtle version of RDF. There's also an XML version of it, which is maybe more programmatically easy to manipulate if you are already familiar with working with XML. And in this particular case, that same n3/turtle representation that's shown there is then expressed here in maybe a slightly more structured and easily readable XML form. You can see, again, we have our subject and then we have a set of predicate object pairs. There's also another technology, which is quite interesting, called RDFA which allows you to put these RDF statements into basic webpages. The reason you might want to do this, is that you could have a webpage which describes something to a human in terms of text and images and such like. But also it has this structured information embedded in it so you could then pass the same webpage to a piece of code. It can extract this information on it and then make use of that structured information for programmatic purposes. So you can have a page which is both for human consumption and machine consumption with the information encoded in both cases in a single go. Instead of having a separate one for the machine, and a separate one for the humans. Now, one reason you might want to have, one thing you can do with all this data when it's out there is this idea of linked data which is the sort of idea of what the semantic web is all about. This was a, a term coined by Tim Berners-Lee, the founder of the Internet. And he at-length, outlined four principles for linked data. First is that you use URIs to identify things that you expose to the web as resources. Well we've already seen that's how RDF works. There's subjects, predicates, and objects are, are mainly identified by URIs. Hopefully you're using HTTP URIs so that everything just works with, with a web address. Instead of having some strange thing where you need to figure out what the beginning of it is, if it's FTP or Gopher or some other obscure system. You provide useful information about the resource when it's URI is dereferenced. What that means is that when you go to that URL, there is a webpage which describes what that resource or whatever the subject, the object, the predicate is actually about. What that URI is being used as a shorthand to, and to also use links to other related URIs in the data you're exposing. And based on those principles, a whole web of information has built up over the last six or seven years called the linked data web. This connects many different regimes right at the heart of it is a thing called DBpedia. We're all familiar with Wikipedia, now when you look at a Wikipedia page, you often see a little section on the right hand which has in a little box and seems to have structured information. If it's a country, it might always list the, the capital, the population, the current ruler, that sort of thing. If it's a famous person, it'll list when they were born and where they were born, when they died, and that sort of information. And there is a formal structure for a lot of the entities in Wikipedia to capture that sort of information. That information can be extracted from Wikipedia and that's what forms the basis of DBpedia. That information is captured in DBpedia in a set of RDF triplets for all of those semi-structured informations you find in the Wikipedia pages in that right-hand side. And that, that body of connected information, of linked information, is the center of the linked data web, and then that links out to other data collections in the web. Film titles or sport scores or genomic or biomedical publication information. Chemical analyses, that are expressed in these similar forms and linked together through these web mechanisms. And the sorts of things that this link data allows you to do is you can ask sophisticated questions of it through particular mechanisms. For example, you could say, well, I want a list of all episodes of the television series Breaking Bad which are ordered by their air date. That's a query you can ask of the linked data web and it will give you that information back because those triplets link to each other and that information is in there. Or you could find the official websites of companies with more than 50,000 employees. Or you could say, find me things close to the Eiffel Tower. One of the main hubs in the linked data web is a set of geopositions, and so the Eiffel Tower subject or object can be resolved into something which has a geospatial location. And then you can compare that to other geospatial locations and find those things which are near to the Eiffel Tower and then, then render that in a list. Or, as I said, you could do things like discover new drugs to treats Alzheimer's. So you could say, what proteins are involved in signals transduction and are related to pyramidal neurons? Now, if you an experiment was done of blindly asking Google this, and it got 223,000 hits from the Google search engine, and, and none of those were real, valid results. But when you ask the same query of linked healthcare data, you get 32 hits and each of those hits is a successful result. It is a protein that's involved in signal transduction and related to pyramidal neurons which helps you do that sort of thing. For this shows the power of actually linking this information together in a, the main knowledge-based meaningful way, in a semantically meaningful way. And that is the power of this, this, this sort of technology. And that's where we will, will, well we'll finish by some of the tools that you can use. If you want to look at linked web data in your browser you can use the Tabulator tool there are specific browsers for doing this, Disco or the Openlink Data browser. There are libraries which you can use to encode and work with this information, there's this old one called the Semantic Web Client library in Java, there are more modern versions and others. If you want to, if you have a relational database that you want to expose in, easily into the linked data web, there is also a tool called d2r that you can use for doing this. It's essentially a, a set of mappings that you have to provide which will translate the database schema you have into something that can be understood in terms of a conceptual scheme. And we'll cover those in the next module. And so that's where we'll leave this module