Hello, my name is Matthew Graham and this is Module Two of two modules on Semantics. In the last module we discussed how to represent knowledge in a machine processable way. We talked about the various different representations of RDF, the technology that can be used to do this. How to embed this information into a web page, and also the linked data idea that makes use of this to create a, a, web of linked information that's accessible through browsers. In this module we're going to be considering how to actually represent domain knowledge in terms of machine processable form through concept schemes and ontologies. So we have a hierarchy of concept schemes that can be used to capture domain knowledge. The idea is in its simplest form we have a controlled vocabulary. So this can just be a closed list of terms that can be used for classification. So it can just be a certain list of words or phrases that we are using for a set of labels, and we will always make use of that closed list. But there's no notion of any sort of relationship between them. There's no notion of a hierarchy of definition within them. So the next level up from a controlled vocabulary, which can just really regard as, as a bag of terms, is what's called a taxonomy and then a taxonomy, you have a controlled vocabulary but you have imposed a hierarchy on it in terms of term and, and sub term. So for example, eh, when we were talking about anatomical structure in the last lecture. We had this notion of looking up the word Hindbrain and then their maybe in our, taxonomy that we're using, a, a, a super term of that would be central nervous system and a sub term of that would be an individual structure that makes up hindbrain. So these hierarchies of, meaning don't necessarily you need to be only in terms of obvious meaning, they can also be in terms of other hierarchical notions whether it's anatomical development or geographical location or things like that. About the taxonomy we then have this more advanced concept of a thesaurus and in a thesaurus we have a taxonomy but there are then notions of not only hierarchical relationships between the individual terms, there are notions that you have broader and narrower terms, that you can have synonymous terms. So these are terms that you can use at a, at an equal level in a hierarchy instead of each other. There is some notion of a top term a, a root term. There is a notion of a scope note, to, to determine that, you know, this is when this particular term is relevant, this is when this particular term is relevant and then there are also notions of related terms. So, that this term is related to this term over in this part of the hierarchy, or this term is related to this term in this other thesaurus. Other ideas or concept schemes that you, you could use when you're, you're programatically manipulating knowledge and information, also things like subject heading list, or terminology, glossaries, faceted classifications. So the idea is that you would employ one of these when you are working with a knowledge base and, and trying to encapture a domain knowledge into a, a programmatic fashion. Now the way you would represent these sort of simple concept schemes is through yet another W3C standard. This one is known as SKOS which stands for the simple knowledge organization system. It's used for expressing knowledge concept schemes in a machine-understandable way, and it's, is itself expressed in terms of RDF statements and a thing called RDFS, which is a, a simple schema language for RDF, which can be used for defining RDF data structures. And the idea there is that you express your entire concept scheme as an RDF graph with content and structure. And there are various words there you can see which identify specific terms within, within the SKOS system relating to things that we've identified in the particular type of data structure, particular type of knowledge structure we may be interested in. So we have scopeNotes and narrower, broader, related, hiddenLabels. Whether something's an OrderedCollection, it's a ConceptScheme that sort of thing preferred labels. So the idea is that I could take one of my vocabularies or taxonomy that I"m using and I can then write it up formally in the form of a, a SKOS vocabulary, and then that there are programmatic tools which I can then use that to ingest. And I could then start programmatically manipulating RDF triples, which are expressed in terms of that SKOS vocabulary, such that they may be the subject or mainly the predicate are terms which are relevant or that are parts of these SKOS vocabularies. And here's an example of, of SKOS, of of of a skos vocabulary. This is actually defining a temperature scale, and it's saying that I'm listing a whole load of different temperature scales here. But the one I'm particularly describing in this example is the absolute temperature scale. I'm saying that this has a preferred label, which is abso, absolute temperature scale in lower case. The definition in my provider, even readable definition here. There's a broader term of it in terms of this concept which is temperature scales, there is a narrower term which is Kelvin, and then we have related, concepts in our scheme to this temperature, just because the absolute temperature scale which is the Celsius temperature scale, and another related concept is, is temperature itself. So we could define an entire skos vocabulary devoted to the physical properties of temperature and measuring temperature. Temperature scales, how that information is encoded, how those different systems are related to each other. And then we could have, a piece of code, which could use that to programatically manipulate, temperature readings for climate data models for example. There is, a query language for RDF called sparql. Again, it's a W3 standard. I'm not going to go into it in any great detail here, just to show you though that there are particular keywords in there which will look familiar to the sorts of keywords that we covered when, in the two modules we did on SQL in the databases. So you will see select and you will see aware and order by, distinct, limit, offset, that sort of thing. This is an example of a Sparql query. In sparql variables are identified by question marks at the start of it. What we're doing here is here is we are, querying a database of elements, and we're asking for the, the name, the chemical symbol, the weight, the atomic weight, the atomic number, and the color from our data collection of linked data or whatever expressed in a machine-readable way. And then we're saying where we have to think we're looking specifically for uranium. And this is the way we would identify the things which are, are transuranic and heavier than uranium and then we order by, them by weight. So this is an example of this sort of thing. We could write this as a similar example if this data was expressed in in a relational database form with the SQL statement. The advantage of sparql however is that the idea is that the data doesn't necessarily reside in a single database. It already has this notion that the data is all linked together out there somewhere, so there's already a notion when you're, you're constructing a sparkle query that it's going to be a distributed query hitting multiple data sources where the data is identifiable and, and resided. And that's given by those URIs in the RDF statements that you're working with. Now, the most elaborate, or the most powerful concept scheme that you can use, is a thing called an ontology. And the formal definition of ontology is, is that a formal specification of a shared conceptualization. What that means is that you have an agreed set of ideas of what the domain knowledge base for a particular de, subdomain or domain is, the particular knowledge base for that is, and that's how you are expressing it in a machine processable way and that's called an ontology. So it is essentially a data model that represents a set of concepts within the domain and all the relationships between those concepts. An ontology allows you to define arbitrary relationships, arbitrary properties, arbitrary concepts so you're not limited to the types of the relationships you can have to just hierarchical or whatever. You can have any type of relationship and you can identify those and, and encode those. So ontology is generally describe individuals which are the ground level objects, the, the individual facts that you may be trying to represent. They will describe classes and if you're familiar with object oriented programming then, then that should make some degree of sense for you to have. You have classes and instances in object oriented programming and in ontology you have individuals in classes. You have attributes that the classes and the individuals have, properties, features, characteristics, those sort of things and then you have relationships which you you use to associate all of those together. And then there can be events which define how those attributes or relationships can change in pre-described ways. Yet again, there is a W3C standard for authoring ontologists for representing ontologists. This is the web ontology language or OWL, as its called. It's again, based on RDF so an ontology is expressed in terms of RDF structures. And this is largely regarded as one of the fundamental technologies underpinning the semantic web because it gives you the richness of expression to offer arbitrarily encode any set of concepts, any piece of domain knowledge into a machine processible fashion. As I said OWL allows descriptions of relations between classes particularly disjointedness and that sort of thing. It allows you to express cardinality constraints on particular things, so you have exactly one. Characteristic of properties and numerator classes, and all sorts of, of, very powerful things. So, owl regards data as being interpreted as a set of individuals which it describes a set of property assertions relating the individuals to each other and a set of axioms placing constraints on, on sets of individuals, on, on classes and the types of relationships allowed between them. So, for example, we may have a family ontology which is defining this notion of what a family is and the concepts that that contains that may have for, example, a, a hasMother relation which is only present between two individuals when one of them, has a parent, is also present. And in our notion of family ontology, we may also have this relationship called HasTypeOBlood. And members of the class of that relationship, HasTypeOBlood, are never related via the hasParent constraint two members of HasTypeABBlood. So in that particular example, we can put biological constraints on our notion of family into this knowledge base and have that as a machine processable form. So it's showing that these are not just simple relationships, there are, there are ways we can encode any sort of information we want into something like that. We can also, make inferences, or say that if Ada has mother Ann, and Ann is, has type O blood, then Ann is not a member of the HasTypeABlood relationship. And one of the powers of things that are represented through ontology is particularly is that we can use logical inferencing to identify statements that are not consistent with the knowledge base that we're, we're using. This is done but using reasoning surfaces, services. This is done using reasoning services. And, what's called first order predicate logic and, the idea is because we've expressed our, our, our knowledge base, in a for most, and rigorous way then we can use mathematical logic to either make further inferences or look for statements that are in congruence with each other. So we can say that sheep only eat grass. We can say that grass is a plant. We can identify plants and parts of plants as disjoint from animals and parts of animals. We can say vegetarians only eat things which are not animals or parts of animals and then having encoded those statements up we can ask our logical inferencing action to make some inferences and it can come up with that statement that sheep are vegetarians. Which is a logically, consistent statement based on the, the set of statements we've already given it, a set of instances we've already given it, plus the relationship rules, and the properties we've defined between those. So that gives a little toy example of the kind of inferencing that you can do. Some of different things examples on one of our smart applications might use. If you remember the, the set of zebra fish, meta data, multimedias data that we talked about as an example in, in the last module. Then we may have some inferences that we're going to run, we have an ontology behind it, we have our data marked up in terms of that ontology, and so we may want to look for and atomical terms which are relevent for a particular stage of development. Or we may want to search for terms which are related to a particular term in some way, from some class whatever. Or we may want to find, two classes, and we're looking for the lowest common ancestor in the hierarchy, because we want to use that as a single search term which would then encompass both terms, in, in a hierarchical fashion that we're interested in doing. And we can use logical inferencing for all finding all those sort of problems. Finally, software that you may want to look at for working with ontologies analysis, there's Jena and Protege- I would highly recommend looking at Protege. It allows you to explore ontologies, it allows you construct your own ontologies, it allows you to reason over ontologies. It's, it's a very powerful tool and then there's a, a, an instance of an infrontention here which is come at, at the bottom, CWM, which you may also want to have a look at. There are other like pellet and hermit which have I think both Java and Python wrappers that you can use in, in your code base as well if you are working with RDF and ontological information. And that is the end of the module. Thank you.