Skip to content

Interacting with RDF Data via DM's Javascript Libraries

Tim Andres edited this page Mar 12, 2014 · 1 revision

The Databroker

sc.data.Databroker is the center of any interaction with RDF data made using Javascript. The databroker owns an instance of an sc.data.Quadstore, which indexes all of the quads (triples with a context, i.e., named graph) and has methods for fast queries using wildcards. While you could interact with the QuadStore directly to do querying, the Databroker provides easier methods of interacting with data through the getResource() method. It also abstracts away querying servers for more data by returning deferred resource objects, and maintains indexed stores of modifications to RDF data made since the last sync with a server.

The Databroker is designed to use ore:isDescribedBy triples whenever possible for finding resources. To take advantage of this functionality, you need to first prime the Databroker with some initial data on where to find resources. To avoid cross-site AJAX request issues, you can provide the Databroker with a proxy url function (which takes a given url and returns a url for requesting that resource from a proxy server) at instantiation, but by default, the Databroker will simply try to make the request. To this end, the databroker has a method called fetchRdf, which takes a url as its first parameter, and an optional callback function as its second parameter. Alternatively, you could manually add the ore:isDescribedBy triples without making a request for serialized triples in this way. Once this data is loaded, you can then request further data about resources for which the Databroker has ore:isDescribedBy information through the deferred resource interface.

Resources

Although RDF is defined in triples (and quads when named graphs are used), this isn't always the most useful way to think about the data when processing its semantics. The sc.data.Resource class used by the Databroker lets you think in terms of how these triples describe given resources (identified by their uris), much like the nested structure of XML and Turtle representations of RDF. As an added benefit, they also understand the semantics of the owl:sameAs predicate.

The basics

When you call databroker.getResource(uri), where databroker is an instance of sc.data.Databroker and uri is the uri of some resource you want to work with, the databroker returns an instance of an sc.data.Resource object. The only real data stored in this object is its uri, and a reference back to the original data broker, and a reference to a sc. data.Graph object with the current context the databroker is using. You can also instantiate a resource by calling its constructor with a reference to a different graph if necessary. When you ask the resource object for data, it always refers back to its graph object, meaning that it always stays up to date.

Let's use a text document as an example.

<http://example.org/TextA>
    dc:title "Text A" ;
    rdf:type
        dctypes:Text ,
        cnt:ContentAsText ;
    cnt:chars "Lorem ipsum dolor..." .

To access this data in javascript, you can use a Resource object like this:

var resource = databroker.getResource('http://example.org/TextA');

var title = resource.getOneProperty('dc:title'); // Returns "Text A"
var types = resource.getProperties('rdf:type'); // Returns a list of the types with fully qualified uris (rather than prefixes)

resource.hasType('dctypes:Text'); // Returns true
resource.hasType('dctypes:Image'); // Returns false

resource.toString(); // Returns a Turtle representation of the resource like the above

Since all data stored in quads uses strings formatted like Turtle syntax (e.g., uris are wrapped with <>, literals are wrapped with "" and optionally followed by <XMLDatatype>), Resource objects try to return data in the way you'd reasonably expect to use it. For example, if you're requesting a property that is a uri, the Resource will return it without the wrapping <>. Likewise, if you're requesting a property that is a literal, the Resource will return it without the wrapping "" or datatype information. If you need this information, you should use the unescaped properties methods - getOneUnescapedProperty() and getUnescapedProperties(). In practice, it's much easier to let the Resource strip away the formatting information for you (provided you're working with a well defined data model).

You'll also notice that you can refer to properties using Turtle and XML like namespace syntax in a raw string. These strings are automatically expanded to a full uri using the prefixes defined in the Databroker's namespaces object. You could also use a fully qualified uri wrapped with <> instead.

Traversing the graph

Of course, the real power of RDF is in its ability to connect resources.

Let's add to the RDF data from the previous example:

<http://example.org/AnnoA>
    rdf:type oa:Annotation ;
    oa:hasBody <http://example.org/TextA> ;
    oa:hasTarget <http://example.org/TextB> .

Now there is an annotation which has our original text as a body. We can discover these relations using Resource objects.

var anno = databroker.getResource('http://example.org/AnnoA');

// You can get the uri of the body like this
var bodyUri = anno.getOneProperty('oa:hasBody');
// but you can also get the body as a resource directly
var body = anno.getOneResourceByProperty('oa:hasBody');

// Likewise, you can walk back up the chain.
// This finds resources which reference body by the oa:hasBody predicate
var bodyAnnos = body.getReferencingResources('oa:hasBody'); // Returns a list containing our original anno

Deferred Resources

What if you need to ask for more data from somewhere else on the web? The Databroker provides a way of abstracting away the need to search for ore:isDescribedBy triples and generate AJAX requests. You can simply request a deferred resource, which uses the jQuery Deferred interface. You can then bind handlers to the progress and done events of the deferred resource, and work with the downloaded data asynchronously.

Example:

var deferred = databroker.getDeferredResource(uri);
deferred.progress(function(resource) {
    // This gets called every time a file is loaded
    // Check if the properties you need are now part of the resource, and if they are, do something with them
    
    if (resource.getOneProperty('dc:title') != null) {
        console.log('title', resource.getOneProperty('dc:title'));
    }
});
deferred.done(function(resource) {
    // Data should be loaded by now
    // Do something with the data here
});
deferred.fail(function(resource) {
    // The resource had ore:isDescribedBy triples, but none of the files could be loaded
});

If you already have a Resource object, but would still like to query the web for updates, you can call the getDeferred() method of the resource as a shortcut to return a deferred resource with the same uri. Because Resource objects always get their data from the Databroker at the time of the request, it doesn't matter whether you use your original Resource object or the one passed as a parameter to your callback functions, as long as the query is performed after the Databroker has received the new data from the web.

Creating New Resources

TODO

Modifying resources

Thanks to their tight link with the Databroker, Resource objects also allow you to modify resources while keeping track of the necessary changes to send to a back-end server.

It's important to note that unlike the getter methods of Resources, the add, delete, and set methods require you to be explicit about datatypes. This means that you need to wrap uris in angle brackets, and literals in quotes (don't forget to escape inline quotes); however, prefix namespace notation is still supported, and sc.data.Resource objects provided as values will automatically be converted into angle bracket wrapped uris. Utility functions for this sort of wrapping exist in sc.util.Namespaces.

You can add triples about a resource by calling the addProperty(predicate, object) on a Resource object. This adds the new triple to the Databroker's main QuadStore, and it also adds it to a store of new triples for the purpose of synchronizing the data with a server.

RDF allows for a resource to have many possible values for one property, like multiple values for rdf:type. However, many data models specify that certain properties only have one value. In a case like this, you can use the setProperty(predicate, object) method. This will delete any and all values for that given property, and then add only the one specified.

You can also delete properties using Resource objects. To delete a given predicate object pair for a resource, call the deleteProperty(predicate, object) method. To delete all values for a given property, simply omit the object parameter to have deleteProperty(predicate). You can also delete all properties of a given resource by calling deleteAllProperties() on that resource, effectively deleting all triples with that resource as their subject. All of these deleted triples are also stored in the Databroker, meaning that they can also be synchronized with a server.

Clone this wiki locally