SyntaxHighlighter

SyntaxHighlighter

Tuesday, November 16, 2010

Adding Foreign Namespace Support to the NITF XML Schema

As previously discussed on this blog, NITF has very limited support for foreign namespaces.  I've been experimenting with ways to remedy this and presented the results at the most recent face-to-face IPTC meeting (in Rome).  I have posted the NITF slides on slideshare.

During the meeting, it was decided to break the NITF 4.0 effort into two:

  • NITF 3.6 release, which would directly address the addition of foreign namespaces and could be approved as early as January 2011
  • NITF 4.0 which would add support for G2 features (such as qcodes) and will likely require significantly more work to develop and approve
It was decided to take this approach, since foreign namespace support addresses pressing needs (in fact, many people do not realize that it isn't legal already to mix and match NITF with other XML schema).  It was also felt that this change is relatively minor and so doesn't merit a major version number change (not sure I agree with that, but ...).


Therefore, I have now created an NITF 3.6 set of schema files, fixing various bugs in the previous experimental schema and adding expansion points that were missing.  So, compared with NITF 3.5, the experimental schema

  • Added any attributes in globalNITFAttributes and commonNITFAttributes
  • Added any element into head, body, docdata, body.head, block, enriched text, after body, media, body.end

I've also versioned the ruby files to 3.6 and altered the comments to discuss NITF 3.6, rather than NITF 3.5. Finally, I've created an NITF instance that exercises the various foreign namespace capabilities.

http://groups.yahoo.com/group/nitf/files/schema/nitf-3-6.xsd
http://groups.yahoo.com/group/nitf/files/schema/nitf-3-6-ruby-include.xsd
http://groups.yahoo.com/group/nitf/files/schema/nitfastronautNSspan.xml

These files are also available via the NITF 3.6 directory on the IPTC website.


Comments?  Questions?  Critiques?

Monday, November 15, 2010

IPTC and Semantic Web Technologies - Linked Data, Metadata and Ontology

At the IPTC's most recent face-to-face meeting in Rome, we reviewed our explorations of semantic web technologies for news. The news standards body has been looking at three major areas:


Linked Data
We discussed our work to turn IPTC's subject codes into Linked Data using SKOS concepts and Dublin Core properties. Michael Steidl (Managing Director of the IPTC) was planning to demo the IPTC Linked Data ... but, sadly, Internet access was not working in the hotel! He was, however, able to discuss the proposed collaboration between the IPTC and MINDS on Linked Data for news.

Linking and Mapping

Much of the discussion about IPTC's Linked Data work turned on the difficulties of mapping. In addition to representing the IPTC subject codes in RDF/XML and RDF/Turtle, there was some work done to map from the 17 top level IPTC terms to dbpedia concepts. We quickly figured out that these top level terms are chiefly umbrella terms and so don't map very well to individual dbpedia concepts. The meeting felt that it would be good to map the second level terms, but the problem is that this is quite a lot of work and - as usual - it isn't clear who will do it! We then explored some of the challenges of creating and maintaining the links in Linked Data - that is where a lot of the value, but also much of the investment, lies.

My slides about IPTC's Linked Data work are available on slideshare:

Metadata in HTML - rNews and hNews
Many news providers have created feeds to supply news using IPTC formats such as NITF and NewsML-G2. However, there are an increasing number of consumers of news who only want to work with "pure" web technologies, i.e. HTML rather than XML. So, the IPTC has been looking at the two major paths to represent metadata in HTML - microformats and RDFa.

hNews
I discussed hNews - the microformat for news that was adopted by the community in late 2009 - which builds upon hAtom by adding a few news-specific fields (such as Source and Dateline). As well as explaining how to add microformats to your HTML templates, I provided some statistics that the Associated Press has gathered on adoption. (As of October 2010, we know of about 1,200 sites using hNews, predominantly in North America). See my Prezi on hNews at http://prezi.com/uo5ggdkll8sa/an-introduction-to-hnews/ for more.

rNews
Evan Sandhaus (Semantic Technologist at The New York Times) described rNews - a proposal for an RDFa vocabulary for news. As the names imply, rNews and hNews are similar in intent (news-specific metadata in HTML) but somewhat different in approach. Whereas hNews went through the microformats process, an RDFa vocabulary can be created by anyone. Evan has created an initial rNews draft based somewhat on the NewsML-G2, NITF and hNews models but it is clearly heavily influenced by the needs of the New York Times.

Members of the IPTC's Semantic Web Yahoo! Group can view Evan's rNews draft and are encouraged to discuss it in that email group. At the Rome face-to-face meeting there was quite a lot of interest, but also several issues raised about the details of the first draft. The meeting generally agreed to continue looking at both hNews and rNews, with a view to making a recommendation on both in 2011.

The benefit of getting rNews and hNews adopted by the IPTC is that greater industry support translates into less work for toolmakers: if many news providers support hNews and/or rNews - and do so in very similar ways - then it is easier to build parsers and tools to extract metadata from HTML.

News Ontology
Benoît Sergent of the European Broadcasting Union discussed the work that he and his colleague Jean-Pierre Evain have been doing to create a news ontology, based upon the NewsML-G2 news model. Benoît described how EBU would like to combine the video content that it produces with content from its member organizations and other third parties. If they can represent this information using a flexible, universal model (the news ontology) they could use off-the-shelf tools (such as a triple store) to query, manipulate and recombine that content.

In many ways, this is the most fundamental piece of the semantic web work that the IPTC is undertaking. It is also the least accessible, for many. Members of the IPTC SemWeb group can view a draft of the news ontology and can comment in that email list.

Thursday, September 2, 2010

Expanding Beyond the Point - Geo Geometries and Features

When adding "geo data" to your content, it is tempting to think you just need to add a latitude and longitude and you're done.  After all, this allows you to plot your data on a map, by applying a push pin or the like.
what's the point?
http://www.flickr.com/photos/darren/35248684/

However, that lat+long pair may indicate the centre of a location in your content, but it doesn't reveal anything about the scale or type of the location.  There are two ways you can do this.  One is to make use of a geometry, the other is to identify the feature type.  And you can use both together.

Geo Geometries

Geometry at the National Mosque
http://www.flickr.com/photos/swamibu/2527486134/
If you look at the various geo standards, you'll see that it is common to use different types of geometry for geographic data.  For example, if you look at http://www.georss.org/simple you'll see

Point
Line
Polygon
Box
Circle

The polygon, box and circle all express areas.  The needs of your application will likely drive which one might be most appropriate.  For example, if you want to display a map at the right "zoom level", it is likely that a point (aka the centroid) and a box (aka the bounding box) are enough.  (In which case, you'll need three points to express them - one for the centroid, two for the corners of the bounding box.  Obviously, the other geometries have different characteristics - you only need the centroid and a radius to express a circle, whereas a polygon must have at least three points, but can have many more).

Geo Features

Pigeon Point / Sky Whale...
http://www.flickr.com/photos/nzdave/303236483/
On the other hand, it could be that what you're aiming to do is to describe what is known as the "geo feature", rather than or in addition to the geometry of an area.  In other words, how is the area classified - is it a city, a country, a park, a farm, a forest, a lake ...?  If so, then there are some geo ontologies that might help.  A popular one is the geonames ontology, which breaks down geo features into subtypes of administrative boundaries, hydrographic, area, populated place, road / railroad, spot, hypsographic, undersea and vegetation http://www.geonames.org/statistics/total.html

Wednesday, July 21, 2010

Linked Data and the World Cup

In January of 2010, I attended the News Linked Data Summit. This was a collection of several organizations involved in news production and distribution, looking to see if there was a way to collaborate on moving forward the Semantic Web (and particularly Linked Data) for news. It was a very interesting discussion, lead mainly by the BBC and The Guardian. At the end of the day, the group decided to move forward together by producing Linked Data for the UK Election. (In January, the election was clearly going to happen in the near future, but not yet announced - it wound up happening in May 2010).

I had a counter proposal for a Linked Data experiment - the World Cup. This made more sense for my employer and seemed to be of interest to at least some others. It was not to be, however. (I later spoke to some folks about whatever happened with the UK Election Linked Data experiment. As far as I can tell, the Guardian did produce some election data. But it isn't clear to me that this turned into anything bigger).

By Shine 2010 http://www.flickr.com/photos/shine2010/3292491879/


However, it turns out that the BBC did use the Semantic Web to power their World Cup “microsite” (actually their World Cup site has more pages than the non-SemWeb Sports section). In "The World Cup and a call to action Around Linked Data", BBC Architect John O'Donovan gives an overview of how they used fine-grained metadata to be able to produce their site with far less need for editorial curation of individual pages. In an accompanying piece, Jem Rayfield describes the technical details of the "dynamic semantic publishing".

Although their descriptions are couched in the terms of the Semantic Web, much of what they describe is more to do with the application of fine-grained metadata than with the particular data formats they use. They do describe how they make use of certain Semantic Web technologies - such as an RDF Triplesore / SPARQL system - to apply derived properties. And it is almost certainly the case that the use of the RDF model has given them more flexibility than can be the case when you use RDBMS systems - or even XML-based schema. However, fundamentally, what they are describing is the potential for what could be done via the application of metadata to content. It is interesting to see it in action.


The England Team didn't do so well at the World Cup - thanks Doug88888! http://www.flickr.com/photos/doug88888/4550561194/

It is also interesting how much of a revelation this is to people. (The first article was widely distributed via Twitter. And the comments are quite breathless in their admiration).

Hopefully, the work that the IPTC is doing on Linked Data and the Semantic Web will help other news organizations (including the AP!) unlock some of the great metadata work that is going on behind the scenes...

Monday, May 17, 2010

Towards NITF 4.0 - Experimental Support for "Foreign" Namespaces

One of the critiques that has been leveled against NITF for many years is that it cannot be customized by including "other" XML namespaces [1]. In NITF 3.5, we completed the step taken towards opening up the NITF 3.4 schema by fixing a bug in namespace support in enriched text [2].

However, we decided that full support for foreign namespaces was such a big change, that this would constitute one of the major pieces of work for NITF 4.0 [3]. I've created an *experimental* NITF XSD with foreign namespace support. It can be downloaded from

http://groups.yahoo.com/group/nitf/files/schema/

I've performed a number of tests with this schema and it seems to me to be going in the right direction. I've also discovered a problem with the NITF 3.5 namespace support, which I've fixed in this experimental XSD [4]. I promise to write a future note about the choices I made in adding foreign namespace support. However, I wanted to get the current version out there, to give people a chance to download and try it out [5].

--
[1] See for example, the excellent discussion of schemas and extensibility by Bob DuCharme and how NITF is too closed http://snee.com/xml/xml2005/industryschemas.html#d50e406

[2] Get the NITF schema at http://www.iptc.org/std/NITF/3.5/specification/

[3] Discussion of the plans for NITF 4.0 (amongst other things) can be reviewed in http://www.slideshare.net/smyles/nitf-2010-spring-working-group

[4] A special prize for the first person to figure out what the bug was

[5] Note that the NITF 4.0 experimental schema contains the documentation I copied over from the NITF 3.5 DTD. I'm still keen to get feedback on this too!

Friday, May 14, 2010

SKOS and Protoge HOWTO

In the IPTC, we are doing some work to figure out how to represent the IPTC Controlled Vocabularies as Linked Data.  We've decided to use SKOS as the RDF Vocabulary.  One of the things we wanted to do was to use a tool that "understands" SKOS.  We decided to look at Protoge for this.  Here are the steps we figured out to make it work (for some values of work):

Download version 4.0.2 from http://protege.stanford.edu/download/registered.html
-          Install it in your PC
-          Add the SKOSed plugin (use Check for plugins... item under File)
-          Add the Pellet Reasoner plugin
-          (you have to restart Protege before the new plugins are active)
-          Add Views to the Individuals tab: SKOSed view -> Inferred Concept Hierarchy + SKOS Usage
-          Then you may load a SKOS vocabulary – but only with narrower and broader relationships, does not work with the ...Transitive variants.
-          Then you should run the Pellet Reasoner against this vocabulary
-          Only then you should see the hierarchy in the Inferred Concept Hierarchy frame.
(JPE later adds "I have declared the narrower and broader properties as "transitive" using protégé and it  works.")

Posting them here, to make it easier for me to find (and maybe to help others).

Friday, May 7, 2010

Recently, I was asked for pointers to introductory material on the Semantic Web, specifically for such topics as N3 and Turtle.  I found "The Semantic Web 1-2-3" the most helpful in starting to get to grips with the mysteries of the SemWeb.  Note, however, that it is somewhat outdated now (it refers to DAML+OIL for example).  But it is still a good foundation and it has more links to great semwebby material than you can shake a stick at, if that's your idea of fun.  A couple of extra links that might be of help are
I can't really find anything good that explains Turtle, other than the formal spec.  But my twitter-length explanation is that Turtle is N3 minus the reasoning extensions but plus internationalization (i.e. it is a more exact rendition of RDF than N3 is).

Other great links out there?