SyntaxHighlighter

SyntaxHighlighter

Saturday, April 16, 2011

rNews - News Metadata in HTML

I chair the IPTC's SemWeb group. This is the third in a series of short posts about IPTC's work on using Semantic Web technologies for news. (The first post discussed where we stand with Linked Data for News and the second post looked at News Ontologies).

The big news at our March 2011 meeting was that the IPTC voted to approve Draft 0.1 of rNews. This kicks off an experimental phase, in which we ask people to learn about rNews and give us feedback via the rNews Forum. We plan to incorporate all the feedback (and fix a couple of little errors that we've found) so that Draft 0.2 of rNews will be ready for the Berlin IPTC meeting in June.
The Path by cuppini
http://www.flickr.com/photos/cuppini/539692848/
rNews and hNews
rNews is a set of specifications and best practices for using RDFa to embed news-specific metadata into HTML documents. It serves a similar purpose to hNews, although it uses a different technical approach. (hNews is a microformat - a set of conventions that make use of standard HTML elements and attributes to convey metadata). Although it is tempting to see hNews and rNews as rivals, I actually see them as supporting one another. And my suspicion is that tools that support one will pretty easily be able to support the other.

We're already starting to get some feedback on rNews, which is great. We've gotten some excellent questions about the rNews design philosophy from those with in-depth knowledge of the Semantic Web but little direct news publishing experience. And publishers who are old hands at the news business are looking at rNews as a way to learn more about semantic markup for their web content. Whatever your background, I encourage you to look at rNews and let us know what you think.

Thank You by theredproject
http://www.flickr.com/photos/theredproject/3302110152/
Getting There
The road to rNews started in the summer of 2010. Like any worthwhile technical standard, it required both articulating a vision of what we wanted to achieve and a lot of long conference calls, poring over the details in spreadsheets. I'd like to take this opportunity to particularly thank Dave Compton, John Evans, Andreas Gebhard, Jayson Lorenzen, Evan Sandhaus and Michael Steidl for their dedication in creating and refining rNews Draft 0.1

rNews Live!
I am excited about rNews and hNews and their potential to stimulate an ecosystem of tools for news on the web. I have some ideas for how this might happen, some of which I talked about in the recently-published interview about rNews with semanticweb.com.
NYC New York Times Building by wallyg
http://www.flickr.com/photos/wallyg/2259318046/
I plan to share some more ideas about rNews at the upcoming New York Semantic Meetup "Meet the IPTC and learn about rNews", hosted by the New York Times. You should come by and find out more.

Monday, April 4, 2011

News Ontology - Large Pieces, Loosely Joined

I chair the IPTC's SemWeb group. This is the second in a short series of short posts about IPTC's work on using Semantic Web technologies for news. (The first post discussed where we stand with Linked Data for News).

News Ontology - or, actually, ontologies
I am more-or-less aware of several projects that are underway to build ontologies relating to news. These various efforts seem more-or-less aware of each other and are, in some cases, working with each other directly. Specifically, I believe that PA, AFP, BBC, EBU, W3C and IPTC are each crafting ontologies - and there are probably more besides.

Wait - Onto What?
The word "ontology" often seems to scare people. But it just means a formal representation of knowledge in a particular domain, expressed using Semantic Web technologies (specifically OWL - the Web Ontology Language).

An OWL or a WOL? by dullhunk
http://www.flickr.com/photos/dullhunk/422076496/

You may also find this discussion of how ontologies relate to controlled vocabularies and taxonomies helpful. (Although, capriciously, I have linked to an ontology definition that is outside of the W3C SemWeb Technology orthodoxy. Or maybe that was deliberate?)

So, What Good is an Ontology?
Mike Atherton, User Experierence Designer at RedUXD, published a well-received presentation that nicely illustrates one powerful benefit of an ontology. However, his 100 slide Beyond the Polar Bear deck doesn't mention any ontologies until the 99th slide. Instead, it discusses the important of domain modeling and how that helps him build compelling user experiences for substantial websites, chock full of different types of media  and content types - he discusses various BBC websites and microsites. (I highly recommend you check it out - don't let the apparent length put you off).
Won't you be my friend? by ucumari
http://www.flickr.com/photos/ucumari/2440393713/
Ontologies in themselves don't directly deliver a benefit. Instead, they are infrastructure: they are a helpful way to structure information that - in combination with other technologies and practices - can deliver significant advantages over shallower ways of working in a particular domain, such as news. And not just for user experience; there are ways to exploit ontologies in other areas, including mining news (such as for sentiment analysis) or for data journalism.

Examples,  Please
Not all of the various ontologies are public at this time, but a few are.

For example, the BBC has ontologies for programmes, wildlife and  sport. The W3C recently published version 1.0 of their Ontology for Media Resources.

The EBU have experimented with an ontology based on IPTC's NewsML-G2, as have others. And at the recent IPTC face-to-face meeting, Paul Kelly of XML Team discussed the potential for Sports and Semantic Technologies.

Can't We Just Have One?
Rather than having all these different, overlapping ontologies, is it possible to just have one super, unified ontology?

In fact, the decisions about what to include, omit, emphasize or downplay in your ontology depend on what you're trying to do. I believe that each ontology therefore reflects a particular point of view, a specific editorial voice, in deciding what is and isn't important within a domain. That is not to say there would be no benefit from a coordinated, standardized news ontology. The work to create (never mind understand and use) an ontology is significant; even when you agree on the key things to model, there are different choices about the best way to express that model (sometimes driven by the limitations in the tools you have to work with). And having a standard model promotes greater interoperability amongst providers and more choice for clients.

One of the key benefits of Semantic Web technologies is the ability to mix and match different ontologies. And modern methods for developing ontologies (such as NeOn) emphasize reuse and composition. (See also the interesting Master Thesis "Analyzing and Ranking Multimedia Ontologies for their Reuse" by Ghislain Auguste Atemezin) So, I see the independent but somewhat coordinated efforts to create news-related ontologies as being a strength.
Mochuelo de hoyo by barloventomagico
http://www.flickr.com/photos/barloventomagico/2435316564/

Get Involved
If you would like to find out more about the work that the IPTC is doing to help standardize the use of Semantic Web technologies for news, then get in touch.

Tuesday, March 29, 2011

Linked Data for News - An Update on IPTC and Semantic Web Technologies

At the IPTC's most recent face-to-face meeting in Dubai, much of the discussions revolved around Semantic Web technologies. The big news was that the IPTC voted to approve Draft 0.1 of rNews, but this wasn't the only matter discussed.

As I blogged about before, the news standards body has been looking at three areas and how they relate to news:

I chair the IPTC's SemWeb group. Here is the first in a short series of posts on where each of these areas stand and how you can get involved.

The Semantic Web
Semantic Web technologies extend today's web with machine readable information and links between data and services. The "Semantic Web" has also been termed "The Giant Global Graph" and "Web 3.0", amongst other names; there are several different theories about exactly what it is and exactly how to get there.

IPTC News Codes using Linked Data
The IPTC has explored the technical aspects of representing news codes using the technologies and conventions of Linked Data.
Link by manel
http://www.flickr.com/photos/manel/315901872/
Linked Data is a set of best practices for publishing data using a subset of Semantic Web technologies. The IPTC News Codes are a set of metadata taxonomies designed for use by the news industry. The codes are already expressed in machine-readable XML, using IPTC's in G2 KnowledgeItem mechanism; it seemed a natural fit to explore expressing the news codes using the Linked Data principles.


IPTC and MINDS
Inspired by this IPTC work, MINDS (an association of European and US news agencies) and the IPTC have been mulling a joint project based on Linked Data for news. At the Dubai meeting, I reviewed the presentation I gave in February 2011 to MINDS about IPTC's Semantic Web and Linked Data work. There's a MINDS meeting in London the week of March 14th, so I expect we will learn more about any joint work.

The Chain by intherough
http://www.flickr.com/photos/intherough/3244476512/
If you'd like to access the IPTC news codes in SKOS (not to mention XHTML and G2) then visit http://www.iptc.org/site/NewsCodes/NewsCodes_Retrieval_in_Different_Formats and find out more about their full content-negotiation glory.

Get Involved
You can get involved with the news codes work by contacting the IPTC; one easy way to do that is to join the news codes Yahoo! email list.

Thursday, March 24, 2011

BBC TV and Video Futures

The boffins at the BBC have been doing a lot of interesting work with an eye to the future of tv and video. And they have been kind enough to share a lot of their work via their blog.
Boffin by ajc1
http://www.flickr.com/photos/ajc1/2368836604/
A Connected Future for TV
Many TVs are able to connect to the Internet, either directly or via a plethora of IP-connected boxes. However, for most viewers, the Net is just a different way to receive programming - an alternative to cable, satellite or over-the-air broadcast channels. In March 2011, Roly Keating, BBC Director of Archive Content, used his keynote speech at the Digital Television Group's Annual Summit to describe his vision for Connected TV.
televsion by waltjabsco
http://www.flickr.com/photos/waltjabsco/684747788/
For Keating, Connected TV means more than just video on demand or using your TV to browse the Internet. He would like to enrich the experience of watching TV by giving viewers the ability to explore topics in greater depth, by providing just-in-time context, by creating links between TV and Internet programming and viewers. To me, this means bringing the power of hypermedia to visual content, just as the web has brought hypermedia to text in recent decades.

Frame Accurate Video in HTML5
Dirk-Willem van Gulik, BBC Chief Technical Architect, reviews how he and his team have worked with the open source community to create frame accurate video editing capabilities using web technologies, specifically HTML5. This means that using off-the-shelf browsers, it will be possible to work on video at professional levels of precision, on virtually any kind of web-connected device.

At this point, the capabilities are only in the bleeding-edge versions of (certain) browsers. And the frame accurate browser facilities are just foundational; the actual tools to perform full-fledged video editing can be built on top of this platform, but don't exist quite yet, it seems. But I think it is great that the BBC have really invested in building out missing capabilities in open source tools that will ultimately benefit not only themselves, but may other publishers, large and small.

BBC's Technology Vision
Spencer Piggott, Head of Technology Direction for BBC Technology, reveals the BBC Technology Strategy. In a comprehensive set of bullet points in the linked powerpoint, he covers the plans for such topics as high definition, content acquisition, development platforms, transcoding, internet distribution, rights management and search, amongst many others.
Vision sign by hamptonroadspartnership
http://www.flickr.com/photos/hamptonroadspartnership/5351621035/
It is great to get this kind of insight into the BBC Technology challenges and plans. It would seem that the BBC boffins are working on lots of interesting things, which is reassuring.
Reassuring by arenamontanus
http://www.flickr.com/photos/arenamontanus/2849737658/

Sunday, February 20, 2011

Recommendations and Randomness: How I Pick Which Books to Read

There was a time when I went to a library or a bookshop and was immediately faced with a problem: how to pick which book to read?  Over time, I have changed how I solve this problem, which has changed what I read.

Maybe I should change again?

How Did I Use to Do It?
One technique would be to find the section that corresponded to an interesting-to-me topic and randomly browse through those books, trying to find one that looked worth reading.  I would look at the table of contents, skim the first page or two.  This was generally pretty successful  - and explains why I would read so many non-fiction books.

The other thing I would do is to see if I could find the latest book by an author I had previously read and liked.  Which would explain why I read so many mystery or science fiction books (genres which thrive on series and name-brand authors) and virtually no literary fiction.  Also, I like mysteries and science fiction.

Obviously, I was not alone in these techniques.  Bookshops and libraries continue to be organized by topic and type (although bookshops in America seem to be evolving into coffee shops and ebook device vendors, with paper books as something of an afterthought).  And the book publishing industry loves when a big name author releases the latest in a series.

However, over time, I decided to change things up a little bit.

Why Did I Feel I Needed to Change My Ways?
The revolution in self-publishing means that a lot more good reading is easily accessible online.  (Together with a lot more bad reading, too, of course - let's not forget Sturgeon's Law).  And it is not only accessibility which has changed.  The good stuff is a lot more discoverable, too.  For me, the advent of blogs (and later Twitter) has meant that I get directed to the best non fiction reading all the time.  More, in fact, than I can read.  (It is possible that I don't actually find out about the best non fiction online via blogs and tweets, but ignorance is bliss).

At the same time, the quality of non fiction books appeared to me to decline.  Too often, it seemed that the essential idea is captured in the title.  The rest of the book is then simply a rehearsing of that idea over and over, typically lacking nuance and depressingly often an over-simplified look at the world through a single lens.

My fiction technique held obvious flaws - how to learn about new authors or read something other than genre fiction?  In the immortal words of the Stranglers - Something Better Change.

I've Got a Little List
Really, the new-to-me technique I adopted is quite old and very obvious.  But here it is.  On the one hand, I extended the techniques for discovering the good stuff online to the world of books.  Once in a while, the blog posts and tweets would mention an interesting-sounding book.  So, I would add the author and title to an online list. (I use backpack).  But the real source of recommendations is The New York Times' Sunday Book Review (I read the dead trees version, but you may prefer it online).  Yes, I find that critics are the best source of information about books.  And, often for non fiction, you get enough of the idea from the review, so that you don't need to read the actual book.  (I try to supplement my source of critical information, including the Guardian, but the NYT is my main source).

Randomness and Routine
So, I now have a set of great recommendations for books and authors in an online list.  (At the time of writing, I have somewhere between one hundred and two hundred entries). But how do I actually select what to read?

Well, my normal routine is: go to the library, consult my list of books to read and randomly trawl through the list, consulting the library's catalogue until I find some that are available. I will also try to balance things out a little bit by trying to read some non-fiction, along with non-fiction. Plus I like to get at least one smallish, paperback book that I can read on the train. But random factors like who else happens to have checked out a book on my list, plus whether the library happens to have decided to stock it in the first place, heavily influence what I read in practice (rather than the theory of my recommendations list).


How is it Working Out?
I've been pretty happy with this set of techniques. I read way more fiction - and different types of fiction - than I used to. I tend not to read all that much non-fiction anymore. Although, having noticed this, I've resolved to do better with that. It means that I almost exclusively read books from the library (although I will read a book if you give or lend me it, too).

Areas where it doesn't work include books that are too obscure or controversial for my local library to stock. (Typically, this limits the non-fiction work I might read). And it also means that the idea of e-books (the Nook, the Kindle, etc.) doesn't make any sense to me. Half of my technique relies on the constraints of what is available at the library. I imagine that e-books have some constraints (maybe not everything is available as an e-book?) but it doesn't seem the same, somehow. Similarly, it means that I don't really visit first-hand bookshops anymore. Unless there's a specific book that I need to buy (generally for someone else) or it is some unique emergency situation (such as in an airport).

How do you pick your books?

Wednesday, February 2, 2011

Do OWL Classes Inherit Properties?

I'm slowly learning about using Semantic Web technologies. Sadly, I'm trying to do this in a rather ad hoc, as needed way, so my understanding is far from complete. I just ran across this interesting question: does an ontology subclass inherit the properties of its superclass?

Searching the web for opinions on this topic leads to a couple of contradictory views:

http://www.semanticoverflow.com/questions/619/rdfs-owl-inheritance-with-josekipellet
The answer to this question says: "Instances of subClasses do not inherit properties from instances of parent classes".

http://eclectic-tech.blogspot.com/2010/05/semantic-web-introduction-part-3-rdf.html
This states "In the example below, Penicillin is declared to be a sub class of both Antibiotic and USRegulatedMedication.It will therefore inherit the properties of those classes."

So, I turned to the W3C RDF Schema Recommendation, to see whether the definition would shed any light.

http://www.w3.org/TR/rdf-schema/#ch_subclassof
"The property rdfs:subClassOf is an instance of rdf:Property that is used to state that all the instances of one class are instances of another."

This reads to me that when C1 rdfs:subClassOf C2 then anything that is a C1 is also a C2.So, that seems to directly support the notion that C1 inherits all the properties of C2.

Certainly, if rdf:subClassOf doesn't mean that a subClass inherits the properties of the parent, then what does it mean? That it is unclear is a bit worrying, though.