SyntaxHighlighter

SyntaxHighlighter

Showing posts with label semweb. Show all posts
Showing posts with label semweb. Show all posts

Monday, July 1, 2013

JSON, Rights and Linked Data for News: the Latest IPTC Meeting

What is the best way to represent news using JSON? How can publishers convey rights metadata, to make automatic publishing more efficient? What role does linked data play in improving the production and consumption of news?

In June 2013, publishers from around the world gathered at the IPTC face-to-face meeting in Paris - graciously hosted by the AFP - to discuss these and other topics.
News in JSON
JSON is a lightweight format which continues to gain in popularity and tool support. The IPTC has therefore undertaken an effort to define the best way to represent key news properties using this technology. Although we could automatically translate from one of IPTC's existing XML standards into JSON, we believe it is better to create a spec for news properties in a way that results in more "natural" JSON. Find out what we're proposing via my latest News in JSON slides, which discuss the current NINJS draft.

Rights Metadata
There is a lot of interest amongst publishers of all sizes in expressing rights metadata, particularly for photos. Presently, most publishers need to have editors read the notes associated with each photo, to find out if there are any restrictions they should observe. The promise of machine-readable rights metadata is to make this a more efficient process: for example, to be able to automatically detect when an editor needs to decide whether to use a particular piece of content, rather than examining every photo, just in case.

I've been leading an effort within the IPTC to create RightsML specifically to support publishing industry requirements, based on the general-purpose W3C Community Group standard ODRL framework. In March 2013, we organized a one day conference with representatives from publishers, news agencies, photographers trade associations, law firms and standards bodies to examine the question: how can technology help assert and protect the rights of content creators? Find out more about that discussion, including video of the presentations. We are now focused on driving adoption of the RightsML standard, with better documentation and examples.
photographer by liz west
http://www.flickr.com/photos/calliope/1430290427/
Embedding Rights in Photos
ODRL and hence RightsML has a data model with a well-defined representation in XML. However, many producers of photos would rather have the rights expressed inside the binaries themselves, rather than - or in addition to - having an XML "sidecar". (That's in part because many photo workflows discard any other files and just work with the photos themselves). The challenge is what format to use? Our experiments with trying to embed either RDF or XML within XMP didn't work.So, now we're looking at expressing ODRL in JSON.

NewsML-G2
Currently, NewsML-G2 is the IPTC's flagship standard for news exchange. It continues to evolve, with a full production release each year (and "developer" releases in between). At this meeting, APA unveiled a new perl library to make it easier to produce G2. And we learnt about a major effort to compare the details of NewsML-G2 production by major providers, with the goal of harmonizing them, to make it easier for our customers.
Harmony Roof Sculpture - Opera Garnier - Palais Opera by ell brown
http://www.flickr.com/photos/ell-r-brown/3772502193/
Linked Data and the Semantic Web
We heard about interesting progress from the BBC on using rNews and the Storyline Ontology to enrich the presentation of news on the web. The AFP's medialab showed innovative news prototypes, also leverage rNews, including one for extracting quotes by politicians on various topics. And we at the AP discussed some of the details behind our Metadata Services offering.

Join Us
As well as updates on standards and practical sharing of industry information, the once-a-year AGM is when the IPTC updates its policies. At the Paris meeting, we decided to add an additional level of membership: now individuals can join the group, making it easier for people to participate in the standards used by the news industry and allowing them to attend these kinds of face-to-face meetings. Contact the IPTC Office for more information.

Wednesday, September 7, 2011

Six Ways You Can Help Us Get to rNews 1.0

The IPTC has been working on rNews since late 2010. We're now getting close to unveiling rNews 1.0 - and you can help get us there!

What is rNews?
rNews is a way to embed news-specific metadata into HTML pages. With a standard for representing that news-specific metadata, providers could consistently apply it to web pages in a way that makes it easier to use for tool makers (such as search engines). One key benefit of making it a standard is that it lets smaller players - both publishers and tool makers - participate on an equal footing with larger players. As I discussed in Seven Ideas for rNews you can think of rNews as being like a news-specific API for webpages to encourage open innovation (but with less technical and managerial overhead than typical APIs).

How has the IPTC been Crafting rNews?
Amongst other things, the IPTC's Semantic Web Working Group has been working steadily on the underlying model for rNews. We did lots of outreach to publishers and people interested in consuming news metadata on the web, including SemWeb meetups in New York, Berlin and London. Yesterday, the IPTC SemWeb group agreed to rNews 0.7, which you can see via this rNews spreadsheet.

What are the Specifics of rNews 0.7?
The spreadsheet includes the rNews classes, properties and definitions. It indicates the differences from the previous version of rNews. It also includes a potential alignment with schema.org, which has a similar model but is aimed at microdata specifically, whereas we see the rNews model as working with multiple different syntaxes, including both microdata and RDFa. The rNews site hasn't yet been updated to reflect all of the changes in version 0.7. Partly, that's because we only just agreed the details yesterday. But in part, that's also because we're about to move to rNews 1.0 - the first full Production Release of rNews!

How Can I Help Get to rNews 1.0?
The IPTC plans to vote on rNews 1.0 in the first week of October, at our next face-to-face meeting. You can help ensure that the rNews 1.0 launch is a success by doing any or all of the following:

1. Examine the rNews 0.7 spreadsheet and point out problems, omissions or contradictions via the rNews forum
2. A great way to get to grips with the details is to try marking up a news webpage using rNews 0.7 (you can see some examples linked from the rNews main page)
3. Try extracting rNews properties from an example page using an RDFa or microdata distiller
4. If you like your rNews example marked-up webpage, consider contributing it to the IPTC, so we can share it with others
5. Think about an innovative service or tool you can build using news markup applied consistently and at scale across the web
6. If you work for a publisher, start to talk to your colleagues about the potential to unlock innovation for news on the web via improved news metadata markup

Since the beginning, the rNews work has been a group effort, greeted with interest and enthusiasm from a variety of people. Now - with your help - we are within sight of reaching a very  significant milestone: the release of rNews 1.0.

Saturday, April 16, 2011

rNews - News Metadata in HTML

I chair the IPTC's SemWeb group. This is the third in a series of short posts about IPTC's work on using Semantic Web technologies for news. (The first post discussed where we stand with Linked Data for News and the second post looked at News Ontologies).

The big news at our March 2011 meeting was that the IPTC voted to approve Draft 0.1 of rNews. This kicks off an experimental phase, in which we ask people to learn about rNews and give us feedback via the rNews Forum. We plan to incorporate all the feedback (and fix a couple of little errors that we've found) so that Draft 0.2 of rNews will be ready for the Berlin IPTC meeting in June.
The Path by cuppini
http://www.flickr.com/photos/cuppini/539692848/
rNews and hNews
rNews is a set of specifications and best practices for using RDFa to embed news-specific metadata into HTML documents. It serves a similar purpose to hNews, although it uses a different technical approach. (hNews is a microformat - a set of conventions that make use of standard HTML elements and attributes to convey metadata). Although it is tempting to see hNews and rNews as rivals, I actually see them as supporting one another. And my suspicion is that tools that support one will pretty easily be able to support the other.

We're already starting to get some feedback on rNews, which is great. We've gotten some excellent questions about the rNews design philosophy from those with in-depth knowledge of the Semantic Web but little direct news publishing experience. And publishers who are old hands at the news business are looking at rNews as a way to learn more about semantic markup for their web content. Whatever your background, I encourage you to look at rNews and let us know what you think.

Thank You by theredproject
http://www.flickr.com/photos/theredproject/3302110152/
Getting There
The road to rNews started in the summer of 2010. Like any worthwhile technical standard, it required both articulating a vision of what we wanted to achieve and a lot of long conference calls, poring over the details in spreadsheets. I'd like to take this opportunity to particularly thank Dave ComptonJohn EvansAndreas GebhardJayson LorenzenEvan Sandhaus and Michael Steidl for their dedication in creating and refining rNews Draft 0.1

rNews Live!
I am excited about rNews and hNews and their potential to stimulate an ecosystem of tools for news on the web. I have some ideas for how this might happen, some of which I talked about in the recently-published interview about rNews with semanticweb.com.
NYC New York Times Building by wallyg
http://www.flickr.com/photos/wallyg/2259318046/
I plan to share some more ideas about rNews at the upcoming New York Semantic Meetup "Meet the IPTC and learn about rNews", hosted by the New York Times. You should come by and find out more.

Monday, April 4, 2011

News Ontology - Large Pieces, Loosely Joined

I chair the IPTC's SemWeb group. This is the second in a short series of short posts about IPTC's work on using Semantic Web technologies for news. (The first post discussed where we stand with Linked Data for News).

News Ontology - or, actually, ontologies
I am more-or-less aware of several projects that are underway to build ontologies relating to news. These various efforts seem more-or-less aware of each other and are, in some cases, working with each other directly. Specifically, I believe that PA, AFP, BBC, EBU, W3C and IPTC are each crafting ontologies - and there are probably more besides.

Wait - Onto What?
The word "ontology" often seems to scare people. But it just means a formal representation of knowledge in a particular domain, expressed using Semantic Web technologies (specifically OWL - the Web Ontology Language).

An OWL or a WOL? by dullhunk
http://www.flickr.com/photos/dullhunk/422076496/

You may also find this discussion of how ontologies relate to controlled vocabularies and taxonomies helpful. (Although, capriciously, I have linked to an ontology definition that is outside of the W3C SemWeb Technology orthodoxy. Or maybe that was deliberate?)

So, What Good is an Ontology?
Mike Atherton, User Experierence Designer at RedUXD, published a well-received presentation that nicely illustrates one powerful benefit of an ontology. However, his 100 slide Beyond the Polar Bear deck doesn't mention any ontologies until the 99th slide. Instead, it discusses the important of domain modeling and how that helps him build compelling user experiences for substantial websites, chock full of different types of media  and content types - he discusses various BBC websites and microsites. (I highly recommend you check it out - don't let the apparent length put you off).
Won't you be my friend? by ucumari
http://www.flickr.com/photos/ucumari/2440393713/
Ontologies in themselves don't directly deliver a benefit. Instead, they are infrastructure: they are a helpful way to structure information that - in combination with other technologies and practices - can deliver significant advantages over shallower ways of working in a particular domain, such as news. And not just for user experience; there are ways to exploit ontologies in other areas, including mining news (such as for sentiment analysis) or for data journalism.

Examples,  Please
Not all of the various ontologies are public at this time, but a few are.

For example, the BBC has ontologies for programmes, wildlife and  sport. The W3C recently published version 1.0 of their Ontology for Media Resources.

The EBU have experimented with an ontology based on IPTC's NewsML-G2, as have others. And at the recent IPTC face-to-face meeting, Paul Kelly of XML Team discussed the potential for Sports and Semantic Technologies.

Can't We Just Have One?
Rather than having all these different, overlapping ontologies, is it possible to just have one super, unified ontology?

In fact, the decisions about what to include, omit, emphasize or downplay in your ontology depend on what you're trying to do. I believe that each ontology therefore reflects a particular point of view, a specific editorial voice, in deciding what is and isn't important within a domain. That is not to say there would be no benefit from a coordinated, standardized news ontology. The work to create (never mind understand and use) an ontology is significant; even when you agree on the key things to model, there are different choices about the best way to express that model (sometimes driven by the limitations in the tools you have to work with). And having a standard model promotes greater interoperability amongst providers and more choice for clients.

One of the key benefits of Semantic Web technologies is the ability to mix and match different ontologies. And modern methods for developing ontologies (such as NeOn) emphasize reuse and composition. (See also the interesting Master Thesis "Analyzing and Ranking Multimedia Ontologies for their Reuse" by Ghislain Auguste Atemezin) So, I see the independent but somewhat coordinated efforts to create news-related ontologies as being a strength.
Mochuelo de hoyo by barloventomagico
http://www.flickr.com/photos/barloventomagico/2435316564/

Get Involved
If you would like to find out more about the work that the IPTC is doing to help standardize the use of Semantic Web technologies for news, then get in touch.

Tuesday, March 29, 2011

Linked Data for News - An Update on IPTC and Semantic Web Technologies

At the IPTC's most recent face-to-face meeting in Dubai, much of the discussions revolved around Semantic Web technologies. The big news was that the IPTC voted to approve Draft 0.1 of rNews, but this wasn't the only matter discussed.

As I blogged about before, the news standards body has been looking at three areas and how they relate to news:

I chair the IPTC's SemWeb group. Here is the first in a short series of posts on where each of these areas stand and how you can get involved.

The Semantic Web
Semantic Web technologies extend today's web with machine readable information and links between data and services. The "Semantic Web" has also been termed "The Giant Global Graph" and "Web 3.0", amongst other names; there are several different theories about exactly what it is and exactly how to get there.

IPTC News Codes using Linked Data
The IPTC has explored the technical aspects of representing news codes using the technologies and conventions of Linked Data.
Link by manel
http://www.flickr.com/photos/manel/315901872/
Linked Data is a set of best practices for publishing data using a subset of Semantic Web technologies. The IPTC News Codes are a set of metadata taxonomies designed for use by the news industry. The codes are already expressed in machine-readable XML, using IPTC's in G2 KnowledgeItem mechanism; it seemed a natural fit to explore expressing the news codes using the Linked Data principles.


IPTC and MINDS
Inspired by this IPTC work, MINDS (an association of European and US news agencies) and the IPTC have been mulling a joint project based on Linked Data for news. At the Dubai meeting, I reviewed the presentation I gave in February 2011 to MINDS about IPTC's Semantic Web and Linked Data work. There's a MINDS meeting in London the week of March 14th, so I expect we will learn more about any joint work.

The Chain by intherough
http://www.flickr.com/photos/intherough/3244476512/
If you'd like to access the IPTC news codes in SKOS (not to mention XHTML and G2) then visit http://www.iptc.org/site/NewsCodes/NewsCodes_Retrieval_in_Different_Formats and find out more about their full content-negotiation glory.

Get Involved
You can get involved with the news codes work by contacting the IPTC; one easy way to do that is to join the news codes Yahoo! email list.

Wednesday, February 2, 2011

Do OWL Classes Inherit Properties?

I'm slowly learning about using Semantic Web technologies. Sadly, I'm trying to do this in a rather ad hoc, as needed way, so my understanding is far from complete. I just ran across this interesting question: does an ontology subclass inherit the properties of its superclass?

Searching the web for opinions on this topic leads to a couple of contradictory views:

http://www.semanticoverflow.com/questions/619/rdfs-owl-inheritance-with-josekipellet
The answer to this question says: "Instances of subClasses do not inherit properties from instances of parent classes".

http://eclectic-tech.blogspot.com/2010/05/semantic-web-introduction-part-3-rdf.html
This states "In the example below, Penicillin is declared to be a sub class of both Antibiotic and USRegulatedMedication.It will therefore inherit the properties of those classes."

So, I turned to the W3C RDF Schema Recommendation, to see whether the definition would shed any light.

http://www.w3.org/TR/rdf-schema/#ch_subclassof
"The property rdfs:subClassOf is an instance of rdf:Property that is used to state that all the instances of one class are instances of another."

This reads to me that when C1 rdfs:subClassOf C2 then anything that is a C1 is also a C2.So, that seems to directly support the notion that C1 inherits all the properties of C2.

Certainly, if rdf:subClassOf doesn't mean that a subClass inherits the properties of the parent, then what does it mean? That it is unclear is a bit worrying, though.

Monday, November 15, 2010

IPTC and Semantic Web Technologies - Linked Data, Metadata and Ontology

At the IPTC's most recent face-to-face meeting in Rome, we reviewed our explorations of semantic web technologies for news. The news standards body has been looking at three major areas:


Linked Data
We discussed our work to turn IPTC's subject codes into Linked Data using SKOS concepts and Dublin Core properties. Michael Steidl (Managing Director of the IPTC) was planning to demo the IPTC Linked Data ... but, sadly, Internet access was not working in the hotel! He was, however, able to discuss the proposed collaboration between the IPTC and MINDS on Linked Data for news.

Linking and Mapping

Much of the discussion about IPTC's Linked Data work turned on the difficulties of mapping. In addition to representing the IPTC subject codes in RDF/XML and RDF/Turtle, there was some work done to map from the 17 top level IPTC terms to dbpedia concepts. We quickly figured out that these top level terms are chiefly umbrella terms and so don't map very well to individual dbpedia concepts. The meeting felt that it would be good to map the second level terms, but the problem is that this is quite a lot of work and - as usual - it isn't clear who will do it! We then explored some of the challenges of creating and maintaining the links in Linked Data - that is where a lot of the value, but also much of the investment, lies.

My slides about IPTC's Linked Data work are available on slideshare:

Metadata in HTML - rNews and hNews
Many news providers have created feeds to supply news using IPTC formats such as NITF and NewsML-G2. However, there are an increasing number of consumers of news who only want to work with "pure" web technologies, i.e. HTML rather than XML. So, the IPTC has been looking at the two major paths to represent metadata in HTML - microformats and RDFa.

hNews
I discussed hNews - the microformat for news that was adopted by the community in late 2009 - which builds upon hAtom by adding a few news-specific fields (such as Source and Dateline). As well as explaining how to add microformats to your HTML templates, I provided some statistics that the Associated Press has gathered on adoption. (As of October 2010, we know of about 1,200 sites using hNews, predominantly in North America). See my Prezi on hNews at http://prezi.com/uo5ggdkll8sa/an-introduction-to-hnews/ for more.

rNews
Evan Sandhaus (Semantic Technologist at The New York Times) described rNews - a proposal for an RDFa vocabulary for news. As the names imply, rNews and hNews are similar in intent (news-specific metadata in HTML) but somewhat different in approach. Whereas hNews went through the microformats process, an RDFa vocabulary can be created by anyone. Evan has created an initial rNews draft based somewhat on the NewsML-G2, NITF and hNews models but it is clearly heavily influenced by the needs of the New York Times.

Members of the IPTC's Semantic Web Yahoo! Group can view Evan's rNews draft and are encouraged to discuss it in that email group. At the Rome face-to-face meeting there was quite a lot of interest, but also several issues raised about the details of the first draft. The meeting generally agreed to continue looking at both hNews and rNews, with a view to making a recommendation on both in 2011.

The benefit of getting rNews and hNews adopted by the IPTC is that greater industry support translates into less work for toolmakers: if many news providers support hNews and/or rNews - and do so in very similar ways - then it is easier to build parsers and tools to extract metadata from HTML.

News Ontology
Benoît Sergent of the European Broadcasting Union discussed the work that he and his colleague Jean-Pierre Evain have been doing to create a news ontology, based upon the NewsML-G2 news model. Benoît described how EBU would like to combine the video content that it produces with content from its member organizations and other third parties. If they can represent this information using a flexible, universal model (the news ontology) they could use off-the-shelf tools (such as a triple store) to query, manipulate and recombine that content.

In many ways, this is the most fundamental piece of the semantic web work that the IPTC is undertaking. It is also the least accessible, for many. Members of the IPTC SemWeb group can view a draft of the news ontology and can comment in that email list.

Wednesday, July 21, 2010

Linked Data and the World Cup

In January of 2010, I attended the News Linked Data Summit. This was a collection of several organizations involved in news production and distribution, looking to see if there was a way to collaborate on moving forward the Semantic Web (and particularly Linked Data) for news. It was a very interesting discussion, lead mainly by the BBC and The Guardian. At the end of the day, the group decided to move forward together by producing Linked Data for the UK Election. (In January, the election was clearly going to happen in the near future, but not yet announced - it wound up happening in May 2010).

I had a counter proposal for a Linked Data experiment - the World Cup. This made more sense for my employer and seemed to be of interest to at least some others. It was not to be, however. (I later spoke to some folks about whatever happened with the UK Election Linked Data experiment. As far as I can tell, the Guardian did produce some election data. But it isn't clear to me that this turned into anything bigger).

By Shine 2010 http://www.flickr.com/photos/shine2010/3292491879/


However, it turns out that the BBC did use the Semantic Web to power their World Cup “microsite” (actually their World Cup site has more pages than the non-SemWeb Sports section). In "The World Cup and a call to action Around Linked Data", BBC Architect John O'Donovan gives an overview of how they used fine-grained metadata to be able to produce their site with far less need for editorial curation of individual pages. In an accompanying piece, Jem Rayfield describes the technical details of the "dynamic semantic publishing".

Although their descriptions are couched in the terms of the Semantic Web, much of what they describe is more to do with the application of fine-grained metadata than with the particular data formats they use. They do describe how they make use of certain Semantic Web technologies - such as an RDF Triplesore / SPARQL system - to apply derived properties. And it is almost certainly the case that the use of the RDF model has given them more flexibility than can be the case when you use RDBMS systems - or even XML-based schema. However, fundamentally, what they are describing is the potential for what could be done via the application of metadata to content. It is interesting to see it in action.


The England Team didn't do so well at the World Cup - thanks Doug88888! http://www.flickr.com/photos/doug88888/4550561194/

It is also interesting how much of a revelation this is to people. (The first article was widely distributed via Twitter. And the comments are quite breathless in their admiration).

Hopefully, the work that the IPTC is doing on Linked Data and the Semantic Web will help other news organizations (including the AP!) unlock some of the great metadata work that is going on behind the scenes...

Friday, May 7, 2010

Recently, I was asked for pointers to introductory material on the Semantic Web, specifically for such topics as N3 and Turtle.  I found "The Semantic Web 1-2-3" the most helpful in starting to get to grips with the mysteries of the SemWeb.  Note, however, that it is somewhat outdated now (it refers to DAML+OIL for example).  But it is still a good foundation and it has more links to great semwebby material than you can shake a stick at, if that's your idea of fun.  A couple of extra links that might be of help are
I can't really find anything good that explains Turtle, other than the formal spec.  But my twitter-length explanation is that Turtle is N3 minus the reasoning extensions but plus internationalization (i.e. it is a more exact rendition of RDF than N3 is).

Other great links out there?

Monday, April 26, 2010

Experimenting with Bull Fighting and the Semantic Web

Paul Kelly and I took a look at what it would take to represent NewsML-G2 style Knowledge Items as SKOS.


 Here's a typical G2 style concept from the IPTC subject code vocabulary:

<concept>
<conceptId qcode="subj:01003000" created="2008-01-01T00:00:00+00:00" />
<type qcode="cpnat:abstract">
<name xml:lang="en-GB">Abstract concept</name>
</type>
<name xml:lang="en-GB">bullfighting</name>
<name xml:lang="de">Stierkampf</name>
<name xml:lang="fr">Tauromachie</name>
<name xml:lang="es">toros</name>
<name xml:lang="es">??</name>
<definition xml:lang="en-GB">Classical contest pitting man against the bull</definition>
<definition xml:lang="de">Klassischer Wettkampf Mann gegen Stier.</definition>
<definition xml:lang="fr">Tauromachie, combat entre un homme et un taureau</definition>
<definition xml:lang="es">Clasico enfrentamiento entre Toro y Hombre.</definition>
<definition xml:lang="es">????????????????</definition>
<broader qcode="subj:01000000" type="cpnat:abstract" />
</concept>

Here's what I think that would render as in SKOS - first in N3:

@prefix dc: <http://purl.org/dc/elements/1.1/>.
@prefix skos: <http://www.w3.org/2004/02/skos/core#>.
@prefix subj: <http://iptcsubj.example.com/>.
@prefix rdf: <http://www.w3.org/1999/02/22-rdf-syntax-ns#>.

subj:01003000 rdf:type skos:Concept;
dc:created "2008-01-01T00:00:00+00:00";
skos:prefLabel "bullfighting"@en-GB;
skos:prefLabel "Stierkampf"@de;
skos:prefLabel "Tauromachie"@fr;
skos:prefLabel "toros"@es;
skos:prefLabel "闘牛"@es;
skos:definition "Classical contest pitting man against the bull"@en-GB;
skos:definition "Klassischer Wettkampf Mann gegen Stier."@de;
skos:definition "Tauromachie, combat entre un homme et un taureau"@fr;
skos:definition "Clasico enfrentamiento entre Toro y Hombre."@es;
skos:definition "人間を雄牛と戦わせる伝統的な競技"@es;
skos:broader subj:01000000.

Some conclusions. G2 concepts are easy to map to SKOS! But there are some choices that need to be made (e.g. do we use skos:broader or owl:broader?) It feels like a G2 conceptset is almost identical to a SKOS Concept Scheme.

And there's a bug in the IPTC mapping of the subject codes (the Japanese definition and name is labeled as Spanish).


¡Olé! by karwik