SyntaxHighlighter

SyntaxHighlighter

Showing posts with label news. Show all posts
Showing posts with label news. Show all posts

Tuesday, April 16, 2019

Standardized Rights Statements for News

At the IPTC Spring Meeting in Lisbon  https://iptc.org/events/spring-meeting-2019/, I proposed IPTC Rights Statements For News, an approach inspired by https://rightsstatements.org/.

This approach would support both efficient filtering of content and sophisticated evaluation of restrictions. It is simple, flexible, accurate, descriptive and transparent.

You can read the full proposal at https://www.slideshare.net/smyles/iptc-rights-statements-for-news but, in a nutshell, the idea is:
1. Create a standard set of rights statements specific to news and media
2. Express each rights statement as a URL, which can be embedded in content and therefore is suitable for filtering
3. Enable each rights statement URL to be dereferencable, meaning it can be evaluated by machines or by people

I think that this approach would work well for expressing rights and restrictions for news and media. Having other industry players work with rights in a compatible way would help with adoption. If  customers are getting rights statements in the same way from several publishers, it is more likely that CMS and MAM vendors will implement the necessary support.

If you're interested in finding out more about this or other IPTC initiatives, feel free to get in touch https://iptc.org/about-iptc/contact-us/

Thursday, October 25, 2018

Rights, Classification and Search Relevance for News - A Short Wrap Up of the IPTC Toronto Meeting

Last week, the IPTC held its Autumn meeting, a chance for people from around the world who have a common interest in news and media technology to discuss standards and learn from each other.

We had an entire day dedicated to news search, classification and descriptive metadata. AP's own Chad Schorr discussed our use of Elastic for robust indexing of content, Veronika Zielinska discussed AP's rules-based automated news classification system, and I reviewed our new automated tagging for images. We also heard from our peers at Bloomberg, New York Times, DPA and NTB on their systems. Finally, we had the opportunity to hear directly from Elastic on their suggestions for the best way to use their tools for news and media content.

A fairly momentous event for IPTC this year was Google Image's agreement to display image credit metadata for photos. This follows many years of discussions between IPTC, CEPIC and many others with the search giant. These talks came to a head during the CEPIC Congress / IPTC Photo Metadata Conference in Berlin in May 2018 and I'm very glad to see that concrete changes to Google Images use of metadata followed swiftly after. During the discussion of IPTC's Rights work for the news and media industry, we discussed ways to build on this progress - centred on driving adoption of the RightsML standard for machine processable expressions of rights and restrictions. We also discussed ways that IPTC could cooperate with other organizations, such as Europeana, to drive adoption of rights metadata.

As with other IPTC face-to-face meetings, there were many other interesting presentations and discussions, including the latest developments in video metadata, sports data and hearing from Civil on their plans for blockchain-backed journalism. For a more complete overview, checkout IPTC's posts on Day 1 and Day 2. And consider joining IPTC's next face-to-face meetings in Lisbon and Paris.

Wednesday, May 9, 2018

News Credibility, Verification and the Madness of Crowds - A Junk News Roundup

As the Associated Press states in our News Values and Principles:

"We have a long-standing role setting the industry standard for ethics in journalism. It is our job — more than ever before — to report the news accurately and honestly."

It is easy to see how AP is taking concrete steps in this area by, for example, our fact checking work (online, on twitter). And the AP Verify project is building a "newsroom tool that will combine artificial intelligence with our editorial expertise to automatically source and verify user-generated content."

I thought it would be interesting to take a look at some efforts going on elsewhere in the areas of credibility, verification and identifying junk news.

Standards Efforts

The IEEE is working on P7011 "Standard for the Process of Identifying and Rating the Trustworthiness of News Sources". The IEEE is a formal standards body, responsible for many of the technical standards which underpin the internet.

The Credibility Coalition describes itself as "an interdisciplinary community committed to improving our information ecosystems and media literacy through transparent and collaborative exploration." It is not, in itself, a standards body. If you examine the CredCo "about" page, you will spot my photo - I attended early meetings.

The Credible Web W3C Community Group describes its mission as "to help shift the Web toward more trustworthy content without increasing censorship or social division." There is a significant overlap between members of the Credibility Coalition and the Credible Web Community Group.  Despite the W3C link, this is not a formal standards effort - Community Groups are open to anyone. There are weekly video conferences to define an informal standard.

The Trust Project describes itself as "a consortium of top news companies" and says it "is developing transparency standards that help you easily assess the quality and credibility of journalism." Again, the Trust Project is not a formal standards body (like IEEE, IPTC or W3C).

Verification Projects

At the recent IPTC meeting, we saw presentations about two European projects aimed at helped to identify the spread of misinformation.

Truly Media is a joint project between ATC and Deutsche Welle. It is a "a web-based collaboration platform developed to support primarily journalists and human rights workers in the verification of digital content," and was developed with funds from EU and the DNI.

InVid aims to develop "a knowledge verification platform to detect emerging stories and assess the reliability of newsworthy video files and content spread via social media." It is an EU-funded project. Their demo was quite sophisticated. They also have a browser plugin which lets you verify news video and images yourself.

Wisdom and Madness

Finally, via Fair Warning, I saw "The Wisdom and Madness of Crowds" - a fun explainer in the form of a game. It walks you through why some crowds turn to madness and some to wisdom, with a focus on the spread of misinformation but also good information. It helps give some insight into the different dynamics at play and even some suggestions for how to reduce the spread of junk news and amplify the spread of verified news.

Monday, April 30, 2018

IPTC Names Brendan Quinn as New Managing Director, Celebrates the Service of Michael Steidl

As well as being Director of Information Management at the AP, I'm also Chairman of the Board of the IPTC, the standards body for the international news and media industry. The IPTC sets technology standards used around the world, including photo metadatavideo metadataand machine readable rights. We also develop technical approaches to the challenges facing the news media, such as identifying junk news, leveraging automation and coping with the impact of GDPR. Together with the other members of the IPTC - including ReutersAFPDPA and the New York Times - I help organize face-to-face meetings and numerous teleconferences so that we can work together and learn about interesting new projects from vendors and academics.
The Managing Director is the sole employee of the IPTC, helping to organize the work, manage the finances and recruit new membership. For the last 15 years, Michael Steidl has held this role. When Michael announced his plan to retire in the summer of 2018, I organized and ran the effort to find and recruit a successor. We talked to many candidates, several of whom were highly qualified. The Board did not take the decision lightly. In the end, we made an offer to Brendan Quinn - and we are thrilled he has accepted. Brendan brings with him a wealth of news technology experience, with organizations from around the world and of all sizes. He even worked at the Associated Press on AP's Video Hub. He has a unique combination of strategic insight into the challenges faced by the news industry and the technical know-how to help guide our work in technical standards and beyond. I look forward to partnering with Brendan in charting the future of the organization and to grow the work and influence of the IPTC.

Celebrating Michael Steidl's 15 Years as IPTC Managing Director

At the IPTC's Spring 2018 Meeting in Athens, we covered many interesting technical news topics - including video metadata, machine-readable rights, news credibility, and the challenges of localization and localisation. We also welcomed our incoming Managing Director, Brendan Quinn, and took some time to celebrate our retiring Managing Director, Michael Steidl.

As Chairman of the IPTC, I was honoured to give a speech, marking Michael's achievements over the last 15 years of service to the organization.

I would like to take a few minutes to celebrate Michael Steidl, our IPTC Managing Director for the past 15 years. Looking back over that time, many things stand out. I think we would all agree that Michael has a comprehensive understanding of all aspects of the IPTC - the standards, the history and - last but not least - the often obscure facets of the rules and regulations. Many times, an enthusiastic IPTC person has suggested some change or new idea, only to be gently reminded, "Well, I remember a discussion in 2008 where we voted on that topic and we decided..." I know that we will all miss Michael's kind but firm insistence that the rules and history of the IPTC be respected. And, in many ways, he acts as the representative of all the member organizations, whether or not they happen to be present in a particular discussion, to ensure that, for example, the large organizations don't dominate at the expense of the smaller organizations.

To give some perspective on Michael's achievements, I thought it would be interesting to look back on what Michael himself said about his role and the work of the IPTC. In December of 2003, the IPTC Spectrum - the old newsletter we used to publish - included Michael's reflections on his first year as Managing Director. I recommend you read the whole thing yourself. But now I'd like to focus on three things he said.

First, Michael described visiting the Van Gogh museum in Amsterdam and drew an analogy between the artist's evolving use of colour and his own intention in how he would work within the IPTC. He said "This might be a metaphor for the big task I jumped into: not to reinvent the wheel of running the IPTC office, well developed, maintained, and handed over by David Allen, but to add some extra shades of colour to IPTC’s image as a major player for standards in the news industry." Of course, looking back over the past 15 years of work, it is clear that on the one hand Michael did succeed in taking over the reins from his predecessor. But, on the other hand, he has done a lot more than simply adding some extra shades of colour. In fact, I would say that Michael's contributions to the IPTC is really more equivalent to an entirely new artistic movement - a sort of Renaissance for the organization - including managing the introduction of entirely new ways of operating the IPTC. When Michael started, there were no teleconferences or video conferences or even development of standards through email lists. There was no internet available during the meetings - which has perhaps been a mixed blessing, since people can keep up with the work back home, but we aren't always as focused.

There were already some hints of these changes in Michael's remarks. The second of the three quotes I want to pick out from the 2003 Spectrum:

We want to "discuss new ways of developing and maintaining our standards. These appear quite necessary to me: in the past decades IPTC usually developed and maintained one to two standards in parallel; IPTC 7901 was succeeded by IIM; and this was followed by NITF over a period of almost 15 years. But now three standards - NITF, NewsML, SportsML - have been developed and approved in a time span of about eight years. These three standards are all currently active and an additional three are under development ProgramGuideML, EventsML and an upcoming weather mark up. So soon we will have six active standards." So, three standards were developed in 15 years. Then three more in eight years. So, the IPTC work was already accelerating. But, just as a reminder, at this meeting in Athens, we discussed six different standards and will vote on three major updates. On top of that, we've discussed ten or more additional work areas - including the VideoDextra initiative, the EXTRA project and everyone's favourite topic of GDPR. Some people might think that standards take a long time to develop - and they do! - but we're no longer producing three standards in 15 years. or even only discussing technical standards anymore.

Now, 15 years later, Michael knows every detail of a full range of standards and an impressive array of initiatives. As was mentioned earlier, Michael has developed an extensive records keeping scheme of - I believe it was - 5,000 file folders stuffed with IPTC information. Now, with Brendan Quinn coming on board as our new Managing Director this summer, it must seem like a daunting task to succeed such an accomplished Managing Director.

So, I want to come to the third Spectrum quote from Michael from back in 2003, to reassure Brendan that Michael was once in the same boat. Michael said: "Yes, I had to learn the ropes first. IPTC operations are complex and it’s like conquering an unknown island: region after region had to be explored and all details of operation had to be made transparent, for me and to others. Preparing and providing the required resources for a meeting, taking minutes that reproduce the key points of the discussions, handling the finances, and last but not least supporting and co-ordinating the technical work of IPTC was occasionally really breathtaking and I have to admit it was a steep learning curve." So, Brendan, don't worry it wasn't easy for Michael either, but it can be done!

Finally, I want to close with an entirely different aspect of Michael's time with the IPTC. I've talked a lot about Michael's work. And, of course, solving news technology problems is the main reason for IPTC's existence. However, Michael has always pointed out in his polite, gentle but firm way that there is more to it than that. The IPTC is also an organization made up of people. It is a unique mix of people who often come from rival organizations and quite different backgrounds, who are able to come together and learn from each other and co-operate to solve problems together. And, in that process, it is often the case that rivals can become colleagues and colleagues can become friends. Michael, many of the people here - and many others around the world - count you as a friend. And so, along with your many work achievements with the IPTC, you should be very proud of all the colleagues and friends you have made.

And now, I'd like to ask all of your colleagues and friends here to join me in raising a glass, thanking you and wishing you a very happy retirement. THANK YOU MICHAEL!

Monday, November 20, 2017

The View from Barcelona - IPTC AGM 2017

I Chair the Board of Directors of IPTC, a consortium of news agencies, publishers and system vendors, which develops and maintains technical standards for news, including NewsML-G2, rNews and News-in-JSON. I work with the Board to broaden adoption of IPTC standards, to maximize information sharing between members and to organize successful face-to-face meetings.

We hold face-to-face meetings in several locations throughout the year, although, most of the detailed work of the IPTC is now conducted via teleconferences and email discussions. Our Annual General Meeting for 2017 was held in Barcelona in November. As well as being the time for formal votes and elections, the AGM is a chance for the IPTC to look back over the last year and to look ahead about what is in store. What follows are a slightly edited version of my remarks at the Barcelona AGM.
IPTC has had a good year - the 52nd year for the organization!
We've updated our veteran standards, Photo metadata - our most widely-used standard - and NewsML-G2 - our most comprehensive XML standard, marking its 10th year of development.
We're continuing to work in partnership with other organizations, to maximize the reach and benefits of our work for the news and media industry. In coordination with CEPIC we organized the 10th annual Photo Metadata Conference, looking to the future of auto tagging and search, examining advanced AI techniques - and considering both their benefits and their drawbacks for publishers. With the W3C we have crafted the ODRL rights standard and are launching plans to create RightsML as the official profile of the ODRL standard, endorsed by both the IPTC and W3C.
We've also tackled problems that matter to the media industry with technology solutions which are founded on standards, but go beyond them. The Video Metadata Hub is a comprehensive solution for video metadata management that allows exchange of metadata over multiple existing standards. The EXTRA engine is a Google DNI sponsored project to create an open source rules based classification engine for news.
We've had some changes in the make-up of IPTC. Johan Lindgren of TT joined the Board. Bill Kasdorf has taken over as the PR Chair. And we were thrilled to add Adobe as a voting member of IPTC, after many years of working together on photo metadata standards. Of course, with more mixed emotions, we have also learnt that Michael Steidl, the IPTC Managing Director, for 15 years will retire next Summer. As has been clear throughout this meeting and, indeed, every day between the meetings on numerous emails and phone calls, Michael is the backbone of the work of the IPTC. Once again, I ask you to join me in acknowledging the amazing contributions and dedications that Michael displays towards the IPTC.
Later today, we will discuss in detail our plans to recruit a successor for the crucial role of the Managing Director. And this is not the only challenge that the IPTC faces. We describe ourselves as "the global standards body of the news media" and that "we provide the technical foundation for the news ecosystem". As such, just as the wider news industry is facing a challenging business and technical environment, so is the IPTC.
During this meeting, we've talked about some of the technical challenges - including the continuing evolution of file formats and supporting technologies, whilst many of us are still working to adopt the technologies from 5 or 10 year ago. We've also talked about the erosion of trust in media organizations and whether a combination of editorial and technical solutions can help.
But I thought I would focus on a particular shift in the business and technical environment for news that may well have a bigger impact than all of those. That shift can be traced back to 2014 which, by coincidence, is when I became Chairman of the IPTC. Last week, Andre Staltz published an interesting and detailed article called "The Web Began Dying in 2014, Here's How". If you haven't read it, I recommend it. The article makes a number of interesting points and backs them up with numerous charts and statistics. I will not attempt to summarize the whole thing, but a few key points are worth highlighting.
Staltz points out that, prior to 2014, Google and Facebook accounted for less than 50% of all of the traffic to news publisher websites. Now those two companies alone account for over 75% of referral traffic. Also, through various acquisitions, Google and Facebook properties now share the top ten websites with news publishers - in the USA 6 of the 10 most popular websites are media properties. In Brazil it is also 6 out of 10. In the UK it is 5 out of 10. The rest all belong to Facebook and Google.
Both Facebook and Google reorganized themselves in 2014, to better focus on their core strengths. In 2014, Facebook bought Whastapp and terminated its search relationship with Bing, effectively relinquishing search to Google and doubling down on social. Also in 2014, Google bought DeepMind and shutdown Orkut, its most successful social product. This, along with the reorganization into Alphabet, meant that Google relinquished social to Facebook and allowing it to focus on search and - even more - artificial intelligence. Thus, each company seems happy to dominate their own massive parts of the web.
But ... does that matter to media companies? Well, Facebook said if you want optimal performance on our website, you must adopt Instant Articles. Meanwhile, Google requires publishers to use its Accelerated Mobile Pages or "AMP" format for better performance on mobile devices. And, worldwide, Internet traffic is shifting from the desktop to mobile devices.
Then, if you add in Amazon, Apple and Microsoft, it is clear that another huge shift is going on. All of the Frightful Five are turning away from the Web as a source of growth and instead turning to building brand loyalty via high end devices. Following the successful strategy of Apple, they are all becoming hardware manufacturers with walled gardens. Already we have Siri, Cortana, Alexa and Google Home. But also think about the investments going on by these companies in AR and VR as ways to dominate social interactions, e-commerce and machine learning over the Internet.
So, just as news companies must confront these shifts in the global business and technology environment, so must the IPTC. During this meeting, we've talked about our initial efforts to grapple with metadata for AR, VR and 360 degree imagery. We've also discussed techniques which are relevant to news taxonomy and classification, including machine learning and artificial intelligence. At the same time, Facebook, Google and others are not totally in control, as they - along with Twitter - found themselves having to explain the spread of disinformation on their platforms and under increased government scrutiny, particular in the EU. So, all of us, whether we describe ourselves as news publishers or not, are dealing with a rapidly changing and turbulent information, technical and business environment.
What does this mean for IPTC? IPTC is a news technology standards organization. But it is also unique in that we are composed of news companies from around the world. We know from the membership survey that both of these factors - influence over technical solutions and access to technology peers from competitors, partners, diverse organizations large and small - are very important to current members. In order to prosper as an organization, IPTC needs to preserve these unique benefits to members, but also scale them up. This means that we need to find ways to open up the organization in ways that preserve the value of the IPTC and fit with the mission, but also in ways that are more flexible. We need to continue to move beyond saying that the only thing we work on is standards and instead use standards as a component of the technical solutions we develop, as we are doing with EXTRA and the Video Metadata Hub. We need to work with diverse groups focused on solving specific business and journalistic problems - such as trust in the media - and in helping news companies learn the best ways to work with emerging technologies, whether it is voice assistants, artificial intelligence or virtual reality.
I'm confident that - working together - we can continue to reshape the IPTC to better meet the needs of the membership and to move us further forward in support of solving the business and editorial needs of the news and media industry. I look forward to working with all of you on addressing the challenges in 2018 and beyond.
Thank you.

Tuesday, August 29, 2017

Emoji, Fake News and 99% Invisible

This morning, I was listening to 99% Invisible, the podcast all about architecture and design.
thinking face
This episode "Person in Lotus Position" was about the process of adding a new emoji to the official set. At one point, they spoke to Jennifer 8. Lee who is on the Unicode Emoji Subcommittee. I know Jenny through Misinfocon. This is a new effort to fight the spread of disinformation on the web via a Knight-funded Credibility Schema Working Group. The goal is to create ways which signal whether a given piece of information on the web is credible.
newspaper
Most of the podcast episode describes the workings of the Unicode committee, which is official standards body for deciding which characters computers and phones will recognize and exchange. It gave a pretty good introduction to the importance and difficulty of this kind of standards work. (As well as being involved in the Credibility Schema Working Group, I'm also the Chairman of the IPTC, the news technology standards body. So, I like to think I have some insight into how these things work).
technologist
If you, like me, are interested in emoji and/or the workings of technical standards groups, then I recommend the episode. (Also, if you're interested in stopping the spread of fake news or in promoting technical standards within the global news industry, feel free to get in touch).

Tuesday, April 4, 2017

EXTRA Progress - Building an Open Source News Classification Engine

Over the last year, I've been leading a project within the IPTC to build an open source rules-based classification engine for news. Dubbed "EXTRA" (shorthand for EXTraction Rules Apparatus), the software will be freely-available under an MIT license. The work is being funded by a grant of €50,000 from Google's Digital News Initiative Innovation Fund.
“Extra” by Jeremy Brooks https://flic.kr/p/4aKH3c
We have drawn up the technical requirements, hired Infalia PC to partner with us on building the software and selected Elasticsearch's percolator as the fundamental technology. We've licensed two news corpora - English from Reuters and German from the Austrian Press Agency. Linguists are creating rules for classifying those corpora with IPTC’s Media Topics using the EXTRA engine.

I'm thrilled to say that the project is on track to deliver a working version of the engine, together with the sample rules, by the summer of 2017. Read more about the project at https://iptc.org/news/extra-iptc-infalia-elasticsearch-open-source-rules-based-classification-engine/ and feel free to contact me for more information.

Wednesday, January 11, 2017

Developing the Digital Marketplace for Copyrighted Works

I recently spoke at "Developing the Digital Marketplace for Copyrighted Works", organized by the Commerce Department’s Internet Policy Taskforce. The goal of the public meeting was to "facilitate constructive, cross-industry dialogue among stakeholders about ways to promote a more robust and collaborative digital marketplace for copyrighted works". My impression of the roughly 80 attendees was of a mix of publishers and lawyers - who tended to be pretty conservatively dressed - and music people - who tended to be in more sparkly outfits. Most of the discussion revolved around music, photo and video, but it turned out that a lot of the problems and potential solutions were quite similar across industries and media types. I've linked to the video of the event at the end of this post.

I've been involved in rights work at the Associated Press, including adding rights and pricing metadata to AP's Image API. I lead the IPTC's Rights Working Group. And I'm working within the W3C's POE group to turn ODRL into an official standard.


I spoke on the first panel with the topic of "Unique Identifiers and Metadata". I was teased a bit about "fake news" (this was in Alexandria. VA on December 9th 2016, so close both in time and place to the U.S. Presidential Election). Amongst other things, I spoke about how apparently simple things - "let's agree on identifiers for photos" - turned out to be quite complicated - since a text item, a photo or a video is not really a single, simple atomic thing, but more like a molecule of information. (You can watch the entire panel - which turned out to be quite lively, despite the early hour - in the video linked below).

I also moderated a round table, with the topic "What are the practical steps to adopting standards for identifying and controlling copyrighted works?". As everyone at my table introduced themselves, they mostly said "oh, I'm just here to learn, I don't have much to contribute" but, in fact, we had a very vigorous discussion, which covered *lots* of topics! I summarized them during the "Plenary" session (again, I've linked to the video below). We talked about three areas. First, was why we need standards - creators and rights holders should be compensated for their work, which could be financial compensation or it could be getting distribution and recognition. Second, we talked about the big barriers - technology, the culture of the Internet and human nature itself. Finally, we talked about concrete steps which the government and other organizations could take to get standards developed and adopted. (For the details, you'll need to watch the video. My segment runs from about 39:30 to about 45:55 but I recommend watching everyone's summary of their individual breakout sessions)

If you're interested in rights, then you should consider coming to London for the week of May 15th. That's because the BBC is hosting a Rights Day on May 15th, the IPTC will be holding its Spring Meeting (including discussing RightsML) on May 16th and 17th and W3C will hold its face-to-face meeting on May 18th and 19th. If you're interested in any or all, contact me and I will put you in touch with the right rights people.

Opening Remarks and Panel Session 1: Unique Identifiers and Metadata
Panel Session 2: Registries and Rights Expression Languages


Panel Session 3: Digital Marketplaces

Plenary Session



Tuesday, November 22, 2016

The View From Berlin - IPTC AGM 2016

I Chair the Board of Directors of IPTC, a consortium of news agencies, publishers and system vendors, which develops and maintains technical standards for news, including NewsML-G2, rNews and News-in-JSON. I work with the Board to broaden adoption of IPTC standards, to maximize information sharing between members and to organize successful face-to-face meetings.

We hold face-to-face meetings in several locations throughout the year, although, most of the detailed work of the IPTC is now conducted via teleconferences and email discussions. Our Annual General Meeting for 2016 was held in Berlin in October. As well as being the time for formal votes and elections, the AGM is a chance for the IPTC to look back over the last year and to look ahead about what is in store. What follows are my prepared remarks at the Berlin AGM.

The Only Constant

It is clear that the news industry is experiencing a great degree of change. The business side of news continues to be under pressure. And, in no small part, this is because the technology involved in the creation and distribution of news continues to rapidly evolve.

However, in many ways, this is a golden age of journalism. The demand for news and information has never been higher. The immediate and widespread distribution of news has never been easier.

The IPTC has been around for 51 years. I've been a delegate to the IPTC since 2000 and Chairman of the Board since June 2014. I'd like to give my perspective on the changes going on within the news industry and how IPTC has and will respond.

We're On a Mission

IPTC is rooted in - and foundational to - the news industry. Our open source standards for news technology enable the operations of hundreds of news and media organizations, large and small. IPTC standards are instrumental in the software used to create, edit, archive and distribute news and information around the world.

We are starting to evolve the scope of our work beyond standards - such as via the EXTRA project to build an open source rules-based classification engine. Much of what we do is relevant to not only news agencies and publishers, but also to photographers, videographers, academics and archivists. By bringing together these diverse groups, we can not only create powerful, efficient standards and technologies, but also learn from each other about what works and what does not.

Ch-ch-changes

We've introduced quite a bit of change within the IPTC since I've become Chairman and that has continued over the last year.

What's Going On?

We're working to improve our existing family of standards by
  • continuing to improve documentation - to make it easier to get going with a standard and simpler to grasp the nuances when you want to expand your implementation
  • making our standards more coherent and consistent - as many organizations need to use a combination
We're extending the reach of the IPTC, both by working with other organizations (including PRISM, IIIF, WAN-IFRA and W3C). But also by engaging in new types of work such as EXTRA and the Video Metadata Hub, which are not traditional standards but are open source projects for the benefit of the community we serve.

Since I've become Chair, we've renewed our efforts to communicate the great work that we do. You can see a big uptick in our engagement via Twitter and LinkedIn, as well as by refreshing the design of our the IPTC website. Plus we're doing a lot more work "out in the open" on Github.

We're continuing to streamline the operations of the IPTC. We've simplified our processes to better reflect the ways we actually operate these days. For example we have dramatically reduced the number of formal votes we take. But we still have sufficient process in place to ensure that the interests of all members are protected. For 2017, we have decided to have two-plus-one face-to-face meetings, rather than our usual three-plus-one. We will hold two full face-to-face meetings (one in London, the other in Barcelona), plus our one day Photo Metadata conference in association with the CEPIC Conference in Berlin. This will allow us to intensify our work on the meetings, with more ambitious and compelling topics and speakers.

Do Better

As I said, we've been changing our processes, particularly for the face-to-face meetings. But what else could we do to simplify our processes whilst at the same time ensuring that there is a balance between the interests of all members? Are there ways for the IPTC to deliver more value to the membership? How do we continue to balance our policy of consensus-driven decision-making with the need to be more flexible and nimble?

IPTC is a membership-driven organization. Membership fees represent the vast majority of the revenue for our organization. As the news industry as a whole continues to feel pressure - including downsizing, mergers and, unfortunately some members going out of business - the IPTC is experiencing downward pressure on its own revenue. So, we are working on ways to reach new members, whilst at the same time ensuring that existing members continue to derive value. We're also open to exploring new ways of generating revenue which fit with our mission - let us know your ideas!

What new areas should the IPTC focus on? Many journalists are experimenting with an array of technologies - Augmented Reality, Virtual Reality, 360 degree photos, drones and bots, to name but a few. And let's not forget about the "Cambrian Explosion" of technologies related to news and metadata on the Web, including AMP, AppleNews, Instant Articles, rNews, Schema.org and OpenGraph. How can IPTC help - negotiating standards? Developing best practices? Navigating the ethics of these technologies?

Happy

If you're happy with the IPTC, then please tell others.

If you're not happy, then please tell me!

I Want to Thank You

Without you, the members of IPTC, literally none of this is possible. So, I'd like to take a moment to thank everyone involved in the organization, particularly everyone involved in all of the detailed work of the IPTC. And I'd like to acknowledge and thank Andreas Gebhard, who is stepping down from the Board, and Johan Lindgren who has been voted on.

Finally, I'd like to extend a special thanks to Michael Steidl, Managing Director of the IPTC, who is personally involved in almost every aspect of what we do.

2017

No doubt, next year will bring us many new and, often, unexpected challenges. I look forward to tackling with all of you, the IPTC.

Thursday, October 6, 2016

Developers Needed For IPTC's EXTRA Rules-based Classification Engine

Over the last several months, I've been working within the IPTC - along with a number of other news organizations - on "EXTRA" (shorthand for EXTraction Rules Apparatus), an open-source source rules based classification engine for news content. I'm thrilled because this week we reached a significant milestone: we started the formal process of looking for developers to implement the EXTRA engine.
“Extra” by Jeremy Brooks https://flic.kr/p/4aKH3c
The IPTC was awarded a grant of 50,000 from Google's Digital News Initiative Innovation Fund to build and freely distribute the initial version of EXTRA. As part of the IPTC, we are working with several news providers to supply sets of news documents, and with linguists to write rules to classify the documents. We've been working on defining the technical requirements and now we’re looking for software developers to design, develop, document and test EXTRA.

Below is the formal announcement. If you know anyone who might be interested, let them know. And if you are interested, please let us know!

Developers Needed For IPTC's EXTRA Rules-based Classification Engine

IPTC https://iptc.org/ is looking for software developers to design, develop, document and test EXTRA https://iptc.github.io/extra/, an open source rules-based classification engine for news. First preference will be given to applications received by 21st October 2016, and review will continue until the positions are filled. Applyhere.

"Classification" means assigning one or more categories to the text of a news document. Rules based classifiers use a set of Boolean rules, rather than machine-learning or statistical techniques, to determine which categories to apply.

EXTRA is the EXTraction Rules Apparatus, a multilingual open-source platform for rules-based classification of news content. IPTC was awarded a grant of €50,000 from the first round of Google’s Digital News Initiative Innovation Fund https://www.digitalnewsinitiative.com/ to build and freely distribute the initial version of EXTRA. DNI granted IPTC €50,000 for the entire project.


We are working with news providers to supply sets of news documents and with linguists to write rules to classify the documents. IPTC is looking for qualified developers to create the rules engine to accurately and efficiently categorize the documents using the rules. mandatory and preferred requirements.

Please consult this page for more information and to let us know if you’re interested in being considered.



Tuesday, July 26, 2016

Making Progress on Rights - W3C Permissions Obligations and Expressions First Public Working Drafts


I've been working within W3C's Permissions & Obligations Expression (POE) Working Group as an Invited Expert. We have just issued our First Public Working Drafts:
"one" by Andre Chinn
https://flic.kr/p/5pGcyx
The W3C POE WG aims to create recommendations for permissions, obligations and licensing statements for digital content. The WG is using the W3C ODRL Community Group specifications as the starting point for its work. These are the same specifications which form the foundation of IPTC's RightsML work.
"poe" by 为民 王
https://flic.kr/p/gp2Bc
If you're interested in digital content, then I recommend looking at - and commenting on - the W3C POE drafts. The ODRL Information Model describes the foundational concepts, entities and relationships of ODRL. The ODRL Vocabulary & Expression describes how to encode the ODRL model in XML, JSON and RDF.
"Use in case of emergency" by Katia Sosnowiez
https://flic.kr/p/5MMhFz
The POE WG has also published the Use Case and Requirements Note. I have contributed one of the Use Cases: News Permissions and Restrictions. Again, the Working Group is looking for feedback on - and contributions to - the Use Cases, so that it can derive a detailed set of requirements for the POE work.

Friday, February 7, 2014

Ban Unknown Properties!

In XML, it is common to define a schema, in part to help with validation. This means that you can take an instance document, which is meant to conform to that XML schema, and test whether it really does, using a validation engine. Such testsing can be very useful to catch otherwise hard-to-spot errors - like misspelling an attribute name or getting the order of elements wrong.
Information Validations by dopey
http://www.flickr.com/photos/dopey/9591636030/
JSON didn't originally have a schema language. However, IETF are developing one. When the IPTC created the standard for News in JSON, we decided to define a schema for NINJS, so that you can check whether your JSON objects conform to the standard.
JSON Card -- Front by equanimity
http://www.flickr.com/photos/equanimity/3762360637/
Currently, in the NINJS schema, we have set "additionalProperties" to false. That's because we want to make it possible to validate a JSON document using that schema. If we didn't set additionalProperties=false, then any property name - whether it is a deliberate provider extension or an inadvertent typo - will pass validation.
Property Line by Whatknot
http://www.flickr.com/photos/whatknot/3401555810/
On the other hand, this means that a provider who wants to add their own properties in the NINJS schema has to create their own local copy. We've explained how to do that on the dev website (http://dev.iptc.org/ninjs-How-To-create-provider-specific-extensions). But it isn't totally obvious that is what you're meant to do. And it is certainly different from other schema that the IPTC has developed using XML (such an NewsML-G2), which instead have extensibility built in from the start. However, with IETF JSON Schema v4 (the latest) that's the only choice we have (http://json-schema.org/latest/json-schema-validation.html#anchor64).
Ban Symbol by uvw916a
http://www.flickr.com/photos/25023895@N02/5350382220/
In the v5 version of JSON Schema, they are introducing "ban unknown properties mode" https://github.com/json-schema/json-schema/wiki/ban-unknown-properties-mode-%28v5-proposal%29. Rather than being something that is built into the schema (like additionalProperties), "ban unknown properties" is a directive to the validation engine. The v5 JSON Schema is due to be finalized Real Soon Now. And at least one JSON Schema validator supports it already https://github.com/geraintluff/tv4.
NEAT by wannaoreo
http://www.flickr.com/photos/wannaoreo/270690022/
This seems to be a much neater solution to the validation problem. And, frankly, without it, the JSON schema isn't nearly as powerful. So, I'd like to see IPTC adopt this for the NINJS schema. And, if you're working on a JSON schema, I recommend you look at it, too.

Monday, July 1, 2013

JSON, Rights and Linked Data for News: the Latest IPTC Meeting

What is the best way to represent news using JSON? How can publishers convey rights metadata, to make automatic publishing more efficient? What role does linked data play in improving the production and consumption of news?

In June 2013, publishers from around the world gathered at the IPTC face-to-face meeting in Paris - graciously hosted by the AFP - to discuss these and other topics.
News in JSON
JSON is a lightweight format which continues to gain in popularity and tool support. The IPTC has therefore undertaken an effort to define the best way to represent key news properties using this technology. Although we could automatically translate from one of IPTC's existing XML standards into JSON, we believe it is better to create a spec for news properties in a way that results in more "natural" JSON. Find out what we're proposing via my latest News in JSON slides, which discuss the current NINJS draft.

Rights Metadata
There is a lot of interest amongst publishers of all sizes in expressing rights metadata, particularly for photos. Presently, most publishers need to have editors read the notes associated with each photo, to find out if there are any restrictions they should observe. The promise of machine-readable rights metadata is to make this a more efficient process: for example, to be able to automatically detect when an editor needs to decide whether to use a particular piece of content, rather than examining every photo, just in case.

I've been leading an effort within the IPTC to create RightsML specifically to support publishing industry requirements, based on the general-purpose W3C Community Group standard ODRL framework. In March 2013, we organized a one day conference with representatives from publishers, news agencies, photographers trade associations, law firms and standards bodies to examine the question: how can technology help assert and protect the rights of content creators? Find out more about that discussion, including video of the presentations. We are now focused on driving adoption of the RightsML standard, with better documentation and examples.
photographer by liz west
http://www.flickr.com/photos/calliope/1430290427/
Embedding Rights in Photos
ODRL and hence RightsML has a data model with a well-defined representation in XML. However, many producers of photos would rather have the rights expressed inside the binaries themselves, rather than - or in addition to - having an XML "sidecar". (That's in part because many photo workflows discard any other files and just work with the photos themselves). The challenge is what format to use? Our experiments with trying to embed either RDF or XML within XMP didn't work.So, now we're looking at expressing ODRL in JSON.

NewsML-G2
Currently, NewsML-G2 is the IPTC's flagship standard for news exchange. It continues to evolve, with a full production release each year (and "developer" releases in between). At this meeting, APA unveiled a new perl library to make it easier to produce G2. And we learnt about a major effort to compare the details of NewsML-G2 production by major providers, with the goal of harmonizing them, to make it easier for our customers.
Harmony Roof Sculpture - Opera Garnier - Palais Opera by ell brown
http://www.flickr.com/photos/ell-r-brown/3772502193/
Linked Data and the Semantic Web
We heard about interesting progress from the BBC on using rNews and the Storyline Ontology to enrich the presentation of news on the web. The AFP's medialab showed innovative news prototypes, also leverage rNews, including one for extracting quotes by politicians on various topics. And we at the AP discussed some of the details behind our Metadata Services offering.

Join Us
As well as updates on standards and practical sharing of industry information, the once-a-year AGM is when the IPTC updates its policies. At the Paris meeting, we decided to add an additional level of membership: now individuals can join the group, making it easier for people to participate in the standards used by the news industry and allowing them to attend these kinds of face-to-face meetings. Contact the IPTC Office for more information.

Wednesday, May 29, 2013

Big Data: the Big Picture

In May 2013, I was invited to be part of a panel at DAM NY 2013 talking about "Big Data". I shared some of what we're doing with Big Data at the Associated Press.



In particular, I've worked on an effort to create a digital archive - all of the text, photo, video, graphics and audio content that AP has ever published - which adds up to hundreds of millions of items. We combine that with our rich taxonomy of people, places, companies, organizations and subjects, available to you as the AP Metadata Services. I explained a bit about how we use that archive of content to provide insight and drive further enrichment, all in support of better products and services.

I feel that the term "Big Data" is a bit of a buzzword that is getting attached to a lot of different efforts and there's some healthy skepticism about the true value of some of what is touted under that particular heading. However, we have really found that there's a lot to be learned from bringing together significant data sets. And, also, that working with large data sets really does require some different techniques.

The World of Big Data in a Single Infographic

The other day, this Wikibon Infographic on Big Data caught my eye. It seems to define the term pretty broadly, but manages to convey some of the key technologies that underpin the field (not just Hadoop) and why there seems to be such an upsurge (a combination of low scaling costs, maturing tools and larger enterprises with a thirst for data driven strategies).

Worth a glance.

Monday, February 25, 2013

Mining for eBooks

In February 2013, the W3C, in partnership with IDPF and BISG, organized a workshop on eBooks, in conjunction with O'Reilly's TOC. I was invited to speak about AP's and IPTC's experience with implementing permissions and restrictions with machine readable rights. (ePub lets you include DRM statements; it seems that some publishers are using ODRL v1; IPTC have selected ODRL v2 for the foundation of RightsML). It was a great experience being on the panel and I got a lot of thoughtful and interesting questions.
eBook Readers Galore by libraryman
http://www.flickr.com/photos/libraryman/5052936803/
eBook Newbie
I'm a bit of a eBook neophyte. However, I learnt a lot from hearing the other publishers talking about their experience, hopes and frustrations with this digital publishing mechanism. And it struck me how similar the news industry is to the book industry. In his opening keynote, Bill McCoy talked about the three main ways that publisher deliver books these days: files, apps and websites. Of course, these are also the three main ways that news is delivered, today, too (not to mention dead trees in both cases for non digital publishing). As various other speakers presented at the workshop, they repeatedly used examples from newspapers and magazines (although, sometimes, as illustrations of what *not* to do). And, it has to be said, both book publishers and news publishers are in the same boat of trying to figure out their digital futures.

For more about both the eBook workshop and TOC, I recommend Ivan Herman's reflections.
Evolution of Readers by jblyberg
http://www.flickr.com/photos/jblyberg/4505413539/

Mining for eBooks
Given how easy it is to create and publish an eBook, it would seem that mining a news archive could yield some interesting books. Some news publishers are already conducting experiments with ebooks in this way. For example, the UK's Guardian have a series of Guardian Shorts. (Martin Belam wrote some quite interesting articles about how he worked with the Guardian archive to create ebooks on the Internet and the Olympics). Similarly, Vanity Fair have also started to play with ebooks.

Of course, ebooks aren't the only way to make use of a rich news archive. The New York Times recently launched their TimesMachine which lets you see browse back issues between 1851 and 1922 (“all the news which was fit to print”).

As software continues to eat the world, it will be interesting to see how formerly different kinds of publishers converge and diverge in their attempts to make their digital ways.

Thursday, January 24, 2013

Machine Readable Rights: A One Day Conference in Amsterdam


I'm helping to organize a free, one day conference on 12th March 2013 in Amsterdam, to discuss "Machine Readable Rights and the News Industry".
Sunset Over Amsterdam (Frontpage) by Werner Kunz
http://www.flickr.com/photos/werkunz/4565900446/
We're aiming to bring together the major players who are interested in the topic, across business, legal, editorial and technical groups. We have quite a few people who have signed up already, but it isn't too late to register if you're interested in attending - or speaking.
Coffee cup by Doug88888
http://www.flickr.com/photos/doug88888/2953428679/
The IPTC has organized similar one day conferences before. We've found that the presentations and panel discussions are always thought provoking. And the less formal introductions and discussions that happen over the coffee breaks, lunch time chats and after meeting drinks are at least as important.

If you're interested in machine readable rights, then don't hesitate to sign up!