Currently Browsing: Open data

Open data and privacy. Should I bother?

Open data and privacy. Should I bother?

Privacy is often mentioned as an obstacle when implementing an open data policy, but never really elaborated on. Should you really bother about privacy when opening up your data? My answer: yes you should.

Alan Westin laid the foundation of our modern conception of information privacy, which focuses on the individual’s right to control what is known about him. The modern European right to information privacy still leans on the notion of privacy as a right to control one’s personal information. Article 8 of the Charter of Fundamental Rights of the European Union gives everyone the right “to the protection of personal data concerning him or her”. This fundamental right to information privacy is further elaborated by the EU Data Protection Directive. The concept of ‘processing personal data’ is the touchstone of this directive. Personal data should be processed fairly and for legitimate and specified purposes.

EU data protection is all about the protection of ‘personal data’. Personal data is “information relating to an identified or identifiable natural person” and an identifiable person is “one who can be identified, directly or indirectly, in particular by reference to an identification number or to one or more factors specific to his physical, physiological, mental, economic, cultural or social identity” (Article 2 of the EU Data Protection Directive). Personal data can thus be both directly and indirectly identifying.

Train times, the location of public toilets and the number of car accidents could all be open data. No open data provider will (hopefully) offer names, addresses, social security numbers, or other data that directly or indirectly identifies natural persons as open data. Open data is at the most anonymized or aggregated data that cannot be related to individuals. The Open Knowledge Foundation visualizes open data and “private data” as two non-overlapping subsets. Unfortunately, in reality this distinction is not so easy to draw.

Even when data has been anonymized or aggregated, data analysis techniques now allow us to re-identify individuals in such data (See Paul Ohm for an overview). For instance, when Netflix offered anonymized data for a contest for the best method to improve its movie recommendations, Arvind Narayanan and Vitaly Shmatikov showed that this data could in fact be used to identify Netflix subscribers.

In particular regarding open data, Andrew Simpson demonstrated that it is relatively easy to link statistical open data to individuals. In one case, names and addresses of councillors, and names, posts and salaries of senior public servants were uncovered by combining data from the British open data portal with other already available public data. The lack of consideration of other data in the public domain prior to publication of statistical open data thus led to the identification of individuals.

Combining datasets is at the core of de-anonymizing and de-aggregating data. Data that is non-identifiable today, may turn out be indirectly identifiable tomorrow. The more computing power and publicly available data, the easier it becomes to identify individuals in data. And when data can be related to individuals, data protection law kicks in.

What does this mean for open data providers? Open data providers should not just consider the identifiability of their open data in isolation. They should also take other publicly available data into account when selecting data that they want to offer as open data. That is a difficult task. Maybe open data is not such a great idea after all?

Open Data Workshop @ Geonovum

Geonovum (a semi-public organization devoting itself to providing better access to geo-information in the public sector) is hosting an open data workshop on November 9, 2011. Location: De Observant in Amersfoort.

Who will be there and what will they be talking about?

  • Marc de Vries (ePSI platform) will try to look into the future of open data.
  • Christopher Dittmann (Shell) will give a talk on the experience of availability/non-availability of open geospatial data.
  • Paul Suijkerbuijk (ICTU) will share his experience with national government open data platform.

Interactive sessions:

  • Johan van Arragon (Province of Zuid-Holland) will talk about the costs and benefits of open data.
  • Paul Hendriks and Peter-Jan Speerstra (Municipality of Rotterdam) will deal with the question of how to implement an open data policy.
  • Jens Riecken (Ministry of the Interior and Local Affairs NordRhein Westfalen, Germany) will explain how to utilize the wisdom of the crowds.
  • Kathleen Janssen will take a step back and will deal with legal, financial and practical issues that need to be tackled. I am particularly interested in this session.
  • Richard Blad will give a talk on how to organize an open data-community.

The full program can be found here: http://www.geonovum.nl/dossiers/kennissessies/opengeodata/programma.

I’ll be there. By the way, in the spirit of the open data philosophy: it’s free!

What a contrast: Google uses London open data for tube and bus directions, while Paris public transport operator kills public transport app

This month brought contrasting news on the openness of public transport information in two EU countries. The Telegraph celebrated Google’s mapping service for adding live public transport information and directions. The mobile version of Google Maps has a function that detects a user’s location and that direct him to the nearest tube station or bus stop. Another great function is the alert-function, which warns users to get off their bus or train when they have reached their destination. What’s Google’s secret?

The service relies on Transport for London’s open data platform, which allows developers direct access to data on public transport in the capital, including up-to-date details of roadworks and tube suspensions. Google did not pay for access to the data, which has been freely available since last June.

Around the same time in France, a similar public transport information was killed by the Paris public transport operator (RATP). CheckMyMetro is a free iPhone and Android app that lets French metro users connect to each other and allows them to share information on inter alia incidents and delays. The Paris public transport operator filed a complaint with Apple arguing that the traffic information in the CheckMyMetro app infringed the operator’s database rights. As a result, Apple asked the creator of CheckMyMetro to remove the app from the App Store:

Dear Sirs,

The RATP is a French public company in charge of Public Transports in the Paris area French.
The RATP is the author of the Paris Metro map and the owner of corresponding French design registration (INPI deposit n°06 5325 –Nov. 17th 2006). French and International law on copyright as well as French law on Design thus protect this map. Moreover, the RATP is the owner of the trademark # (INPI deposit n°92402043 – January 21st 1992).

The RATP is concerned with the application “Check my metro” proposed for downloading by the publisher LittleSphere on the App Store and the iTunessince we did not authorize any reproduction or distribution of the said design and trademark.

Moreover, this app embeds the traffic information of our wap site without prior authorization which constitutes an infringement on our rights as producer of database conferred by the French law.

Such reproductions and diffusions may then be considered as counterfeiting acts, and the RATP is entitled to enforce its rights within the French jurisdictions.

Consequently, we ask you to remove the application “Check my metro” by LittleSphere of the App Store and iTunes and to inform the publisher in the same way.

The app is back in the App Store, however, the public transport information has been removed.

Although the French government has started an open data initiative called ETALAB, the Paris public transport information is outside of the realm of the open data initiative because it is in the hands of the public transport operator. The creators of CheckMyMetro, however, are not waiting for the information to be open. They have started their own OpenStreetMap-like project for the Paris Metro at www.checkmymap.fr.

 

Update, 16:00h: I’ve replaced ‘public transport authority’ with ‘public transport operator‘.