Thursday, 26 July 2018

O'Reilly Ideas blog: Doing good data science

The first in a new series of posts on data ethics at O'Reilly Ideas is an easy-to-read taste of the issues, including how data science relates to corporate culture:

https://www.oreilly.com/ideas/doing-good-data-science
“Data scientists, data engineers, AI and ML developers, and other data professionals need to live ethical values, not just talk about them.”

Monday, 23 July 2018

"FAIR in practice" - Jisc report now available


To provide a better understanding of current practice with regard to the use of FAIR principles in the UK and the potential these have to enhance research operations and output, Jisc assigned work to Robert Allen and David Hartland of Hapsis Innovation Ltd. They worked in close collaboration with esteemed experts in the UK such as Simon Coles, Susanna Sansone and Melissa Terras and investigated practices and views in various research disciplines.

The results have now been published in the report ‘FAIR in practice - Jisc report on the Findable Accessible Interoperable and Reuseable Data Principles’, that can be found using DOI 10.5281/zenodo.1245567

A blog on the report, titled ‘Open science is all very well but how do you make it FAIR in practice?’ by Rachel Bruce and Bas Cordewener, is available here.

Wednesday, 18 July 2018



Open Science By Design: Realizing a Vision for 21st Century Research released today by the National (USA) Academies of Sciences, Engineering and Medicine.

http://sites.nationalacademies.org/pga/brdi/open_science_enterprise/index.htm 

Monday, 16 July 2018

A funder-imposed data publication requirement seldom inspired data sharing - Plos article

Abstract


Growth of the open science movement has drawn significant attention to data sharing and availability across the scientific community. In this study, we tested the ability to recover data collected under a particular funder-imposed requirement of public availability. We assessed overall data recovery success, tested whether characteristics of the data or data creator were indicators of recovery success, and identified hurdles to data recovery. Overall the majority of data were not recovered (26% recovery of 315 data projects), a similar result to journal-driven efforts to recover data. Field of research was the most important indicator of recovery success, but neither home agency sector nor age of data were determinants of recovery. While we did not find a relationship between recovery of data and age of data, age did predict whether we could find contact information for the grantee. The main hurdles to data recovery included those associated with communication with the researcher; loss of contact with the data creator accounted for half (50%) of unrecoverable datasets, and unavailability of contact information accounted for 35% of unrecoverable datasets. Overall, our results suggest that funding agencies and journals face similar challenges to enforcement of data requirements. We advocate that funding agencies could improve the availability of the data they fund by dedicating more resources to enforcing compliance with data requirements, providing data-sharing tools and technical support to awardees, and administering stricter consequences for those who ignore data sharing preconditions.

Thursday, 12 July 2018

Data infrastructure literacy

http://journals.sagepub.com/doi/10.1177/2053951718786316

Jonathan Gray, Carolin Gerlitz, Liliana Bounegru. Data infrastructure literacy. Big Data & Society (2018) 5(2) https://doi.org/10.1177/2053951718786316
"A recent report from the UN makes the case for “global data literacy” in order to realise the opportunities afforded by the “data revolution”. Here and in many other contexts, data literacy is characterised in terms of a combination of numerical, statistical and technical capacities. In this article, we argue for an expansion of the concept to include not just competencies in reading and working with datasets but also the ability to account for, intervene around and participate in the wider socio-technical infrastructures through which data is created, stored and analysed – which we call “data infrastructure literacy”. We illustrate this notion with examples of “inventive data practice” from previous and ongoing research on open data, online platforms, data journalism and data activism. Drawing on these perspectives, we argue that data literacy initiatives might cultivate sensibilities not only for data science but also for data sociology, data politics as well as wider public engagement with digital data infrastructures. The proposed notion of data infrastructure literacy is intended to make space for collective inquiry, experimentation, imagination and intervention around data in educational programmes and beyond, including how data infrastructures can be challenged, contested, reshaped and repurposed to align with interests and publics other than those originally intended."