Wednesday, 28 March 2018


Public Lecture with Trish Greenhalgh & Anne Kelso
Measuring the Impact of Research

This Webcast was held on Monday, 19th March 2018


Hundreds of universities targeted in global data steal - University World News

http://www.universityworldnews.com/article.php?story=20180327132912436

Cambridge Analytica controversy must spur researchers to update data ethics

A scandal over an academic’s use of Facebook data highlights the need for research scrutiny.
Nature 555, 559-560 (2018)
doi: 10.1038/d41586-018-03856-4 

Monday, 26 March 2018




UK and US researchers ‘less likely to share research data’


Large worldwide survey suggests scholars in continental Europe more readily provide data alongside results



Wednesday, 21 March 2018

Shifting to Data Savvy: The Future of Data Science In Libraries report

dScholarship@Pitt
Report release - 19 March, 2018

The Data Science in Libraries Project is funded by the Institute for Museum and Library Services (IMLS) and led by Matt Burton and Liz Lyon, School of Computing & Information, University of Pittsburgh; Chris Erdmann, North Carolina State University; and Bonnie Tijerina, Data & Society. The project explores the challenges associated with implementing data science within diverse library environments by examining two specific perspectives framed as ‘the skills gap,’ i.e. where librarians are perceived to lack the technical skills to be effective in a data-rich research environment; and ‘the management gap,’ i.e. the ability of library managers to understand and value the benefits of in-house data science skills and to provide organizational and managerial support.

This report primarily presents a synthesis of the discussions, findings, and reflections from an international, two-day workshop held in May 2017 in Pittsburgh, where community members participated in a program with speakers, group discussions, and activities to drill down into the challenges of successfully implementing data science in libraries. Participants came from funding organizations, academic and public libraries, nonprofits, and commercial organizations with most of the discussions focusing on academic libraries and library schools.

What is Data Savvy? 
A family of data science roles has been identified, which can be characterised by the real-world requirements for actual positions, as described in two related smallscale studies (Lyon & Mattern, 2017; Lyon, Mattern, Acker, & Langmead, 2015). The six roles are: data archivist, data curator, data librarian, data analyst, data engineer, and data journalist. While all of these roles have been framed as data science roles, other framing which comes from the corporate sector has tended to describe only data analyst-type roles as data scientists. However, in reality, there are a wide gamut of roles—‘data savvy’ roles—that orbit within and around the world of data scientists. Data savvy librarians gain familiarity with the datasets, understand technical methods and techniques, and speak multiple disciplinary languages allowing them to work more closely with researchers or the public. Some librarians engage more deeply, becoming technically proficient in data preparation and analysis, allowing them to work with data, automate workflows, and become fully embedded in research teams. In other words, data science exists more or less on a spectrum, depends on an institution’s size and mission, and spans work requiring the deep statistical and software engineering skills, to work that focuses on advocacy, policy, communication, and data management. Being data savvy is an essential ingredient of all of these roles.

http://d-scholarship.pitt.edu/33891/

http://d-scholarship.pitt.edu/33891/1/Shifting%20to%20Data%20Savvy.pdf

Friday, 16 March 2018

Nature Career feature - Data management made simple

Nature magazine has a Career feature article focused on data management:

Data management made simple
Keeping your research data freely available is crucial for open science — and your funding could depend on it.

It includes:

  • 12 tips for writing a data management plan
  • Who needs them?
  • Where do I get help?
  • Do plans vary across disciplines?
  • Will they improve my science?

Friday, 9 March 2018

What's going on in this graph

The NYTimes has a dedicated unit for producing educational materials built upon their articles. One interesting output from that unit ("The Learning Network") is "What's going on in this graph?", which may be a useful example for developing resources around or teaching digital/data/visual literacy.

This is an offshoot of their "What's going on in this picture" feature, where they present part of the information--e.g., a picture without a caption--and ask students to consider the context of the image, and comment in a public forum. In the case of the "Graph" version, one or more graphs are shown separate to the article in which they occurred. In addition to this, bits of information (e.g., labels) are sometimes removed and the students are asked to discuss what information might be shown by the graph. Later in the week, the missing information is revealed. As a learning activity, there are offline and online components, and the discussion and final "reveal" is supported by online forums with mediators.

For further background, see this interview at the "Data Stories" podcast.

Thursday, 8 March 2018

Data Storage Finder will help Cornell researchers evaluate options

This is pretty impressive! Possibly a little too involved for most researchers, but very useful for research support staff.

https://data.research.cornell.edu/content/data-storage-finder-will-help-cornell-researchers-evaluate-options

Cornell University - We are pleased to announce the launch of the Data Storage Finder, a self-service, interactive tool to help discover and evaluate data storage options. Cornell researchers can answer questions about their data needs to identify services based on features important to them, choose just those services they want to learn more about, or explore and compare them all, in one easy-to-use webpage: 
https://finder.research.cornell.edu/storage.

Screenshot of Data Storage Finder with some services selected and some deselected to show how the comparison chooser works.

Monday, 5 March 2018

Leeds blog post on publishers and data access statements

---------- Forwarded message ----------
From: Nick Sheppard <N.Sheppard@leeds.ac.uk>
Date: Sat, Mar 3, 2018 at 12:02 AM
Subject: Publishers and data access statements
To: RESEARCH-DATAMAN@jiscmail.ac.uk


Hi all

We’ve posted a blog about some of the issues we’ve encountered with publishers and data access statements, grateful for any comment:


Thanks

Nick





Nick Sheppard
Research Data Management (RDM) Advisor I Research Data Leeds
Leeds University Library
0113 343 8956

PLOS Criteria for Recommended Data Repositories

Posted March 1, 2018 by PLOS ONE Editors in News & Policy

The Blog post from PLOS outlines a set of criteria that explain how PLOS ONE assess repositories for inclusion in their list of recommended repositories. These criteria follow FAIR principles on data openness, which PLOS encourages repository owners to adhere to when setting up and managing their repositories. 

http://blogs.plos.org/everyone/2018/03/01/criteria-for-recommended-data-repositories/

Friday, 2 March 2018

Amnesia, the data anonymization tool of OpenAIRE that allows to remove identifying information from data, is up and running in beta mode

https://www.openaire.eu/amnesia-is-up-and-running

"Amnesia is a flexible data anonymization tool that allows to remove identifying information from data."

"Amnesia does not only remove direct identifiers like names, SSNs, etc., but also transforms secondary identifiers like birth date and zip code so that individuals cannot be identified in the data. Amnesia supports k-anonymity and km-anonymity."

COEL Specification: a privacy-by-design framework for the collection and processing of behavioural data

"We are pleased to announce that the Classification of Everyday Living Version 1.0 from the OASIS COEL TC was approved as an OASIS Committee Specification on 25th February 2018.
Digital technologies are becoming integral to our daily lives. They promise huge benefits to society, but the resulting data creates many challenges for an individual's privacy.
The COEL Specification provides a privacy-by-design framework for the collection and processing of behavioural data. It is uniquely suited to the transparent use of dynamic data for personalised digital services, IoT applications where devices are collecting information about identifiable individuals and the coding of behavioural data in identity solutions."


https://17d7qj42wrunagbc02q1r561-wpengine.netdna-ssl.com/wp-content/uploads/2018/03/OASIS-COEL-CS-announcement.pdf


  1. A wide range of different services and processes can qualify as implementations of the COEL Specification. Examples include, but are not limited to:
    • A Data Engine could use the COEL Specification to structure and manage the dynamic personal data it is set up to receive and process;
    • A Personal Data Store could use the COEL Specification to structure the dynamic personal data it is designed to manage;
    • A Personal Electronic Device could use the COEL Specification to communicate the personal event data that it can output;
    • An Internet of Things (IoT) Device which interacts with identifiable individuals could use the COEL Specification to communicate the personal event data that it can output;
    • An Identity Authority could use the COEL Specification to deliver and check unique pseudonymised keys;
    • A Data Portability Exchange could use the COEL Specification to translate personal data stored in another format into compliant Behavioural Atom data;
    • A User Interface could use the COEL Specification to code and interpret interactions with an individual;
    • A Customer Relationship Manager (CRM) could use the COEL Specification to code, store and analyse interactions with an individual.