Links, stories, images and articles about data-intensive research from Australia and beyond. An unofficial blog created by the Australian Research data Commons (Formerly: Australian National Data Service (ANDS), Nectar & RDS)) for the purpose of recording and sharing external content. ARDC does not endorse the content posted.
Wednesday, 28 March 2018
Hundreds of universities targeted in global data steal - University World News
http://www.universityworldnews.com/article.php?story=20180327132912436
Monday, 26 March 2018
Wednesday, 21 March 2018
Shifting to Data Savvy: The Future of Data Science In Libraries report
dScholarship@Pitt
Report release - 19 March, 2018
The Data Science in Libraries Project is funded by the Institute for Museum and Library Services (IMLS) and led by Matt Burton and Liz Lyon, School of Computing & Information, University of Pittsburgh; Chris Erdmann, North Carolina State University; and Bonnie Tijerina, Data & Society. The project explores the challenges associated with implementing data science within diverse library environments by examining two specific perspectives framed as ‘the skills gap,’ i.e. where librarians are perceived to lack the technical skills to be effective in a data-rich research environment; and ‘the management gap,’ i.e. the ability of library managers to understand and value the benefits of in-house data science skills and to provide organizational and managerial support.
This report primarily presents a synthesis of the discussions, findings, and reflections from an international, two-day workshop held in May 2017 in Pittsburgh, where community members participated in a program with speakers, group discussions, and activities to drill down into the challenges of successfully implementing data science in libraries. Participants came from funding organizations, academic and public libraries, nonprofits, and commercial organizations with most of the discussions focusing on academic libraries and library schools.
Report release - 19 March, 2018
The Data Science in Libraries Project is funded by the Institute for Museum and Library Services (IMLS) and led by Matt Burton and Liz Lyon, School of Computing & Information, University of Pittsburgh; Chris Erdmann, North Carolina State University; and Bonnie Tijerina, Data & Society. The project explores the challenges associated with implementing data science within diverse library environments by examining two specific perspectives framed as ‘the skills gap,’ i.e. where librarians are perceived to lack the technical skills to be effective in a data-rich research environment; and ‘the management gap,’ i.e. the ability of library managers to understand and value the benefits of in-house data science skills and to provide organizational and managerial support.
This report primarily presents a synthesis of the discussions, findings, and reflections from an international, two-day workshop held in May 2017 in Pittsburgh, where community members participated in a program with speakers, group discussions, and activities to drill down into the challenges of successfully implementing data science in libraries. Participants came from funding organizations, academic and public libraries, nonprofits, and commercial organizations with most of the discussions focusing on academic libraries and library schools.
What is Data Savvy?
A family of data science roles has been identified, which
can be characterised by the real-world requirements
for actual positions, as described in two related smallscale
studies (Lyon & Mattern, 2017; Lyon, Mattern,
Acker, & Langmead, 2015). The six roles are: data
archivist, data curator, data librarian, data analyst,
data engineer, and data journalist. While all of these
roles have been framed as data science roles, other
framing which comes from the corporate sector has
tended to describe only data analyst-type roles as
data scientists. However, in reality, there are a wide
gamut of roles—‘data savvy’ roles—that orbit within
and around the world of data scientists. Data savvy
librarians gain familiarity with the datasets, understand
technical methods and techniques, and speak multiple
disciplinary languages allowing them to work more
closely with researchers or the public. Some librarians
engage more deeply, becoming technically proficient
in data preparation and analysis, allowing them to
work with data, automate workflows, and become fully
embedded in research teams. In other words, data
science exists more or less on a spectrum, depends
on an institution’s size and mission, and spans work
requiring the deep statistical and software engineering
skills, to work that focuses on advocacy, policy,
communication, and data management. Being data
savvy is an essential ingredient of all of these roles.
http://d-scholarship.pitt.edu/33891/
http://d-scholarship.pitt.edu/33891/1/Shifting%20to%20Data%20Savvy.pdf
http://d-scholarship.pitt.edu/33891/
http://d-scholarship.pitt.edu/33891/1/Shifting%20to%20Data%20Savvy.pdf
Friday, 16 March 2018
Nature Career feature - Data management made simple
Nature magazine has a Career feature article focused on data management:
Data management made simple
Keeping your research data freely available is crucial for open science — and your funding could depend on it.
It includes:
Data management made simple
Keeping your research data freely available is crucial for open science — and your funding could depend on it.
It includes:
- 12 tips for writing a data management plan
- Who needs them?
- Where do I get help?
- Do plans vary across disciplines?
- Will they improve my science?
Labels:
academic publishing,
data,
data management,
data management plans,
DMP,
funder mandates,
funders,
Nature
Friday, 9 March 2018
What's going on in this graph
The NYTimes has a dedicated unit for producing educational materials built upon their articles. One interesting output from that unit ("The Learning Network") is "What's going on in this graph?", which may be a useful example for developing resources around or teaching digital/data/visual literacy.
This is an offshoot of their "What's going on in this picture" feature, where they present part of the information--e.g., a picture without a caption--and ask students to consider the context of the image, and comment in a public forum. In the case of the "Graph" version, one or more graphs are shown separate to the article in which they occurred. In addition to this, bits of information (e.g., labels) are sometimes removed and the students are asked to discuss what information might be shown by the graph. Later in the week, the missing information is revealed. As a learning activity, there are offline and online components, and the discussion and final "reveal" is supported by online forums with mediators.
For further background, see this interview at the "Data Stories" podcast.
This is an offshoot of their "What's going on in this picture" feature, where they present part of the information--e.g., a picture without a caption--and ask students to consider the context of the image, and comment in a public forum. In the case of the "Graph" version, one or more graphs are shown separate to the article in which they occurred. In addition to this, bits of information (e.g., labels) are sometimes removed and the students are asked to discuss what information might be shown by the graph. Later in the week, the missing information is revealed. As a learning activity, there are offline and online components, and the discussion and final "reveal" is supported by online forums with mediators.
For further background, see this interview at the "Data Stories" podcast.
Labels:
data literacy,
data visualisation,
digital literacy
Thursday, 8 March 2018
Data Storage Finder will help Cornell researchers evaluate options
This is pretty impressive! Possibly a little too involved for most researchers, but very useful for research support staff.
https://data.research.cornell.edu/content/data-storage-finder-will-help-cornell-researchers-evaluate-options

https://data.research.cornell.edu/content/data-storage-finder-will-help-cornell-researchers-evaluate-options
Cornell University - We are pleased to announce the launch of the Data Storage Finder, a self-service, interactive tool to help discover and evaluate data storage options. Cornell researchers can answer questions about their data needs to identify services based on features important to them, choose just those services they want to learn more about, or explore and compare them all, in one easy-to-use webpage:
https://finder.research.cornell.edu/storage.
Monday, 5 March 2018
Leeds blog post on publishers and data access statements
---------- Forwarded message ----------
From: Nick Sheppard <N.Sheppard@leeds.ac.uk>
Date: Sat, Mar 3, 2018 at 12:02 AM
Subject: Publishers and data access statements
To: RESEARCH-DATAMAN@jiscmail.ac.u k
From: Nick Sheppard <N.Sheppard@leeds.ac.uk>
Date: Sat, Mar 3, 2018 at 12:02 AM
Subject: Publishers and data access statements
To: RESEARCH-DATAMAN@jiscmail.ac.u
Hi all
We’ve posted a blog about some of the issues we’ve encountered with publishers and data access statements, grateful for any comment:
Thanks
Nick
Nick Sheppard
Research Data Management (RDM) Advisor I Research Data Leeds
Leeds University Library
0113 343 8956
PLOS Criteria for Recommended Data Repositories
Posted March 1, 2018 by PLOS ONE Editors in News & Policy
The Blog post from PLOS outlines a set of criteria that explain how PLOS ONE assess repositories for inclusion in their list of recommended repositories. These criteria follow FAIR principles on data openness, which PLOS encourages repository owners to adhere to when setting up and managing their repositories.
http://blogs.plos.org/everyone/2018/03/01/criteria-for-recommended-data-repositories/
The Blog post from PLOS outlines a set of criteria that explain how PLOS ONE assess repositories for inclusion in their list of recommended repositories. These criteria follow FAIR principles on data openness, which PLOS encourages repository owners to adhere to when setting up and managing their repositories.
http://blogs.plos.org/everyone/2018/03/01/criteria-for-recommended-data-repositories/
Friday, 2 March 2018
Amnesia, the data anonymization tool of OpenAIRE that allows to remove identifying information from data, is up and running in beta mode
https://www.openaire.eu/amnesia-is-up-and-running
"Amnesia is a flexible data anonymization tool that allows to remove identifying information from data."
"Amnesia does not only remove direct identifiers like names, SSNs, etc., but also transforms secondary identifiers like birth date and zip code so that individuals cannot be identified in the data. Amnesia supports k-anonymity and km-anonymity."
"Amnesia is a flexible data anonymization tool that allows to remove identifying information from data."
"Amnesia does not only remove direct identifiers like names, SSNs, etc., but also transforms secondary identifiers like birth date and zip code so that individuals cannot be identified in the data. Amnesia supports k-anonymity and km-anonymity."
COEL Specification: a privacy-by-design framework for the collection and processing of behavioural data
"We are pleased to announce that the Classification of Everyday Living Version 1.0 from the
OASIS COEL TC was approved as an OASIS Committee Specification on 25th February
2018.
Digital technologies are becoming integral to our daily lives. They promise huge benefits to society, but the resulting data creates many challenges for an individual's privacy.
The COEL Specification provides a privacy-by-design framework for the collection and processing of behavioural data. It is uniquely suited to the transparent use of dynamic data for personalised digital services, IoT applications where devices are collecting information about identifiable individuals and the coding of behavioural data in identity solutions."
https://17d7qj42wrunagbc02q1r561-wpengine.netdna-ssl.com/wp-content/uploads/2018/03/OASIS-COEL-CS-announcement.pdf
Digital technologies are becoming integral to our daily lives. They promise huge benefits to society, but the resulting data creates many challenges for an individual's privacy.
The COEL Specification provides a privacy-by-design framework for the collection and processing of behavioural data. It is uniquely suited to the transparent use of dynamic data for personalised digital services, IoT applications where devices are collecting information about identifiable individuals and the coding of behavioural data in identity solutions."
https://17d7qj42wrunagbc02q1r561-wpengine.netdna-ssl.com/wp-content/uploads/2018/03/OASIS-COEL-CS-announcement.pdf
-
A wide range of different services and processes can qualify as implementations of the COEL
Specification. Examples include, but are not limited to:
-
A Data Engine could use the COEL Specification to structure and manage the dynamic personal
data it is set up to receive and process;
-
A Personal Data Store could use the COEL Specification to structure the dynamic personal data
it is designed to manage;
-
A Personal Electronic Device could use the COEL Specification to communicate the personal
event data that it can output;
-
An Internet of Things (IoT) Device which interacts with identifiable individuals could use the
COEL Specification to communicate the personal event data that it can output;
-
An Identity Authority could use the COEL Specification to deliver and check unique
pseudonymised keys;
-
A Data Portability Exchange could use the COEL Specification to translate personal data stored
in another format into compliant Behavioural Atom data;
-
A User Interface could use the COEL Specification to code and interpret interactions with an
individual;
-
A Customer Relationship Manager (CRM) could use the COEL Specification to code, store and
analyse interactions with an individual.
Subscribe to:
Posts (Atom)