Wednesday, 17 June 2020

Complex data workflows contribute to the reproducibility crisis in science

Via O'Reilly Data Newsletter:

Complex data workflows contribute to the reproducibility crisis in science

70 independent teams, given the same data and a common analysis challenge, came to very different conclusions in a recent Stanford study that highlights the challenges of data analysis in an era of huge datasets and highly flexible processing workflows.

"Simply put, no two groups of researchers are necessarily crunching data the same way. And with so much data to get through and so many ways to process it, researchers can arrive at totally different conclusions."

Friday, 24 April 2020

Why Sharing Academic Publications Under “No Derivatives” Licenses is Misguided

From: ttps://creativecommons.org/2020/04/21/academic-publications-under-no-derivatives-licenses-is-misguided/


...
In this blog post, we explain that applying restrictive licenses to academic publications is a misguided approach to addressing concerns over academic integrity. Specifically, we make it clear that using Creative Commons “No Derivatives” (ND) licenses on academic publications is not only ill-advised for policing academic fraud but also and more importantly unhelpful to the dissemination of research, especially publicly-funded research. We also show that the safeguards in place within truly open licenses (like CC BY or CC BY-SA) are well-suited to curbing malicious academic behavior, above and beyond other existing recourses for academic fraud and similar abuses. 

...
Read the blog post: https://creativecommons.org/2020/04/21/academic-publications-under-no-derivatives-licenses-is-misguided/

Tuesday, 3 March 2020

Nature technology feature: Find a home for every imaging data set

Repositories let researchers store, share and access life‑science images — and maybe even extract new findings.

https://www.nature.com/articles/d41586-020-00594-4

"... researchers shouldn’t simply drop their data sets into small, project-specific archives or generic cloud storage. “Just dumping the data somewhere doesn’t mean people can use it,” Ellenberg explains. “You need to organize the data, you need to annotate it, and curate it.” "

Wednesday, 26 February 2020

What makes for good data "health"?

In a blog post on Towards Data Science, Barr Moses poses the question "How come we know everything about how well our data infrastructure is performing, but so little about whether the data is right?" and answers it by identifying five "data observability pillars" of "data health":
  • Freshness: is the data recent? When was the last time it was generated? What upstream data is included/omitted?
  • Distribution: is the data within accepted ranges? Is it properly formatted? Is it complete?
  • Volume: has all the data arrived?
  • Schema: what is the schema, and how has it changed? Who has made these changes and for what reasons?
  • Lineage: for a given data asset, what are the upstream sources and downstream assets which are impacted by it? Who are the people generating this data, and who is relying on it for decision making?

https://towardsdatascience.com/good-pipelines-bad-data-e55d9ba17920

Wednesday, 19 February 2020

What makes a machine learning paper reproducible?

From a blog post:
... we hear warnings that Artificial Intelligence (AI) and Machine Learning (ML) face their own reproducibility crises. This leads us to ask: is it true? It would seem hard to believe, as ML permeates every smart-device and intervenes evermore in our daily lives. From helpful hints on how to act like a polite human over email, to Elon Musk’s promise of self-driving cars next year, it seems like machine learning is indeed reproducible. How reproducible is the latest ML research, and can we begin to quantify what impacts its reproducibility?

https://thegradient.pub/independently-reproducible-machine-learning/

Licence friction: when two datasets collide

Licence Friction: A Tale of Two Datasets

OpenStreetmap want to include data from a dataset published by Sport England. Will it be possible? It seems the licensing situation is complex, even though both datasets are "open".
... we have two organisations both aiming to publish and use data for the public good. But, because of complexities around derived data and licence compatibilities, data that might otherwise be used in new, innovative ways is instead going unused.

https://blog.ldodds.com/2020/01/24/licence-friction-a-tale-of-two-datasets/

Thursday, 30 January 2020

Google Dataset Search out of beta

Google's Dataset Search (https://datasetsearch.research.google.com/) is now out of beta, with some new features added.
Across the web, there are millions of datasets about nearly any subject that interests you. If you’re looking to buy a puppy, you could  find datasets compiling complaints of puppy buyers or studies on puppy cognition. Or if you like skiing, you could find data on revenue of ski resorts or injury rates and participation numbers. Dataset Search has indexed almost 25 million of these datasets, giving you a single place to search for datasets and find links to where the data is. Over the past year, people have tried it out and provided feedback, and now Dataset Search is officially out of beta.

https://blog.google/products/search/discovering-millions-datasets-web/

Tuesday, 14 January 2020

Data management for data intensive research (i.e. eResearch) in a nutshell

Eleven tips for working with large data sets


“It’s a mindset,” says Teal, “treating data as a first-class citizen.

https://www.nature.com/articles/d41586-020-00062-z

Thursday, 5 December 2019

IEEE Spectrum: Documenting algorithm designs for machine learning

From IEEE Spectrum:

Hey, Data Scientists: Show Your Machine-Learning Work
Documenting software development is standard practice—the same should hold for algorithm design

In the last two years, the U.S. Food and Drug Administration has approved several machine-learning models to accomplish tasks such as classifying skin cancer and detecting pulmonary embolisms. But for the companies who built those models, what happens if the data scientist who wrote the algorithms leaves the organization?
In many businesses, an individual or a small group of data scientists is responsible for building essential machine-learning models. Historically, they have developed these models on their own laptops through trial and error, and pass it along for production when it works. But in that transfer, the data scientist might not think to pass along all the information about the model’s development. And if the data scientist leaves, that information is lost for good.
That potential loss of information is why experts in data science are calling for machine learning to become a formal, documented process overseen by more people inside an organization.

Read the rest of the article at:
https://spectrum.ieee.org/computing/software/hey-data-scientists-show-your-machinelearning-work

Monday, 25 November 2019

Nature: Google health-data scandal spooks researchers

"Academics seeking health data for research rather than commercial purposes must typically get approval from an ethical-review committee before they can start a project. The researchers also often strip identifying information from the records they work with. Commercial uses of personal data don’t necessarily undergo the same review"

https://www.nature.com/articles/d41586-019-03574-5


Tuesday, 29 October 2019

Nature News: Venice ‘time machine’ project suspended amid data row

https://www.nature.com/articles/d41586-019-03240-w

Disagreements among international partners leave plans to digitize the Italian city’s history in limbo.

Two key partners have suspended the Venice Time Machine project after reaching an impasse over issues surrounding open data and methodology. The State Archive of Venice and the Swiss Federal Institute of Technology in Lausanne (EPFL) say they have had to pause data collection, and the archive’s director has raised questions about the usability of the 8 terabytes of information that have already been collected.

Tuesday, 1 October 2019

DataCite brochure

DataCite have released a brochure describing who and what they are, as well as what DOIs are how you get them etc...

https://datacite.org/assets/DataCite_Brochure.pdf


Thursday, 12 September 2019

Everything a Data Scientist Should Know About Data Management

From towardsdatascience.com:

https://towardsdatascience.com/everything-a-data-scientist-should-know-about-data-management-6877788c6a42

By Phoebe Wong & Robert Bennett

To be a real “full-stack” data scientist, or what many bloggers and employers call a “unicorn,” you’ve to master every step of the data science process — all the way from storing your data, to putting your finished product (typically a predictive model) in production. But the bulk of data science training focuses on machine/deep learning techniques; data management knowledge is often treated as an afterthought. Data science students usually learn modeling skills with processed and cleaned data in text files stored on their laptop, ignoring how the data sausage is made. Students often don’t realize that in industry settings, getting the raw data from various sources to be ready for modeling is usually 80% of the work ...

Thursday, 8 August 2019

Microsoft releases data use agreements for open data

Microsoft has released three data use agreements for open data, and is asking for public comments/feedback before the end of September 2019:
  • Open Use of Data Agreement (O-UDA)
    • "Designed for use with open datasets which don't include personal data or data owned by a data provider. It is the most open and least restricted of the three first proposals."
  • Computational Use of Data Agreement (C-UDA)
    • "The goal of the C-UDA is to define a use of data sets for AI training purposes that contain third party materials, in a manner consistent with law."
  • Data Use Agreement for Open AI Model Development (DUA-OAI)
    • "The goal of the DUA-OAI is to create a template agreement that parties might use to share data to train an artificial intelligence (AI) model and then to make that trained AI model publicly available through an open source licensing structure."
From ZDNet:

Microsoft looks to 'do for data sharing what open source did for code'
https://www.zdnet.com/article/microsoft-looks-to-do-for-data-sharing-what-open-source-did-for-code/

The agreements are listed here:
https://news.microsoft.com/datainnovation/#data-use-agreements

Thursday, 16 May 2019

Decoding ‘Game of Thrones’ by way of data science

“With the final season of the television series ‘Game of Thrones’ upon us it is a good opportunity to take a closer look at the books that the series is based on. We will discover how a numerical processing of the books can help us reveal patterns that lie hidden in ‘A Song of Ice and Fire’.”

https://blog.usejournal.com/decoding-a-game-of-thrones-by-way-of-data-science-fd81e66d1255

Technical debt for data scientists

“We take shortcuts in developing a solution without an understanding of the risks and costs of those shortcuts, and without a realistic plan for how we’re going to pay back the debt ... The result is that data science projects become expensive or impossible to maintain as time goes on.”

This blog post explains how to tell if your projects are too deeply in debt and what to do about it:

https://blog.shotwell.ca/posts/2019-04-19-technical-debt-in-data-science/

Wednesday, 15 May 2019

Nature Feature: Data sharing and how it can benefit your scientific career

Open science can lead to greater collaboration, increased confidence in findings and goodwill between researchers.

https://www.nature.com/articles/d41586-019-01506-x

"The key ... is to practise “open science by design” ... For example, many researchers now keep data, computer code and other materials in web-based, interactive tools such as the popular Jupyter electronic notebook, which makes online archiving much easier."

"For early-career scientists who prefer producing data to managing it, Mons has this advice: “Go to a university that takes data stewardship seriously.”"

Thursday, 7 March 2019

BBC News report on AAAS meeting: Machine learning 'causing science crisis'

Interesting report on the BBC News web site, covering a presentation given at last month's meeting of the AAAS about how the (mis)use of machine learning is contributing to the “reproducibility crisis” in science:

AAAS: Machine learning 'causing science crisis'
By Pallab Ghosh, Science correspondent, BBC News, Washington

Techniques used by thousands of scientists to analyse data are producing results that are misleading and often completely wrong.
Dr Genevera Allen from Rice University in Houston said that the increased use of such systems was contributing to a “crisis in science”.
She warned scientists that if they didn’t improve their techniques they would be wasting both time and money. Her research was presented at the American Association for the Advancement of Science in Washington.

Read the full story:
https://www.bbc.com/news/science-environment-47267081

Tuesday, 19 February 2019

Elsevier releases report "Research Futures: drivers and scenarios for the next decade"

Elsevier has released a report explaining its part in improving “the information system supporting research” over the next decade...

https://www.elsevier.com/connect/elsevier-research-futures-report

Wednesday, 13 February 2019

Life sciences data steward function matrix - article

Life sciences data steward function matrix

 Salome Scholtens; Petronella Anbeek;  Jasmin Böhmer; Mirjam Brullemans-Spansier; Marije van der Geest;  Mijke Jetten;  Christine Staiger;  Inge Slouwerhof;  Celia W G van Gelder
Sufficient, high quality data steward expertise and capacity in projects and institutes is one of the necessities for FAIR data management in life- sciences and personalised medicine research. In a ZonMw funded project of UMCG, UMCU, Radboudumc, Radboud University and DTL, supported by the relevant national stakeholders, we are working to make the data steward function concrete, to create consensus on the function and required competencies and to develop tailored education. The overall project aim is to professionalise the data steward function within the life-sciences domain, with a special focus on the implementation of the FAIR data principles. All documents related to this project can be found on the Zenodo Collection “Towards a community-endorsed data steward profession description for life-science research”.
This publication contains the first project deliverable: a matrix, that may function as the basis for a common job description of a data steward that is broadly supported within the Dutch life-sciences community. In the next phase of the project, this matrix will be complemented by knowledge, skills and competencies of a data steward, which will be translated into concrete learning objectives. These in turn will be used to develop an education line and training material for data stewards (including a design for an eLearning module). Sustainable implementation and alignment with existing education will be ensured.


from: 
https://zenodo.org/record/2561723#.XGNSD88zauM