Wednesday, 17 June 2020

Complex data workflows contribute to the reproducibility crisis in science

Via O'Reilly Data Newsletter:

Complex data workflows contribute to the reproducibility crisis in science

70 independent teams, given the same data and a common analysis challenge, came to very different conclusions in a recent Stanford study that highlights the challenges of data analysis in an era of huge datasets and highly flexible processing workflows.

"Simply put, no two groups of researchers are necessarily crunching data the same way. And with so much data to get through and so many ways to process it, researchers can arrive at totally different conclusions."

Friday, 24 April 2020

Why Sharing Academic Publications Under “No Derivatives” Licenses is Misguided

From: ttps://creativecommons.org/2020/04/21/academic-publications-under-no-derivatives-licenses-is-misguided/


...
In this blog post, we explain that applying restrictive licenses to academic publications is a misguided approach to addressing concerns over academic integrity. Specifically, we make it clear that using Creative Commons “No Derivatives” (ND) licenses on academic publications is not only ill-advised for policing academic fraud but also and more importantly unhelpful to the dissemination of research, especially publicly-funded research. We also show that the safeguards in place within truly open licenses (like CC BY or CC BY-SA) are well-suited to curbing malicious academic behavior, above and beyond other existing recourses for academic fraud and similar abuses. 

...
Read the blog post: https://creativecommons.org/2020/04/21/academic-publications-under-no-derivatives-licenses-is-misguided/

Tuesday, 3 March 2020

Nature technology feature: Find a home for every imaging data set

Repositories let researchers store, share and access life‑science images — and maybe even extract new findings.

https://www.nature.com/articles/d41586-020-00594-4

"... researchers shouldn’t simply drop their data sets into small, project-specific archives or generic cloud storage. “Just dumping the data somewhere doesn’t mean people can use it,” Ellenberg explains. “You need to organize the data, you need to annotate it, and curate it.” "

Wednesday, 26 February 2020

What makes for good data "health"?

In a blog post on Towards Data Science, Barr Moses poses the question "How come we know everything about how well our data infrastructure is performing, but so little about whether the data is right?" and answers it by identifying five "data observability pillars" of "data health":
  • Freshness: is the data recent? When was the last time it was generated? What upstream data is included/omitted?
  • Distribution: is the data within accepted ranges? Is it properly formatted? Is it complete?
  • Volume: has all the data arrived?
  • Schema: what is the schema, and how has it changed? Who has made these changes and for what reasons?
  • Lineage: for a given data asset, what are the upstream sources and downstream assets which are impacted by it? Who are the people generating this data, and who is relying on it for decision making?

https://towardsdatascience.com/good-pipelines-bad-data-e55d9ba17920

Wednesday, 19 February 2020

What makes a machine learning paper reproducible?

From a blog post:
... we hear warnings that Artificial Intelligence (AI) and Machine Learning (ML) face their own reproducibility crises. This leads us to ask: is it true? It would seem hard to believe, as ML permeates every smart-device and intervenes evermore in our daily lives. From helpful hints on how to act like a polite human over email, to Elon Musk’s promise of self-driving cars next year, it seems like machine learning is indeed reproducible. How reproducible is the latest ML research, and can we begin to quantify what impacts its reproducibility?

https://thegradient.pub/independently-reproducible-machine-learning/

Licence friction: when two datasets collide

Licence Friction: A Tale of Two Datasets

OpenStreetmap want to include data from a dataset published by Sport England. Will it be possible? It seems the licensing situation is complex, even though both datasets are "open".
... we have two organisations both aiming to publish and use data for the public good. But, because of complexities around derived data and licence compatibilities, data that might otherwise be used in new, innovative ways is instead going unused.

https://blog.ldodds.com/2020/01/24/licence-friction-a-tale-of-two-datasets/

Thursday, 30 January 2020

Google Dataset Search out of beta

Google's Dataset Search (https://datasetsearch.research.google.com/) is now out of beta, with some new features added.
Across the web, there are millions of datasets about nearly any subject that interests you. If you’re looking to buy a puppy, you could  find datasets compiling complaints of puppy buyers or studies on puppy cognition. Or if you like skiing, you could find data on revenue of ski resorts or injury rates and participation numbers. Dataset Search has indexed almost 25 million of these datasets, giving you a single place to search for datasets and find links to where the data is. Over the past year, people have tried it out and provided feedback, and now Dataset Search is officially out of beta.

https://blog.google/products/search/discovering-millions-datasets-web/

Tuesday, 14 January 2020

Data management for data intensive research (i.e. eResearch) in a nutshell

Eleven tips for working with large data sets


“It’s a mindset,” says Teal, “treating data as a first-class citizen.

https://www.nature.com/articles/d41586-020-00062-z