Wednesday, 26 February 2020

What makes for good data "health"?

In a blog post on Towards Data Science, Barr Moses poses the question "How come we know everything about how well our data infrastructure is performing, but so little about whether the data is right?" and answers it by identifying five "data observability pillars" of "data health":
  • Freshness: is the data recent? When was the last time it was generated? What upstream data is included/omitted?
  • Distribution: is the data within accepted ranges? Is it properly formatted? Is it complete?
  • Volume: has all the data arrived?
  • Schema: what is the schema, and how has it changed? Who has made these changes and for what reasons?
  • Lineage: for a given data asset, what are the upstream sources and downstream assets which are impacted by it? Who are the people generating this data, and who is relying on it for decision making?

https://towardsdatascience.com/good-pipelines-bad-data-e55d9ba17920

Wednesday, 19 February 2020

What makes a machine learning paper reproducible?

From a blog post:
... we hear warnings that Artificial Intelligence (AI) and Machine Learning (ML) face their own reproducibility crises. This leads us to ask: is it true? It would seem hard to believe, as ML permeates every smart-device and intervenes evermore in our daily lives. From helpful hints on how to act like a polite human over email, to Elon Musk’s promise of self-driving cars next year, it seems like machine learning is indeed reproducible. How reproducible is the latest ML research, and can we begin to quantify what impacts its reproducibility?

https://thegradient.pub/independently-reproducible-machine-learning/

Licence friction: when two datasets collide

Licence Friction: A Tale of Two Datasets

OpenStreetmap want to include data from a dataset published by Sport England. Will it be possible? It seems the licensing situation is complex, even though both datasets are "open".
... we have two organisations both aiming to publish and use data for the public good. But, because of complexities around derived data and licence compatibilities, data that might otherwise be used in new, innovative ways is instead going unused.

https://blog.ldodds.com/2020/01/24/licence-friction-a-tale-of-two-datasets/