Wednesday, 26 February 2020

What makes for good data "health"?

In a blog post on Towards Data Science, Barr Moses poses the question "How come we know everything about how well our data infrastructure is performing, but so little about whether the data is right?" and answers it by identifying five "data observability pillars" of "data health":
  • Freshness: is the data recent? When was the last time it was generated? What upstream data is included/omitted?
  • Distribution: is the data within accepted ranges? Is it properly formatted? Is it complete?
  • Volume: has all the data arrived?
  • Schema: what is the schema, and how has it changed? Who has made these changes and for what reasons?
  • Lineage: for a given data asset, what are the upstream sources and downstream assets which are impacted by it? Who are the people generating this data, and who is relying on it for decision making?

https://towardsdatascience.com/good-pipelines-bad-data-e55d9ba17920

No comments:

Post a Comment