A Data Management Record (DMR) is an active record of working data that can be curated and form part of a Data Management Plan (DMP).
Why are we placing the emphasis on planning for good data management for possible future projects, when that time could be spent curating working research data from projects you are currently working on?
Data Management Plans (DMP's) have become synonymous with
research data management best-practice, often with the view of the more
comprehensive the better. Universities and organisations have created lengthy
checklists of questions which over time have grown into labyrinthine forms
requiring a significant amount of time to complete with any kind of rigour. The
Australian researcher faced with no mandate or reward to complete a lengthy DMP
resorts to writing a throw-away DMP or none at all.
It is thus unsurprising that no study has shown any benefit
from completing traditional DMPs. In fact, when asked, the largest DMP provider
in the UK mentioned that researchers rarely go back to visit the documentation
created to meet funder mandates. Researchers may write DMPs to satisfy funders,
but they don't refer back to them or keep them active or updated. There is no
inherent reason to because DMP's are typically not linked to any reward or
outcome. They become static artefacts which, due to the very nature of
research, quickly become out of date.
This lack of up-front planning of data management leads to
multiple and potentially serious problems for the researcher, their colleagues,
institution and indeed science in general.
This is evidenced by the reproducibility crisis we are now facing[1] A recent survey of 1500 participants
regarding this problem concluded that 34% of researchers do not have
established procedures in order to increase reproducibility[2].
So, what is the right alternative to present researchers
with? Arguably not no DMP at all, but perhaps something more relevant to the
research process. Something practical that can minimise the risks related to
“bad” data management practice. Researchers need to actively manage data at the
project level, with adequate documentation at all stages of the data lifecycle,
not just the planning stage. Therefore, we would like to propose new
terminology for a more relevant, useful, actionable metadata record. A Data
Management Record (DMR) can be further broken down as either active (aDMR) or
curated (cDMR).
By shifting the wording from plan to record the result
is dramatic. The DMP becomes more than a plan. It becomes a critical record
through which the researcher directly manages their data and potentially
engages with other university systems. Given that multiple parts of the
proposed DMR relate to systems and processes around data much of it can also be
auto-populated as the data lifecycle progresses.
DMRs are associated with research projects and as such there
is potential for institutions to utilise them for multiple purposes. They can
be the means by which a researcher manages storage for a research project. They
could control who gets access to the data at any point in time and collect
relevant metadata about the research project throughout the entire data lifecycle,
from conception or idea stage to publication and archiving. This information
likely exists for the majority of research projects already, but is scattered
across multiple systems with little linking information. The linking of this
data via an identifier such as RAiD[3] has significant benefits for
long term provenance, and as such, the DMR should also contain other persistent
identifiers (PIDs) such as ORCID, GRID and DOI where relevant.
Any and all metadata entered into the DMR will be stored and
add value beyond the planning stage. Indeed, it is envisaged to be published
with the final data record in the institutional repository at the completion of
a project or when a publication results from the data.
It is time - we need to grow beyond talking about
traditional DMPs. What we have now is the next generation of DMPs which require
and deserve a name that accurately describes that they do more for everyone involved in the research process. Data
Management Records are project based active and actionable online “forms” that
ultimately turn into curated Data Management Records upon publication of the
metadata they hold, along with the associated PID to the research data
generated by the project.
By creating DMR's at the start of a research project and
then adding information as the project progresses, it gives us the opportunity
to formalise the data chain of command from day one. This chain of command can
be built based upon who has access, which department the data is associated
with, where it is or has been stored and which institutions have oversight.
This will dramatically ease data access problems when they arise as this
information will be ready and to hand to parties that require it. This approach
will reduce the propensity towards orphaned data, and data silos as it
challenges attitudes towards data ownership by associating all data with
projects rather than individuals.
Our take is that all projects likely already have a DMR, be
it written down or not. It's somewhere. By formalising the notion of DMRs and
assisting researchers to capture it as part of normal process everyone’s a
winner!
Helen Morgan - helen.morgan@uq.edu.au
Dr Andrew Janke - andrew.janke@uq.edu.au
