Wednesday, 29 July 2015

Institute of Medicine's Landmark Report on Sharing Clinical Trial Data

In January this year, the Institute of Medicine in the U.S. released a landmark report on sharing clinical trail data. 
"Although clinical trials generate vast amounts of data, a large por­tion is never published or made available to other researchers. Data sharing could advance scientific discovery and improve clinical care by maximizing the knowl­edge gained from data collected in trials, stimulating new ideas for research, and avoiding unnecessarily duplicative trials. In response to 23 public- and private-sector sponsors in the United States and abroad, the Institute of Medicine (IOM) assembled a committee to develop guiding principles and a practical framework for the responsible sharing of clinical trial data. In its report, Sharing Clinical Trial Data: Maxi­mizing Benefits, Minimizing Risk, the committee concludes that sharing data is in the public interest, but a multi-stakeholder effort is needed to develop a culture, infrastructure, and policies that will foster responsible sharing—now and in the future."
http://iom.nationalacademies.org/Reports/2015/Sharing-Clinical-Trial-Data.aspx
 
Although clinical trials generate vast amounts of data, a large por­tion is never published or made available to other researchers. Data sharing could advance scientific discovery and improve clinical care by maximizing the knowl­edge gained from data collected in trials, stimulating new ideas for research, and avoiding unnecessarily duplicative trials. In response to 23 public- and private-sector sponsors in the United States and abroad, the Institute of Medicine (IOM) assembled a committee to develop guiding principles and a practical framework for the responsible sharing of clinical trial data. In its report, Sharing Clinical Trial Data: Maxi­mizing Benefits, Minimizing Risk, the committee concludes that sharing data is in the public interest, but a multi-stakeholder effort is needed to develop a culture, infrastructure, and policies that will foster responsible sharing—now and in the future. - See more at: http://iom.nationalacademies.org/Reports/2015/Sharing-Clinical-Trial-Data.aspx#sthash.5nFatiOk.dpuf
Although clinical trials generate vast amounts of data, a large por­tion is never published or made available to other researchers. Data sharing could advance scientific discovery and improve clinical care by maximizing the knowl­edge gained from data collected in trials, stimulating new ideas for research, and avoiding unnecessarily duplicative trials. In response to 23 public- and private-sector sponsors in the United States and abroad, the Institute of Medicine (IOM) assembled a committee to develop guiding principles and a practical framework for the responsible sharing of clinical trial data. In its report, Sharing Clinical Trial Data: Maxi­mizing Benefits, Minimizing Risk, the committee concludes that sharing data is in the public interest, but a multi-stakeholder effort is needed to develop a culture, infrastructure, and policies that will foster responsible sharing—now and in the future. - See more at: http://iom.nationalacademies.org/Reports/2015/Sharing-Clinical-Trial-Data.aspx#sthash.5nFatiOk.dpuf
Although clinical trials generate vast amounts of data, a large por­tion is never published or made available to other researchers. Data sharing could advance scientific discovery and improve clinical care by maximizing the knowl­edge gained from data collected in trials, stimulating new ideas for research, and avoiding unnecessarily duplicative trials. In response to 23 public- and private-sector sponsors in the United States and abroad, the Institute of Medicine (IOM) assembled a committee to develop guiding principles and a practical framework for the responsible sharing of clinical trial data. In its report, Sharing Clinical Trial Data: Maxi­mizing Benefits, Minimizing Risk, the committee concludes that sharing data is in the public interest, but a multi-stakeholder effort is needed to develop a culture, infrastructure, and policies that will foster responsible sharing—now and in the future. - See more at: http://iom.nationalacademies.org/Reports/2015/Sharing-Clinical-Trial-Data.aspx#sthash.5nFatiOk.dpuf
Although clinical trials generate vast amounts of data, a large por­tion is never published or made available to other researchers. Data sharing could advance scientific discovery and improve clinical care by maximizing the knowl­edge gained from data collected in trials, stimulating new ideas for research, and avoiding unnecessarily duplicative trials. In response to 23 public- and private-sector sponsors in the United States and abroad, the Institute of Medicine (IOM) assembled a committee to develop guiding principles and a practical framework for the responsible sharing of clinical trial data. In its report, Sharing Clinical Trial Data: Maxi­mizing Benefits, Minimizing Risk, the committee concludes that sharing data is in the public interest, but a multi-stakeholder effort is needed to develop a culture, infrastructure, and policies that will foster responsible sharing—now and in the future. - See more at: http://iom.nationalacademies.org/Reports/2015/Sharing-Clinical-Trial-Data.aspx#sthash.5nFatiOk.dpuf
Although clinical trials generate vast amounts of data, a large por­tion is never published or made available to other researchers. Data sharing could advance scientific discovery and improve clinical care by maximizing the knowl­edge gained from data collected in trials, stimulating new ideas for research, and avoiding unnecessarily duplicative trials. In response to 23 public- and private-sector sponsors in the United States and abroad, the Institute of Medicine (IOM) assembled a committee to develop guiding principles and a practical framework for the responsible sharing of clinical trial data. In its report, Sharing Clinical Trial Data: Maxi­mizing Benefits, Minimizing Risk, the committee concludes that sharing data is in the public interest, but a multi-stakeholder effort is needed to develop a culture, infrastructure, and policies that will foster responsible sharing—now and in the future. - See more at: http://iom.nationalacademies.org/Reports/2015/Sharing-Clinical-Trial-Data.aspx#sthash.5nFatiOk.dpuf

Thursday, 23 July 2015

"Tools and Services over Data" Workshop

On 24 June 2015, the Australian National Data Service (ANDS) facilitated a workshop “Tools and Services over Data” at the ANU University House to initiate a discussion between data providers and data service providers about how to more efficiently and effectively connect data to tools and services over data, and to identify common activities that could be coordinated across NCRIS capabilities. The motivation to hold this workshop was twofold: 1) NCRIS-funded capabilities are currently planning the next 1-2 years' activities, so this is a good time to look for synergies 2) At the same time, all of the NCRIS e-Research capabilities, and a number of the NCRIS data-intensive capabilities are working on activities with an explicit data focus. Thus, ANDS organised and facilitated a one day workshop  to discuss these issues, bringing together twenty-two people from NCRIS capabilities (ANDS, National eResearch Collaboration Tools and Resources (NeCTAR), Terrestrial Ecosystem Research Network (TERN), National Computational Infrastructure (NCI), Integrated Marine Observing System (IMOS), Bioplatforms Australia (BPA), Atlas of Living Australia (ALA), NeCTAR Virtual Labs (Biodiversity and Climate Change Virtual Laboratory (BCCVL) , Virtual Geophysics Laboratory (VGL), Characterisation Virtual Laboratory (CVL), and Genomics Virtual Laboratory (GVL)), Geoscience Australia (GA), Australian eResearch Organisations (AeRO), and Queensland Cyber Infrastructure Foundation (QCIF). 
The workshop had three major components: presentations, group breakout sessions, and reporting back / identifying next steps.  In the presentation session, participants each gave 3 minutes’ presentations on: 1) the main services/tools over data that they have provided to their users (or funded in the case of NeCTAR), 2) success stories of other groups using their services/tools, and 3) examples of more coordinated approaches to tools/services over data that they are aware of.  Six of the common topics emerging from these presentations were then selected for the group breakout sessions. Two rounds of breakout sessions were held with three topics discussed in each session. During the group breakout sessions, each group first discussed possible pilot activities and what participants (and others) could do collectively. A representative of each group reported back to the large group.  The workshop ended with a reporting back / next steps session to identify future actions and timelines.

The six topics selected for group discussion, together with highlights extracted from the group discussion and reporting back sessions, were:   
  1. Add-on services - semantic web, provenance, vocabularies, linked data & processing
    • The Semantic Web tries to achieve true discovery, there are already are a large number of value-adding tools such as provenance, proper use of URLs, and vocabulary services.
    • There are new services and tools available in the semantic space, although at different maturing stages.  Provenance may be the next to pickup and work with. ANDS, NeCTAR and other NCRIS capabilities have funded provenance-related elements in different projects, but we need a centralised discussion of the ongoing services in the provenance space.
    • We are getting more and more dependent on vocabulary service and provenance, we need to provide an ongoing network of vocabulary services in conjunction with data.gov.au and other related agencies.
  2. Data formats, services and discovery
    • Data should be stored in open and flexible formats where good conversion tools exist that retain all metadata, proprietary formats are OK where they are well documented and an API or conversion tool exists if this is the only way to maintain all metadata.
    • It is too expensive to maintain data  in multiple formats and at multiple sites. Costs include not only storage but also support etc.
    • There should be a continuous discussion about Data as a Service, and use of common services to access common data formats (think of data delivery formats, not data storage formats).
  3. Sampling subsets of data
    • Caching subset of data leads to needing alerts when the source data changes. Provenance information can help to determine when an alert is needed.
    • A cache of data close to the compute resource can physically integrate data from multiple sources, but this leads to a challenge of how to propagate aggregated provenance information with aggregated data.
    • There is a need to aggregate data across different sensors measuring the same observation in different ways.  Providing aggregated data is a challenge due to differences in calibration. How to trust aggregated data and provide all the provenance information is also non-trivial.
  4. Data fusion challenges
    • Data to be integrated may be from multiple platforms, disciplines, model-observations, license regimes,  citations, and different measurement protocols/standards/resolutions etc.
    • One challenge is how to efficiently describe the data that allows interpretation by others (people and machines)
    • Licensing can become a show-stopper when fusing data under different licensing regimes. There is a need to sort out a national licensing framework for data to make this less likely.
  5. Access to data and storage issues - should we make data locally or remotely accessible?
    • Some data are hosted overseas, for example, the Cancer Genome Atlas. Setting up a proxy on request to mirror the data may not solve the problem unless there is enough computing capacity close to the data for researchers’ customised analysis pipelines.
    • ALA caches spatial data for local indexing. It would be good to cache spatial data at NCI so that everyone could use it and ALA could run their indexer there. What would be best is a centralised authoritative storage system that manages data files and index layers.
    • BCCVL takes a copy of climate data irregularly, and also a direct copy of ALA to allow more complex queries; but don’t want to maintain the mirror of ALA.
  6. Where is the User Interface - should we provide users tools and services over their own desktop, cloud, or virtual desktop?
    • There is no single best solution; different users have different needs. For example, Biologists use Galaxy, Bioinformaticians want to use Desktop in the cloud. Having the data pre-attached makes a huge difference.
    • We may need different efficiency  criteria/measures: NCI uses use of cores - not working for everything. People will use cloud inefficiently, which is part of the cost. Running local IT is also often inefficient.

Finally, the following possible activities were identified in the reporting back / next steps session (organisation responsible for next steps in parentheses):
  • Vocabulary: it is good to have a national project to bring national agencies together. To complement that, this is to be done in conjunction with data.gov.au, the Australian Linked data group, and other research organisations. We need to think about what a national vocabulary is service. (ANDS)
  • Provenance: get people from provenance projects together to think about where we are going next. (ANDS)
  • Software registry: how to describe software and setup a software registry. (RDS may run a workshop on this topic)
  • Look at what data formats and services from community. What are data services for common format? How to discover service? (RDS, VLs)
  • A framework for data licenses applicable when fusing data together (ANDS + AusGoal)
  • Education program for data providers. (ANDS+NeCTAR).  
  • Define minimum description for M2M negotiation including license requirement. (Hamish Holewa from BCCVL will put up a case study.)
  • Improve core allocation process. (NeCTAR)

In conclusion, this was a great information gathering session. Attendees indicated interest of getting together at this year’s eResearch Australasia conference. ANDS is planning to organise a workshop on provenance in late August to bring people from provenance projects together to discuss about where we are going next.

If you would like to contribute to the discussion of one or more of the above topics, please contact Stefanie Kethers (stefanie.kethers@ands.org.au) or Hamish Holewa (hholewa@gmail.com).

Recent Brief on Data Management Policies by Canadian Funding Agencies

A recent comprehensive brief prepared by Kathleen Shearer on behalf of the three funding agencies in Canada suggests that the United Kingdom and the United States are most advanced in their research data management policies.  But the report also highlights that
"Despite the lack of strong policies, Australia is still considered one of the leaders in RDM in that it has made major investments in its services and infrastructure. In 2007, the Australian Government through the National Collaborative Research Infrastructure Strategy Program created the Australian National Data Service (ANDS). ANDS invests in and hosts a number of local and centralized services, including Research Data Australia, a national discovery service to promote visibility of research data collections."
 http://www.science.gc.ca/default.asp?lang=En&n=1E116DB8-1 

Monday, 20 July 2015

Pre-print: The role of Twitter in the life cycle of a scientific publication

Abstract
Twitter is a micro-blogging social media platform for short messages that can have a long-term impact on how scientists create and publish ideas. We investigate the usefulness of twitter in the development and distribution of scientific knowledge. At the start of the 'life cycle' of a scientific publication, twitter provides a large virtual department of colleagues that can help to rapidly generate, share and refine new ideas. As ideas become manuscripts, twitter can be used as an informal arena for the pre-review of works in progress. Finally, tweeting published findings can communicate research to a broad audience of other researchers, decision makers, journalists and the general public that can amplify the scientific and social impact of publications. However, there are limitations, largely surrounding issues of intellectual property and ownership, inclusiveness and misrepresentations of science ‘sound bites’. Nevertheless, we believe twitter is a useful social media tool that can provide a valuable contribution to scientific publishing in the 21st century.
 


Friday, 17 July 2015

Birth and Death Records Help Create Macroscopic View of Cultural History

Art historian Maximilian Schich and his co-authors created a macroscopic view of cultural history by analysing birth and death records of 150,000 notable individuals over a period of 2000 years.   They sourced their data from Freebase.com, the General Artist Lexicon, and the Getty Union List of Artist Names.  The findings were published in "A network framework of cultural history", Science 1 August 2014: Vol. 345 no. 6196 pp. 558-562 DOI: 10.1126/science.1240064.  A video created in conjunction with the paper and the same underlying data has received more than a million view on YouTube.  Free access to the paper and the video is available through:


This is yet another example of how access to data can advance the state of research, innovation, and multi-disciplinary collaboration.
 
Editor's Summary:
Sociologists and anthropologists study the growth and evolution of human culture, but it is hard to measure cultural interactions on a historical time scale. Schich et al. developed a tool for extracting information about cultural history from simple but large sets of birth and death records. A network of cultural centers connected via the birth and death of more than 150,000 notable individuals revealed human mobility patterns and cultural attraction dynamics. Patterns of city growth over a period of 2000 years differed between countries, but the distribution of birth-to-death distances remained unchanged over more than eight centuries.


 

Scholarly profiles

Impact of Social Sciences – What will the scholarly profile page of the future look like? Provision of metadata is enabling experimentation.

A nice analysis, with a table comparing the features, of scholarly profile sites like ORCID, Researchgate etc.

http://blogs.lse.ac.uk/impactofsocialsciences/2015/07/16/scholarly-profile-of-the-future/

U.K.'s Economic and Social Research Council Research Data Policy Updated

The U.K.'s Economic and Social Research Council research data policy has been updated earlier in March this year.  This is clearly relevant to Australian researchers collaborating with researchers funded through the ESRC program.

"All data created or repurposed during the lifetime of an ESRC grant must be made available for re-use or archiving within three months of the end of the grant. Grant holders must provide metadata for resource discovery via the UK Data Service to maximise the discoverability of ESRC data assets."
http://www.esrc.ac.uk/about-esrc/information/data-policy.aspx
 

Wednesday, 15 July 2015

European Commission's View on Increasing Access to Research Information

In this Communication to the European Parliament, the European Commission highlights the importance of increasing public access to research information:
"It is estimated that government investments of $3.8 billion in the Human Genome Project, a U.S. co-ordinated research endeavour including major European contributions, have had an economic impact worth $796 billion, created 310 000 jobs and launched the genome revolution."
https://ec.europa.eu/research/science-society/document_library/pdf_06/era-communication-towards-better-access-to-scientific-information_en.pdf



Tuesday, 14 July 2015

ODI 2 short videos

  1. What can open data do for you?
  2. Open shared and closed data explained
These 2, two minute videos explain simply the power and uses of open data.

Saturday, 11 July 2015

ARC's Presentation on ORCID, Open Access and Research Data

The Australian Research Council's recent presentation by Mr Justin Withers (Director, Policy and Integrity) reiterates ARC's support in the use of ORCID,  the commitment to the Open Access policy, and the commitment to maximize the benefits from ARC-funded research through greater access to research data.

http://www.arc.gov.au/sites/default/files/filedepot/Public/Media%20&%20Resources%20Centre/Presentations/2015_Presentations/p20150702_ORCID_UoN_JustinWithers.pdf


Friday, 10 July 2015

The U.K. Jisc-ARMA ORCID Pilot Project

The U.K. has recently concluded their ORCID Pilot Project.  The final report, ORCID Institutional Implementation Cost Benefit Analysis Report is now available online.   A checklist summary of lessons learned is also available.  A national consortium for ORCID is also underway in the U.K.

http://repository.jisc.ac.uk/6025/2/Jisc-ARMA-ORCID_final_report.pdf

http://repository.jisc.ac.uk/6025/3/Jisc-ARMA-ORCID-checklist-for-HEIs.pdf 

https://www.jisc.ac.uk/news/national-consortium-for-orcid-set-to-improve-uk-research-visibility-and-collaboration-23-jun


The Role of Research Data in Public Health Policies - the Case of Tamiflu

In this episode of Catalyst produced by the ABC, Dr Maryanne Demasi examines issues surrounding access to research data relating to Tamiflu.


Elsevier and US National Cancer Institute implement reciprocal linking between research articles and datasets

STM publisher Elsevier and the US National Cancer Institute (NCI) have implemented two-way linking between research articles on ScienceDirect and datasets stored in NCI's cancer Nanotechnology Laboratory (caNanoLab) data portal. The NCI is part of the US National Institutes of Health (NIH).
<snip>

Reciprocal linking between articles and data is one of Elsevier's initiatives for sharing research data. Elsevier collaborates with more than forty data repositories, and is continually looking to collaborate with other relevant organisations.
More at: http://www.prnewswire.com/news-releases/elsevier-and-us-national-cancer-institute-implement-reciprocal-linking-between-research-articles-and-datasets-512790661.html

Monday, 6 July 2015

Italy to implement ORCID on a national scale, signs three-year consortium membership agreement

The Open Research and Contributor ID (ORCID) has announced that Italy will be implementing ORCID on a national scale, and has signed a three-year consortium membership agreement with ORCID.
More at: http://orcid.org/blog/2015/06/19/italy-launches-national-orcid-implementation
 

Data Curation Bibliography v5


Digital Scholarship has released Version 5 of the Research
Data Curation Bibliography. This selective bibliography
includes over 350 English-language articles, books, and
technical reports that are useful in understanding the
curation of digital research data in academic and other
research institutions.

 

Wiley, Figshare partnership to support authors to openly share data

Publisher John Wiley & Sons, Inc. has announced a partnership with London- based data repository organisation, Figshare. To support authors who wish to openly share their data, Wiley has embarked on this partnership with Figshare to integrate data sharing within existing journal workflows and article publication.

More at: http://au.wiley.com/WileyCDA/PressRelease/pressReleaseId-119082.html

The BMJ becomes first general medical journal to require data sharing for all submitted trials

The BMJ, a weekly peer-reviewed medical journal, requires sharing of individual patient data for all clinical trials, effective July 1, 2015. This means that trials will be considered for publication only if the authors agree to make the relevant anonymised patient level data available on reasonable request.

According to Elizabeth Loder, The BMJ's acting head of research, The BMJ is the first general medical journal to require data sharing for all trials, extending its initial policy on sharing data for trials of drugs or devices, which took effect in January 2013.

In an editorial to mark the launch of the new policy, she explains that the initial policy focused on trials of drug and devices 'because many high profile, serious allegations of selective or non-reporting of trial results related to such products.' However, she says, growing experience and evidence show that reporting problems are not limited to the corporate sector, but affect academic and government sponsored trials as well.

Today's announcement follows initiatives by the US Institute of Medicine (IOM), the World Health Organization (WHO), and the Nordic Trial Alliance, to encourage data transparency.

For instance, a recent IOM report called for a transformation of existing scientific culture to one where 'data sharing is the expected norm' while WHO has said the main results of clinical trials should be made publicly available and submitted for journal publication within a year of study completion.

The efforts of industry, too, must be acknowledged, says Loder. In particular, Medtronic's cooperation with the Yale University Open Data project and GlaxoSmithKline's leadership on data disclosure efforts stand out.

Making anonymised patient level data from clinical trials available for independent scrutiny allows other researchers to replicate key analyses, reduces the possibility that studies will be unnecessarily duplicated, and maximises use of the information from trials - an important moral obligation to trial participants, she writes.

She acknowledges that an initial investment of time and money is needed to prepare trial data for sharing, 'but after the first use there are few additional costs; in essence, the value of the data increases with each use,' she concludes.
More at: http://www.bmj.com/content/350/bmj.h2373

PLOS Recommended Data Repositories


In line with their updated Data Policy, PLOS announce their PLOS Data Repository Recommendation Guide.

Everyone | PLOS One community Blog - Post by Daniella Lowenberg July 2, 2015 at:
http://blogs.plos.org/everyone/2015/07/02/plos-recommended-data-repositories/

Wednesday, 1 July 2015

A Place to Stand: e-Infrastructures and Data Management for Global Change Research

Finalized: Tuesday, June 30, 2015