Thursday, 23 July 2015

"Tools and Services over Data" Workshop

On 24 June 2015, the Australian National Data Service (ANDS) facilitated a workshop “Tools and Services over Data” at the ANU University House to initiate a discussion between data providers and data service providers about how to more efficiently and effectively connect data to tools and services over data, and to identify common activities that could be coordinated across NCRIS capabilities. The motivation to hold this workshop was twofold: 1) NCRIS-funded capabilities are currently planning the next 1-2 years' activities, so this is a good time to look for synergies 2) At the same time, all of the NCRIS e-Research capabilities, and a number of the NCRIS data-intensive capabilities are working on activities with an explicit data focus. Thus, ANDS organised and facilitated a one day workshop  to discuss these issues, bringing together twenty-two people from NCRIS capabilities (ANDS, National eResearch Collaboration Tools and Resources (NeCTAR), Terrestrial Ecosystem Research Network (TERN), National Computational Infrastructure (NCI), Integrated Marine Observing System (IMOS), Bioplatforms Australia (BPA), Atlas of Living Australia (ALA), NeCTAR Virtual Labs (Biodiversity and Climate Change Virtual Laboratory (BCCVL) , Virtual Geophysics Laboratory (VGL), Characterisation Virtual Laboratory (CVL), and Genomics Virtual Laboratory (GVL)), Geoscience Australia (GA), Australian eResearch Organisations (AeRO), and Queensland Cyber Infrastructure Foundation (QCIF). 
The workshop had three major components: presentations, group breakout sessions, and reporting back / identifying next steps.  In the presentation session, participants each gave 3 minutes’ presentations on: 1) the main services/tools over data that they have provided to their users (or funded in the case of NeCTAR), 2) success stories of other groups using their services/tools, and 3) examples of more coordinated approaches to tools/services over data that they are aware of.  Six of the common topics emerging from these presentations were then selected for the group breakout sessions. Two rounds of breakout sessions were held with three topics discussed in each session. During the group breakout sessions, each group first discussed possible pilot activities and what participants (and others) could do collectively. A representative of each group reported back to the large group.  The workshop ended with a reporting back / next steps session to identify future actions and timelines.

The six topics selected for group discussion, together with highlights extracted from the group discussion and reporting back sessions, were:   
  1. Add-on services - semantic web, provenance, vocabularies, linked data & processing
    • The Semantic Web tries to achieve true discovery, there are already are a large number of value-adding tools such as provenance, proper use of URLs, and vocabulary services.
    • There are new services and tools available in the semantic space, although at different maturing stages.  Provenance may be the next to pickup and work with. ANDS, NeCTAR and other NCRIS capabilities have funded provenance-related elements in different projects, but we need a centralised discussion of the ongoing services in the provenance space.
    • We are getting more and more dependent on vocabulary service and provenance, we need to provide an ongoing network of vocabulary services in conjunction with data.gov.au and other related agencies.
  2. Data formats, services and discovery
    • Data should be stored in open and flexible formats where good conversion tools exist that retain all metadata, proprietary formats are OK where they are well documented and an API or conversion tool exists if this is the only way to maintain all metadata.
    • It is too expensive to maintain data  in multiple formats and at multiple sites. Costs include not only storage but also support etc.
    • There should be a continuous discussion about Data as a Service, and use of common services to access common data formats (think of data delivery formats, not data storage formats).
  3. Sampling subsets of data
    • Caching subset of data leads to needing alerts when the source data changes. Provenance information can help to determine when an alert is needed.
    • A cache of data close to the compute resource can physically integrate data from multiple sources, but this leads to a challenge of how to propagate aggregated provenance information with aggregated data.
    • There is a need to aggregate data across different sensors measuring the same observation in different ways.  Providing aggregated data is a challenge due to differences in calibration. How to trust aggregated data and provide all the provenance information is also non-trivial.
  4. Data fusion challenges
    • Data to be integrated may be from multiple platforms, disciplines, model-observations, license regimes,  citations, and different measurement protocols/standards/resolutions etc.
    • One challenge is how to efficiently describe the data that allows interpretation by others (people and machines)
    • Licensing can become a show-stopper when fusing data under different licensing regimes. There is a need to sort out a national licensing framework for data to make this less likely.
  5. Access to data and storage issues - should we make data locally or remotely accessible?
    • Some data are hosted overseas, for example, the Cancer Genome Atlas. Setting up a proxy on request to mirror the data may not solve the problem unless there is enough computing capacity close to the data for researchers’ customised analysis pipelines.
    • ALA caches spatial data for local indexing. It would be good to cache spatial data at NCI so that everyone could use it and ALA could run their indexer there. What would be best is a centralised authoritative storage system that manages data files and index layers.
    • BCCVL takes a copy of climate data irregularly, and also a direct copy of ALA to allow more complex queries; but don’t want to maintain the mirror of ALA.
  6. Where is the User Interface - should we provide users tools and services over their own desktop, cloud, or virtual desktop?
    • There is no single best solution; different users have different needs. For example, Biologists use Galaxy, Bioinformaticians want to use Desktop in the cloud. Having the data pre-attached makes a huge difference.
    • We may need different efficiency  criteria/measures: NCI uses use of cores - not working for everything. People will use cloud inefficiently, which is part of the cost. Running local IT is also often inefficient.

Finally, the following possible activities were identified in the reporting back / next steps session (organisation responsible for next steps in parentheses):
  • Vocabulary: it is good to have a national project to bring national agencies together. To complement that, this is to be done in conjunction with data.gov.au, the Australian Linked data group, and other research organisations. We need to think about what a national vocabulary is service. (ANDS)
  • Provenance: get people from provenance projects together to think about where we are going next. (ANDS)
  • Software registry: how to describe software and setup a software registry. (RDS may run a workshop on this topic)
  • Look at what data formats and services from community. What are data services for common format? How to discover service? (RDS, VLs)
  • A framework for data licenses applicable when fusing data together (ANDS + AusGoal)
  • Education program for data providers. (ANDS+NeCTAR).  
  • Define minimum description for M2M negotiation including license requirement. (Hamish Holewa from BCCVL will put up a case study.)
  • Improve core allocation process. (NeCTAR)

In conclusion, this was a great information gathering session. Attendees indicated interest of getting together at this year’s eResearch Australasia conference. ANDS is planning to organise a workshop on provenance in late August to bring people from provenance projects together to discuss about where we are going next.

If you would like to contribute to the discussion of one or more of the above topics, please contact Stefanie Kethers (stefanie.kethers@ands.org.au) or Hamish Holewa (hholewa@gmail.com).

No comments:

Post a Comment