Thursday, 31 July 2014



I Know Where Your Cat Lives, A Data Experiment Featuring Public Pictures of Cats on a World Map Based on Metadata

by  at LaughingSquid.com


I Know Where Your Cat Lives is a data experiment by Owen Mundy featuring public pictures of cats on a world map based on latitude and longitude coordinates obtained in their metadata. Mundy used at supercomputer atFlorida State University to run the 1 million pictures of cats gathered from publicly available APIs through a series of clustering algorithms in order to prepare them for the site. Mundy is currently raising funds to pay for hosting via a Kickstarter crowdfunding campaign.

Tuesday, 29 July 2014

Living with Data: Stories that Make Data More Personal

"The main idea is that we need more stories that ground data in personal, everyday experience. We need personal data stories make data uses intelligible and impacts personal."
...
"So while data is making our behaviors, habits, and interests more legible to firms and governments, as consumers we haven’t yet developed the critical literacies to understand what our data is saying about us and more importantly how it is shaping our experience."
...
Great blog post and videos from Sara Watson of the Berkman centre - Harvard
http://www.saramwatson.com/blog/video-living-with-data-stories-that-make-data-more

15 insights into open data supply, use and impacts

A snapshot of 15 insights or provocations for policy-makers and practitioners drawn out from the  Open Data in Developing Countries project case study reports.
These are just the first stage of the synthesis work to be carried out in the ODDC project

Monday, 28 July 2014


Value and cost benefit of open data policy in USA
Some recent US views:

 
 
 

4 limitations of Google Scholar Profiles & how to get around them


An ImpactStory blogpost of interest to those wanting to keep an eye on the impact and influence of their publications using Google Scholar. Lists the points below as the main issues, explains why they are an issue and what to do about it if you want to use Google Scholar Profiles.

1. Google Scholar Profiles include dirty data2. Google Scholar Profiles may not last3. Google Scholar won’t allow itself to be improved upon
4. Google Scholar Profiles only measure a narrow kind of scholarly impact

Friday, 25 July 2014

API Provides Open Access to FDA Recall Data

http://www.healthdatamanagement.com/news/API-Provides-Open-Access-to-FDA-Recall-Data-48453-1.html

As part of the Food and Drug Administration’s recently launched openFDA initiative, the regulatory agency is for the first time offering an application programming interface providing web developers and researchers direct access to millions of reports on drug adverse events and medication errors that have been submitted to the FDA since 2004.

Thursday, 24 July 2014

Sharing is caring, but should it count?

Interesting blog post about incentives to share data.  The author concludes:

"...there’s a big perception gap on the importance of ‘my data to my research’, and the importance of ‘my data to someone else’s research’. Closing this gap could go a long way to increasing data sharing. [Also]...the tenure and promotion system is a complicated, political mechanism and trying to leverage it as a way to incentivize data sharing is not easy or straightforward"

http://datapub.cdlib.org/2014/07/23/sharing-is-caring-but-should-it-count/

DOIs and citations for software and data

Exciting to see DOIs and citations for related software and data in this recently released paper in "Plant Methods".   Scroll down the reference list to entries 32 & 33 to see data citation in action.

http://www.plantmethods.com/content/10/1/23

Interesting that the journal has chosen not to conform to the DataCite preferred citation format, but the presence of a  DOI will help with tracking reuse and citation metrics.

Monday, 21 July 2014

How Data Will Transform Science

http://www.forbes.com/sites/gregsatell/2014/07/18/how-data-will-transform-science/

Tuesday, 15 July 2014

Historical closed data leads to poor science

Hi Everyone,

This is a story I thought that Data Scientists would appreciate, despite being predominately skewed to the biological sciences.

Quick background: Males of Drosophila melanogaster (Fruit Fly) 'sing' to their mates by vibrating their wings.

In 1980, researchers identified rhythms in the 'wing songs' that could be related back to particular alleles.

    Kyriacou, C.P. & Hall, J.C., 1980. Circadian rhythm mutations in
    Drosophila melanogaster affect short-term fluctuations in the
    male’s courtship song. Proceedings of the National Academy of
    Sciences, 77(11), pp.6729–6733.

The paper and its research findings enter scientific legend and are reiterated and cited over and over, despite only a few being able to reproduce their results.

Recent work has verified that the rhythms reported were actually artifacts of the way the data was analyzed.

    Stern, D.L., 2014. Reported Drosophila courtship song rhythms are
    artifacts of data analysis. BMC Biology, 12(1), p.38.

A cautionary tail or a case for the release of raw data so everyone can draw their own conclusions?

NOTE: This was an email posted to the school-of-data mailing list

school-of-data@lists.okfn.org
https://lists.okfn.org/mailman/listinfo/school-of-data

Monday, 14 July 2014

Maximising the value of research data: developing incentives and changing cultures

The value of sharing research data is widely recognised by the research community and funders are setting in place stronger policy requirements for researchers to share data. But the costs to researchers in sharing their data can be considerable and the incentives are sometimes few and far between. A recent report from the cross-disciplinary Expert Advisory Group on Data Access (EAGDA) highlights the need for a shift in cultures to provide greater support for researchers in sharing data and greater recognition for those who do it well. Dave Carr and Natalie Banner, from the Wellcome Trust, highlight some of the key findings and recommendations emerging from this work.  

http://blogs.lse.ac.uk/impactofsocialsciences/2014/07/01/maximising-value-research-data-wellcome-trust/

Thursday, 10 July 2014

NASA opens earth science data, cloud computing to the public as part of new contest

SUMMARY:
NASA is launching a new challenge, hosted on Amazon Web Services, that gives the public access to a trove of earth sciences data and computational resources in the name of discovering new uses for all that information.
nasa
NASA launched a new contest Tuesday for imagining and then building new uses for the space agency’s trove of earth sciences data. The challenge — actually two of them, broken down into the imagination and building stages — kicks off on July 1 and runs through Nov. 15, and utilizes the NASA Open Earth Exchange platform. The exchange’s datasets and informational material, as well as the computing resources for the challenge, are hosted on Amazon Web Services.

http://gigaom.com/2014/06/24/nasa-opens-earth-science-data-cloud-computing-to-the-public-as-part-of-new-contest/




Wednesday, 2 July 2014

Article-Level metrics


Executive Summary
Article-Level Metrics (ALMs) are being discussed, shared, and used. ALMs can be employed in conjunction with existing metrics, which have traditionally focused on the long-term  impact  of  a  collection  of articles (i.e., a journal) based on the number of citations generated. This primer is designed to give campus leaders and other interested parties an overview of what ALMs are, why they matter, how they complement established utilities, and how they can be used in the tenure and promotion process. Among this resource’s key takeaways are the following:

http://www.sparc.arl.org/sites/default/files/sparc-alm-primer.pdf


Altmetrics: Rethinking the Way We Measure

Article from

Serials Review

Volume 39, Issue 1, 2013


Altmetrics: Rethinking the Way We Measure

Altmetrics: Rethinking the Way We Measure


Abstract

Altmetrics is the focus for this edition of “Balance Point.” The column editor invited Finbar Galligan who has gained considerable knowledge of altmetrics to co-author the column. Altmetrics, their relationship to traditional metrics, their importance, uses, potential impacts, and possible future directions are examined. The authors conclude that altmetrics have an important future role to play and that they offer the potential to revolutionize the analysis of the value and impact of scholarly work.

http://www.tandfonline.com/doi/full/10.1080/00987913.2013.10765486#.U6ywDfmSytY

Journal data sharing – model framework

A model framework from the Journal Research Data (JoRD) Project makes 21 recommendations for journal publishers to help increase the availability of supporting data.
In the first instance the framework would require journals to both establish and publish data sharing policies. These policies should clearly state what data is to be deposited and where, and whether it should be linked to the journal article. It acknowledges that guidance must be given for the selection of data from large data sets (where the entire set is not relevant to an article). It would require a statement of the reviewing status of data (i.e. that the data will or will not be reviewed) and whether an embargo on making the data available is acceptable.
It also provides guidance on data citation style (see the Data citation project above).
The framework has emerged from the (now completed) Journal Research Data (JoRD) Project which was undertaken at the Centre for Research Communications(CRC) at the University of Nottingham, funded by JISC. The project is now closed, but its outcomes and this model framework is reviewed in Research Data Sharing: Developing a Stakeholder-Driven Model for Journal Policies (preprint available).

Tuesday, 1 July 2014

New guide released : Depositing shareable survey data

Those who own and manage large-scale surveys now have a compact and detailed guide to help them make their data more widely used by researchers.
The guide 'Depositing shareable survey data' was developed by a specialist team at the UK Data Service with extensive input from UK government departments, academic survey owners and survey producers.
The 16-page handbook takes the reader through the full data journey, from fieldwork planning to eventual user access. Correspondingly, all content is organised into five stages of the journey: Plan, Prepare, Negotiate, Deposit and Ingest. While the guide was specifically developed to support new depositors of large-scale surveys, the principles apply to a wide range of significant deposits.
The guide can feature as a useful check for owners and producers; it also suggests wording for commissioning tenders and briefs on considering any data sharing requirements. The aim is to encourage commissioning departments to include a standard paragraph requiring archiving with the UK Data Service in a timely manner and pointing to existing protocols for doing this.
"This handy guide encompasses all relevant guidance we provide on depositing survey data – written and verbal - in a nutshell", says Louise Corti, who heads the UK Data Service Collections Development and Producer Support team. "I expect the guide will find a place pinned to the notice board of every survey commissioner and manager."
The guide is available free from the UK Data Service website at the link below.  Hard copies also available!