A Blog post by Libby Liggins(2019 WDS Data Stewardship Award Winner)
For over four decades, scientists have been collecting genetic DNA sequence data for thousands of the world’s species. In the biodiversity and eco-evolutionary sciences, these data are generated to describe new species, define their evolutionary relationships, determine the levels of dispersal among populations, and assess levels of genetic diversity across a species range. The rate at which we accrue these DNA sequences has increased over time as the use of genetic data has diversified, and the sequencing technologies used to decode the DNA sequences of organisms have become faster, cheaper, and much higher through-put. As this trend continues into the future, it is anticipated that we may soon have more DNA sequences in a digital form than we have existing in the natural world.
This massive and growing data resource could now be consolidated for multiple species and populations and reused to better understand the world’s biodiversity at the genetic level. Genes are recognized as a fundamental component of the biodiversity hierarchy, but have received less attention than species- and ecosystem-level measures of biodiversity. In part, this may be due to synthetic analyses of genetic data being challenging and sometimes impossible, as there has been no concerted effort towards the curation and stewardship of this valuable data resource. While funding agencies and publishers advocate deposition of DNA sequence data in open-access repositories (such as the National Center for Biotechnology Information; and the European Bioinformatics Institute), they do not require the deposition of standardized metadata such as the sampling location, date, and habitat of the sampling event (Pope et al. 2015). This ‘metadata gap’ means that information essential for multispecies analyses to better understand biodiversity and evolutionary patterns across our globe, has not been readily available.
The Genomic Observatories MetaDatabase (GEOME; Deck et al. 2017) has recently provided a solution to this metadata gap. GEOME links ecologically and evolutionarily relevant metadata with DNA sequences uploaded to open-access repositories. The metadatabase incorporates the latest international standards for biodiversity and genomic data, and helps researchers store and access genetic data relevant to studies concerning large scale biodiversity and conservation problems. In conjunction with the open-access DNA sequence repositories, GEOME ensures that researchers and projects generating genetic data can adhere to the FAIR Principles (Findable, Accessible, Interoperable, Reusable; Wilkinson et al. 2016), promoting research community best-practice.
The Ira Moana Project logo. The Māori phrase Ira Moana could be interpreted as meaning ‘ocean genes’ or ‘dot in the ocean’. Both seem appropriate when thinking about the scale of DNA in the vastness of the ocean. The use of te reo Māori (Māori language) resonates with the project objectives that are uniquely New Zealand, as is the Māori language. Yet, moana is used to describe the ocean by many Pacific nations, reminding us of the connections that New Zealand’s biodiversity has with the wider Pacific region.
The Ira Moana Project has partnered with GEOME both to enable a collaborative network of researchers to adhere to these standards in community best-practice, and deliver a searchable metadatabase for the genetic data of Aotearoa New Zealand’s marine organisms. The Project aims to build and maintain the most comprehensive national database of marine genetic data in the world, ensuring kaitiakitanga (guardianship and stewardship) and creating opportunities for data synthesis to inform New Zealand’s future research directions and conservation decisions. The Ira Moana Project builds on the success of the Diversity of the Indo-Pacific Network (DIPnet) that through the use of GEOME and multi-national collaboration, has created the largest population genetic database in the world. DIPnet consolidated over 200 genetic datasets for Indo-Pacific marine organisms, and is now delivering novel biodiversity insights for the Indo-Pacific Ocean (e.g., Crandall et al. 2018), which is the largest and one of the most threatened biogeographic regions on our globe.
The Ira Moana Project is similarly founded in concern for the marine environment. New Zealand is a marine nation—we have one of the largest exclusive maritime economic zones in the world, which sustains our marine and tourism industries, and provides significant recreational and social benefits for New Zealanders. Nationally, and as global citizens, we are under pressure to make informed decisions regarding commercial and recreational activities, and how they can be balanced with the protection of our marine ecosystems. Such decisions of environmental, economic, and societal impact need to be transparent and based on robust information, as well as including knowledge about biodiversity that stretches from ecosystems to genes. The Ira Moana Project has established that there are over 430 genetic datasets for New Zealand marine organisms, and is now working to consolidate these data for the benefit of future researchers and generations of New Zealanders.
The data lifecycle in genetic research. DNA sequence data is routinely deposited into open-access genetic data repositories (under OUTPUTS). Despite metadata being accrued at every step of research (*), starting with COLLECTION, the practice of depositing metadata into repositories such as the Genomics Observatory Metadatabase (GEOME) is very recent. The Ira Moana Project is one of the project’s using the infrastructure provided by GEOME. Stewardship of metadata alongside DNA sequence data ensures that genetic research in the biodiversity, ecological, and evolutionary sciences can be reproducible, the genetic data can be re-used, and that the provenance of the genetic data and the rights of the local communities involved in the research are maintained.
As the first national project to make use of the GEOME infrastructure, the Ira Moana Project has worked with GEOME to extend the capability of the metadatabase to additionally acknowledge indigenous rights. It has become apparent that what is considered fair and equitable research practice within the research community, may not be fair and equitable within broader society. Through collaboration with Local Contexts and Te Mana Rauranga (the Māori Data Sovereignty Network), the Ira Moana Project and GEOME are now beta-testing the capacity for researchers to add Notices (such as the Traditional Knowledge Notice; TK Notice) and new Biocultural Labels as metadata for DNA sequence data. Notices signal that there are accompanying Indigenous rights needing further attention for any responsible and equitable future use of the data. Biocultural Labels further allow the addition of provenance information and community expectations for future use based on Indigenous Data Sovereignty principles—including the CARE Principles (Collective Benefit, Authority to Control, Responsibility, Ethics) launched by the Global Indigenous Data Alliance—thereby enabling Indigenous stewardship and persistent recognition of Indigenous rights within an international framework (complying with the Nagoya Protocol to the Convention on Biological Diversity). The implementation of Notices and Biocultural Labels using GEOME infrastructure is a first for a biological resource and for genetic data, establishing new ethical standards in this research community.
Workshops and datathons for New Zealand researchers have encouraged uptake and use of the metadata infrastructure provided through the Ira Moana Project and GEOME. There are now greater than 85 researchers who have joined the Ira Moana Project Network; being part of the network means being ‘on-board’ both with the things that the Ira Moana Project is trying to achieve for New Zealand, and the metadata standards that GEOME is accommodating for researchers worldwide. As there is a global community of researchers who generate genetic data, it will be some time before there is universal uptake of these newly recognized standards of best-practice. Nonetheless, we should be encouraged by the fact that as a community, we have made similar transformations in our practice in the past; since the introduction of the Joint Data Archiving Policy, it has been considered standard practice to deposit genetic data into open-access repositories. As such, we anticipate that the Ira Moana Project metadatabase will continue to grow and serve New Zealander’s, and there will be increasing uptake of the services that GEOME provides to the research and wider community.
Literature cited – Crandall ED, Riginos C, Bird CE, Liggins L, Treml E, Beger M, Barber PH, Connolly SR, Cowman PF, DiBattista JD, et al. 2019. The molecular biogeography of the Indo-Pacific: Testing hypotheses with multispecies genetic patterns. Global Ecology and Biogeography. 58(5):403–418. – Deck J, Gaither MR, Ewing R, Bird CE, Davies N, Meyer C, Riginos C, Toonen RJ, Crandall ED. 2017. The Genomic Observatories Metadatabase (GEOME): A new repository for field and sampling event metadata associated with genetic samples. PLoS Biology. 15(8):e2002925. – Pope LC, Liggins L, Keyse J, Carvalho SB, Riginos C. 2015. Not the time or the place: the missing spatio‐temporal link in publicly available genetic data. Molecular Ecology. 24(15):3802-9. – Wilkinson MD, Dumontier M, Aalbersberg IJ, Appleton G, Axton M, Baak A, Blomberg N, Boiten JW, da Silva Santos LB, Bourne PE, Bouwman J. 2016. The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data. 3
Citizen science is getting more and more attention worldwide; in particular, there is a growing interest in involving citizens in data collection due to its capability to complement the acquisition of data classically accomplished through existing complex instrumentation networks. Scientists have experimented with multiple forms of citizen science projects, which have been successfully implemented in many fields. The value of using citizen contributions has been proven—or at least explored—in almost all scientific domains, and its potential is currently also being investigated in the processes of decision- and policy-making.
There are many definitions of citizen science. The definition most often used is that of Buytaert et al. (2014): The participation of the general public (i.e. non-scientists) in the generation of new knowledge. In this blog post, I focus on citizen science from the perspective of data collected by citizens and the use of these data, but there is also much research looking into how to involve citizens, and consequently, how they are participating in the collection of data. Taking the latter viewpoint, there is now lots of terminology that can be found in the literature; for example, citizen observatory (CO), citizen sensing, trained volunteers, crowdsourcing, community-based monitoring, volunteered geographic information, eyewitnesses, and so on.
As mentioned in the title, I would like to spend the remainder of this blog post briefly introducing four Horizon 2020 funded projects that have used innovative technologies for collecting data with the help of citizen scientists. The projects ran from the second half of 2016 until mid-2019, and were clustered under WeObserve, which examines the challenges faced by COs in terms of awareness, acceptability, and sustainability. They shared the specific goal that their final (analyzed and processed) data products would not only complement existing data elements within the Global Earth Observation System of Systems (GEOSS), but also become new GEOSS contributions.
SCENT (Smart Toolbox for Engaging Citizens into a People-Centric Observation Web)
Citizens were engaged in environmental monitoring of land-cover/use changes using their smartphones and tablets, enabling them to become the ‘eyes’ of the policymakers. In particular, the project looked at two pilots—the urban case of the Kifisos river in Attica, Greece and the rural case of the Danube Delta in Romania—where the citizen-collected data were used to assess flood models and flooding patterns. You can read more about this project here.
LANDSENSE (Connecting citizens with satellite imagery to transform environmental decision making)
The focus of this project was on the potential of Earth observations taken by citizen scientists to augment and improve the way we see, map, and understand the world. Three main areas of application were selected as demonstrators: urban landscape dynamics, agricultural land use, and forest and habitat modelling. Read more about LANDSENSE here.
Groundtruth2.0 (How to impact decision making with citizen observatories)
The interaction was investigated between people and technology when it comes to setting up a successful system for land and natural resources management. The project combined the social dimensions of COs and enabling technologies so that the implementation of each observatory was tailored to its envisaged societal and economic impacts with a specific emphasis on flora and fauna, as well as water availability and quality. Find out more about the project here.
GROW (Grow Observatory)
In this project, citizen scientists collected information on land, soil, and water resources to answer a long-standing challenge for space science; namely, the validation of soil moisture detection from satellites. Read more here.
Reference Buytaert, W., et al: Citizen science in hydrology and water resources: opportunities for knowledge generation, ecosystem service management, and sustainable development, Front. Earth Sci., 2, 26, doi: 10.3389/feart.2014.00026, 2014.
A ‘challenge call’ is made public a few months before the conference alongside a deadline for the result papers, which are then evaluated by a jury. Introduced by Prof Paulo Carvalho from Coimbra University in Portugal and Prof Ratko Magjarevic from Zagreb University in Croatia, the challenges have proved quite successful, with the participation of 20–30 groups of young researchers responding to the first call.
A major problem for the organizers, however, has been to find adequate datasets containing well-documented biomedical data, such as respiratory measurements, electroencephalography recordings, electro-cardiac recordings, and the like. While many state that Big Data is widely accessible and available, well-documented and consistent biomedical datasets are difficult to find. This has resulted in the IFMBE having to actually sponsor teams to collect appropriate datasets of biomedical measurements specifically for the ‘scientific challenge’ competitions!
To address such issues, programmes are now being started that encourage universities and research groups in the Biomedical Sciences to make their datasets public while taking adequate precautions to protect the privacy of patients when such datasets are linked to physical persons. IFMBE is currently exploring practical ways to constitute collections of well-documented biomedical datasets that comply with the FAIR principles and that are made publicly available to researchers via a repository. Moreover, it is encouraging member societies at large to take up similar schemes either themselves or via universities.
Universities are inherently multidisciplinary and often hold a wide variety of research datasets. This makes them an ideal place to develop and test systems to manage, host, and access multidisciplinary and heterogeneous research datasets. However, the existence of such datasets and how they are preserved is not always well known. At Kyoto University, a survey was conducted by the Academic Data Innovation Unit* to gain a basic understanding of this information towards the planning of a new research data management system. The survey was sent to all researchers at Kyoto University, more than 3,000 of them, in December 2018 and we collected their responses until the end of January 2019. Although the survey was not mandatory, valid responses were received from 244 researchers ranging across the disciplines in Figure 1. From the results, we see that the largest proportion of datasets are held by the Life Sciences. This may not be the reality, however, since we received an unexpectedly low response from the Technology departments, which form the largest group at Kyoto University.
Figure 1: Responses by discipline
Figure 2 indicates the level of openness for each of the datasets identified by researchers. As can be seen, the majority of datasets are shared within a research group only and are not open to others (or even open at all). The implication is that the principle use case we need to account for on campus when developing a data management system is the sharing of data among members within each research group rather than making the data completely open.
Figure 2: Number of open and closed datasets
Despite the above, we believe that it should be possible for some researchers to make their datasets open to all if they are provided with appropriate technical support. Proper education and training on open data and data management will also assist in this process. In particular, around 20 data repositories—mostly hosted by research institutes within Kyoto University—are of especially high quality, and we would expect that about half of them could potentially become CoreTrustSeal-certified WDS Regular Members.
*The Academic Data Innovation Unit is a virtual organization at Kyoto University and is currently chaired by Prof Shoji Kajita. One of its main tasks is to propose a research data management system to accommodate the needs of all researchers at Kyoto University.
The WDS-ECR Network was set up in September 2017 to promote scientific data stewardship, share best practices, and foster better communication among ECRs. As a co-lead of the Network, alongside Sabrina Delgado Arias (Science Systems and Applications, Inc.) and Ivan Pyshnograiev (Igor Sikorsky Kyiv Polytechnic Institute), I coordinate events, speaker series, and periodic teleconferences, as well as liaise with other ECR communities to share ideas on future data practices.
In 2018, the WDS-SC invited a representative of the ECR Network to take part in their meetings by opening up a one-year rolling seat on the Committee. This initiative facilitates communication with the next generations of data managers and enables WDS to develop activities targeting ECR’s interests. I represented the Network on the WDS-SC from July 2018 to June 2019. Being a member of the WDS-SC was an amazing experience and opportunity.
All SC members are working pro bono to share their ideas on how to best shape the future of data stewardship for better science. This is very exciting! SC members meet each month via teleconference, and then twice a year in person where most of the plans for actions are validated. In order to reach out to different communities, face-to-face meetings of the WDS-SC are often co-located with other WDS events such as regional conferences. I attended two such meetings, one in November 2018 in Cape Town, and another in May 2019 in Beijing. During the very intense two-day meetings, SC members present their ideas and discuss the tasks for WDS to undertake in the following months. I got to meet exceptional data experts from around the world, and took part in the decisions, and strategic actions and activities of WDS.
In particular, I participated in the preparation of a training workshop targeting ECRs that is sponsored by a grant of the European Geosciences Union. I saw how much work is involved in setting up such events, and I am sure it will be very rewarding for PhD students and Post-docs to learn more about Research Data Management. The training workshop is a great opportunity for those attending and crucial for the future of science. Being part of the WDS-SC provided me with the chance to share my inputs when necessary. I really appreciated seeing that my suggestions were valued. I thank all the SC members for their warm welcome and for the trust I was given. I also encourage all early career scientists and researchers who work with data to join the WDS-ECR Network. It might be you representing the ECR Network on the WDS-SC in the future!
Talking of which...Sabrina Delgado Arias will represent the WDS ECR Network on the WDS-SC from July 2019 to June 2020. We wish her all the best!
In this piece, the authors begin to describe intersections of information maintenance and care ethics in ways that are real and meaningful for information maintainers (i.e., those who manage, maintain, and preserve information systems).
Contributors to this document have varied experiences with information maintenance: community organizers and facilitators, archivists, repository managers, project managers, designers, librarians, researchers, grantmakers, educators, and more. The authors invite those in occupations and roles who understand that the relationship is especially valuable between information maintenance and an ethic of care to read, react, share, and engage with this potluck of ideas. Please circulate widely!