Duplication of Occurrences in Biodiversity Databases

Duplication of occurrences in Biodiversity Databases can occur when data of the same occurrence is submitted via two or more channels of input. An example of this is can be when an observation is both posted on iNaturalist and also it is submitted within a Bee Atlas initiative.

Below is a screenshot from an iNaturalist observation of a bee on the Web Browser Platform.  Data was submitted with images of the pinned bee which will be submitted later to a Bee Atlas initiative. This observation has reached Research Grade (A). Research Grade is a level of iNaturalist community agreement that occurs, usually with species level identification. In this observation, 2 users have agreed (B) with the initial suggested ID (~determination). The community consensus is when 2/3 or more of the identifiers equivalent or agree on the suggested ID (C) - two out of two, two to three out of three, four to six out of six, and so on. Early on in the development of iNaturalist, Research Grade was a term given to observations that reached a status of identification whereby if permissions by the observation data owner was given, the data would be passed on to biodiversity databases such as the Global Biodiversity Information Facility (GBIF). In this particular observation the copyright permissions are set to share with GBIF as indicated by the CC-BY-NC choice of permissions (D) giving only some rights reserved. iNaturalist users are given the opportunity to have a default setting for permissions depending on their individual desires. This observation has icons indicating that observation information has been shared with biodiversity databases (E). The iNaturalist Help pages states "GBIF only accepts datasets that are licensed for reuse, so they only accept content with a CC BY or CC BY-NC license, or with the CC0 declaration". iNaturalist believes that GBIF updates their dataset from iNaturalist about once a week.

 

The question then is how does one change this duplication from occurring knowing that GBIF will only accept certain permissions. On the observation in question, if one selects the drop down menu designated by the caron symbol, ▼(F), a menu will popup that gives a choice to Edit License (G). Once selected, another internal window will popup with with License choices.


This window below shows the license choices with explanations. Your default license choice will already have a blue dot (H). Deselect this (X) and choose No License (all rights reserved)(I). With this choice, no data information will be shared with GBIF. Don't forget to Set License (J). This will just be for this observation unless you wish to choose other settings as a default or going forward.

By completing the above step. the occurrence data is no longer shared with GBIF.

The screenshot from GBIF below shows the occurrence data as it was shared with GBIF.

 
Below shows an instance of an iNaturalist observation having all rights reserved license placed after already having the observation data shared with GBIF. One can see the GBIF shared icon below in the right in this observation. One can also notice it states above this that it is "not licensed for re-use and will not be shared with data repositories that respect license choices".

 If one selects the GBIF icon, one will find that there is a record of the occurrence stating "The data publisher has removed this record from the GBIF index, but the last version is shown below. Records no longer in a published dataset are removed from the GBIF index. Publishers may sometimes have reasons to remove individual records or an entire dataset or assign new local identifiers". A search of GBIF records through the normal channels will not show this occurrence.

Below is a screenshot from an iNaturalist observation of a pinned bee that will be shared with a Bee Atlas initiative. In this particular instance the observation data copyright was with all rights reserved as indicated by the first arrow - this occurred prior to becoming Research Grade. If one looks below that the second arrow highlights that this observation is not being shared with GBIF and states "not licensed for re-use and will not be shared with data repositories that respect license choices".

As stated, submitting bee specimens to a Bee Atlas or Museum Collections and also adding observation images of the same specimen to one's own life list on platforms such as iNaturalist can result in duplicate occurrences in biodiversity databases. Below are some of the possible issues with duplication of occurrences.

Data duplication can lead to inflated records, redundant data, and increased processing load, potentially skewing biodiversity studies and ecological models.

Data integrity can be questioned by conflicting or inconsistent information from different sources creating uncertainty and record accuracy determination challenges. For example inconsistencies can occur in location, dates, observer name (Bob instead of Robert or RW or some other User Name), or specimen determination by a community of citizen scientists versus a taxonomist.

Complications in data use with duplicates can mislead analysis, complicate data merging, and result in biased findings or misguided policy decisions.

Reputation and credibility with duplication can be impacted by frequent duplicates which can undermine user confidence in data which can lead to decreased usage and trust in platforms.

A question is then why would an individual choose to post an observation of a pinned bee on a site such as iNaturalist. Much the same as birders having life lists of the species they have seen, others are engaged in keeping life lists of other fauna and/or flora. Some are generalists and some will pick a phylum, class, order, family, or genus to focus on. iNaturalist is a great resource for these individuals who can grow their lists and learn to identify new species. The feedback can often be quite immediate. As the image data grows in volume, the ability to draw on this to refine one's identification skills improves. By contributing to the images with more refined images of poorly documented species, one is both increasing their own life lists but also improving identification resources for others in the community. 

As immediate and accessible as these shared resources on iNaturalist are, one needs to consider the greater impact of how the information is shared. Using all rights reserved copyright to restrict the sharing of data with other biodiversity platforms may help to mitigate the duplication of occurrences.

One important thing to note: Since uploading this blog, ust recently for some reason I have not figured out, on occasion the observation will revert to  CC BY-NC. I need to contact the iNat help desk and see if this just happens when someone agrees to my identification thereby moving it to research grade. Because of this it becomes an active process.

Comments

  1. Hey Bob,

    This is an interesting topic and a great review of how to keep your records from going to GBIF if they're destined for another repository and you don't want to have duplicated data. I have a couple of thoughts on the issue of data duplication in species occurrence data... I'll start by saying that I long wished that the there was a share function for iNat like there is for eBird. I thought it was lame to have to duplicate a record of something just so that both my friend and I could add it to our lists. But I don't think that anymore.

    I'm inclined to challenge the idea that these are actually issues of duplications in species occurrence datasets. I'll go through your list of issues point by point so that I don't ramble too much:

    i) Redundant data and increased processing speed: I don't think that this is really an issue of data duplication, I think that its an issue (well more of a function) of the dataset as a whole. iNat is a huge dataset and the data that you pull from it will be unwieldy for analysis. I don't think (though I admit that I don't have the data to back this up) that duplications will have any tangible impact on this considering the amount of data that's included. You could fix the problem of too much "useless data" in more efficient ways than painstakingly removing duplicated records... although you can remove duplicates by date and location with 1 line of code in R.

    ii) Data integrity can be questioned: Data integrity SHOULD be questioned if there's conflicting or inconsistent information. Certainly for inconsistent location, date, or specimen determination. If those do not align then that's a big problem that should be fixed/investigated. For observer name inconsistencies, I guess that's a problem but I think that can be fixed by adding an ORCiD to your profile or just by using a consistent name wherever you contribute.

    iii) Duplicates can mislead analysis: I don't think that this is an issue with duplication, I think it's an issue with using inappropriate analyses. Species occurrence datasets like iNat, GBIF, and (frankly) museum collections should really only be analyzed as presence-only datasets (ie a collection of 1s or NAs). When you work with presence-only data it doesn't matter if there are 50 records in a place or just 1... if it's not an NA it can only be a 1. And it's NEVER a 0. These datasets can't show species absence, nor can they show abundance. If you need to show either then you shouldn't be using iNat data. If duplicates are influencing a model then it's an inappropriate model for the data. Even if though there's a share function on eBird there are still >100 checklists of Temminck's Stint at Panama Flats (and no one thinks that it's more than one bird).

    [1/2]

    ReplyDelete
  2. Hi Bob,

    I wrote a long winded comment but it all deleted when I hit "publish" so I'll be more concise this time (please don't take my brevity as rudeness).

    I disagree that these issues are actually issues caused by duplications between the datasets. I think it's more an issue of the final user misusing the data.

    iNat data is presence-only data (1s and NAs), nothing more. No mater if there are 100 records or 1 record in an area of interest it must only be considered as a 1 in your analyses. If duplicates in a presence-only dataset are messing with your statistical analyses then you're probably using the wrong model.

    I doubt that the best way to reduce the computational load of a dataset like iNat/GBIF is by painstakingly removing duplicates. There's a lot of mess on both that needs cleaning. But if it is a big deal, I can share a single line of R code that removes duplicates based on date and location from your downloaded dataframe.

    The bad reputation and credibility of iNat is an issue of people misunderstanding what iNat is and what it can do, not that there are duplicates (it's presence-only data after all). When taxonomic experts get upset about the problems of iNaturalist it can normally be fixed by reminding them what iNat is and -- more importantly -- what it isn't.

    "Data integrity [SHOULD] be questioned by conflicting or inconsistent information from different sources". If location, date, or species determination data don't align, that's a problem that needs to be resolved. The Robert/Bob issue doesn't seem like as big of a problem but I think it can be resolved by either using the same name across platforms or by adding your ORCiD to your profile.

    Finally, why not just add a field to the Bee Atlas data upload portal (or a Bee Atlas ID as an Observation Field in iNat) so that you can filter the duplications after you've pulled the data? That seems like a better fix than limiting the use of high quality records that might be duplicates to your colleagues, but certainly aren't to anyone who isn't using Bee Atlas data. You're basically removing the best bee records on GBIF because of an issue that I don't really agree exists. That's a bummer...

    I'm sure that I've made mistakes above (especially in my quick second try) so please critique my points. I really want to know how I can be a better iNat user and I certainly can learn a lot from your expertise.

    Cheers,
    Nathan

    ReplyDelete

Post a Comment

Popular posts from this blog

Catch a Buzz Links

Filling Field