--- title: "Requesting GBIF data for a species list" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Requesting GBIF data for a species list} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r setup, include = FALSE} knitr::opts_chunk$set( collapse = TRUE, comment = "#>", eval = FALSE ) ``` GBIF (the Global Biodiversity Information Facility) assigns each taxon in its backbone an integer identifier called a *taxon key*. taxify uses these keys to request occurrence records for a species list. `taxify()` matches names against a local copy of the GBIF backbone. `gbif_request()` submits the keys to GBIF, and `gbif_fetch()` waits for the download and imports the records with the input names attached. GBIF caps a download query at about 1,000 keys, so `gbif_request()` splits a longer list into several downloads, each with its own DOI ([details](#long-lists-are-split-across-downloads)). Name matching runs offline. Requests and downloads require [rgbif](https://CRAN.R-project.org/package=rgbif). The default download method also requires a free GBIF account; see [Setting up rgbif](#setting-up-rgbif). ```{r} library(taxify) spp <- c("Quercus robur", "Pinus sylvestris", "Bellis perennis", "Morus alba") ``` ## Match first, then request Match the names, submit the request, fetch the records and retrieve the citations: ```{r} matched <- taxify(spp, backbone = "gbif") # resolve the names to GBIF keys dl <- gbif_request(matched) # ask GBIF for every key recs <- gbif_fetch(dl) # wait, download, attach the input names cite(recs) # the download's DOI ``` The matching result is a data frame. Inspect it before submitting the request: ```{r} matched <- taxify(spp, backbone = "gbif") matched[, c("input_name", "accepted_name", "accepted_id", "taxonomic_status", "n_ids")] #> input_name accepted_name accepted_id taxonomic_status n_ids #> 1 Quercus robur Quercus robur 2878688 ACCEPTED 3 #> 2 Pinus sylvestris Pinus sylvestris 5285637 ACCEPTED 6 #> 3 Bellis perennis Bellis perennis 3117424 ACCEPTED 1 #> 4 Morus alba Morus alba 5361889 ACCEPTED 1 #> Sources: GBIF 2023.08 | cite() for full citations ``` `accepted_id` contains the selected GBIF key. `n_ids` gives the number of keys associated with the name; `taxify_ids()` lists them individually. By default, `gbif_request()` requests all associated keys. GBIF prepares the download in the background, and `gbif_fetch()` waits for it to finish: ```{r} dl <- gbif_request(matched) recs <- gbif_fetch(dl) ``` The returned `input_name` column links records to the queried names for counting or grouping. `cite(recs)` reports the package, backbone and download citations; see [Citing the download](#citing-the-download). ## Match and request in one call `gbif_request()` also accepts a vector of names. It calls `taxify()` internally, passing through matching arguments such as `fuzzy`, `kingdom` and `region`: ```{r} dl <- gbif_request(spp) # match, then ask GBIF for every key recs <- gbif_fetch(dl) # wait, download, import, attach the input names ``` The matched table is available afterwards as `attr(dl, "taxa")`. Call `taxify()` separately if you want to inspect the matches before submitting a request. Use `dry_run = TRUE` to inspect the keys and estimated record count without contacting GBIF: ```{r} gbif_request(spp, dry_run = TRUE) #> 11 GBIF key(s) from 4 matched name(s). About 4,090,085 occurrence record(s) across them. #> [1] 2878688 7911626 7586523 5285637 7718215 8116613 9079676 8251433 7393648 #> [10] 3117424 5361889 ``` ## Why one name sometimes gives several keys A name can have several GBIF keys because it is a homonym shared by different taxa, or because the backbone contains doubtful or duplicate entries alongside an accepted entry. `taxify_ids()` returns one row per name and key, including its occurrence count: ```{r} taxify(spp, backbone = "gbif") |> taxify_ids() #> input_name accepted_id accepted_name taxonomic_status n_occurrences is_pick #> 1 Quercus robur 2878688 Quercus robur ACCEPTED 1725528 TRUE #> 2 Quercus robur 7911626 Quercus robur DOUBTFUL 2 FALSE #> 3 Quercus robur 7586523 Quercus robur DOUBTFUL 9 FALSE #> 4 Pinus sylvestris 5285637 Pinus sylvestris ACCEPTED 1204947 TRUE #> 5 Pinus sylvestris 7718215 Pinus sylvestris DOUBTFUL 0 FALSE #> 6 Pinus sylvestris 8116613 Pinus sylvestris DOUBTFUL 0 FALSE #> 7 Pinus sylvestris 9079676 Pinus sylvestris DOUBTFUL 0 FALSE #> 8 Pinus sylvestris 8251433 Pinus sylvestris DOUBTFUL 0 FALSE #> 9 Pinus sylvestris 7393648 Pinus sylvestris DOUBTFUL 0 FALSE #> 10 Bellis perennis 3117424 Bellis perennis ACCEPTED 1107989 TRUE #> 11 Morus alba 5361889 Morus alba ACCEPTED 51610 TRUE ``` `is_pick` marks the key reported in `accepted_id`. The other rows are further records of the same name. Requesting only the pick would have missed the eleven records under the two doubtful *Quercus robur* keys. `n_occurrences` is the count taken when the backbone was built. For an accepted key it includes the records of that taxon's synonyms and descendants, which is what a download by that key returns, so it gives the approximate size of a request. ## Choosing what to request Every key is requested by default. `strict = TRUE` narrows the request to the single key reported in `accepted_id`: ```{r} gbif_request(spp, strict = TRUE, dry_run = TRUE) #> 4 GBIF key(s) from 4 matched name(s). About 4,090,074 occurrence record(s) across them. #> [1] 2878688 5285637 3117424 5361889 ``` For this list the two choices differ by 11 records out of 4.09 million, so requesting every key costs almost nothing and keeps records that would otherwise be dropped. For a list with many homonyms the difference is larger, and `taxify_ids()` shows it before choosing. ## Setting up rgbif Sending the request needs rgbif: ```{r} install.packages("rgbif") ``` `method = "search"` works with nothing further. `method = "download"` is an authenticated call and needs a GBIF account, which is free: register at [gbif.org](https://www.gbif.org/). rgbif reads that account from three environment variables. Its own documentation recommends keeping them in `.Renviron`, which also keeps them out of scripts: ```{r} usethis::edit_r_environ() ``` Add the three lines, with no quotes and no spaces around the `=`: ``` GBIF_USER=yourname GBIF_PWD=yourpassword GBIF_EMAIL=name@example.org ``` Save, then restart R so the file is read. rgbif also accepts the lower-case names `gbif_user`, `gbif_pwd` and `gbif_email` in `.Rprofile` as an alternative; see `?rgbif::occ_download`. To confirm GBIF accepts the account, ask it for the account's past downloads. That call is authenticated, so it only succeeds when the credentials are right: ```{r} rgbif::occ_download_list(limit = 5) ``` If a variable is missing, `gbif_request()` stops before contacting GBIF and names which one. If the password is wrong, GBIF answers 401. ## Sending the request `method = "search"` sends unauthenticated searches and returns the records directly. It needs no account, and the GBIF search API limits how deep it will page, so it suits a look at a few taxa: ```{r} recs <- gbif_request(spp, method = "search", strict = TRUE, limit = 50) nrow(recs) #> [1] 200 ``` The records come back as one data frame with a `taxon_key_requested` column, so a row can be traced back to the key that asked for it. `method = "download"` is the default. It submits an asynchronous GBIF download, which has no record cap and is issued a citable DOI. It needs a GBIF account: ```{r} dl <- gbif_request(spp, method = "download") recs <- gbif_fetch(dl) ``` `gbif_fetch()` waits for the download, retrieves it, imports it, and attaches the queried names. All of the keys go into one download: they are sent as a single `taxonKey IN (...)` predicate, so the whole list comes back as one dataset under one DOI. GBIF prepares a download of four million records on its own servers, which takes a while, so the call returns as soon as the request is queued. ### Long lists are split across downloads GBIF caps a download query at 12,000 characters. A taxonKey predicate costs about 10 characters per key, so roughly 1,187 keys fit, and a checklist of a few hundred names can pass that once its homonyms are counted. Above the limit `gbif_request()` splits the keys into chunks of 1,000 and submits them through rgbif's `occ_download_queue()`, which respects GBIF's rule of three concurrent downloads per user: ```{r} dl <- gbif_request(long_species_list, method = "download") #> 2431 keys exceed GBIF's query limit; splitting into 3 downloads of up to #> 1000 keys. Each gets its own DOI. ``` `dl` is then a vector of download keys, and `gbif_fetch()` waits for all of them and stacks the records, so fetching a split request looks the same as fetching a single one: ```{r} recs <- gbif_fetch(dl) ``` The cap applies to the query naming the keys; a single download can hold tens of millions of records. ## Linking the records back to the queried names The records carry GBIF taxon keys. Linking them to the names asked for is a second step, and `gbif_backmatch()` does it, adding `requested_key`, `input_name` and `accepted_name`: ```{r} recs <- gbif_request(spp, method = "search", limit = 100) recs <- gbif_backmatch(recs, recs) table(recs$input_name, useNA = "ifany") ``` `gbif_fetch()` calls it itself, so this is only needed when the records were fetched some other way or with `backmatch = FALSE`: ```{r} recs <- gbif_fetch(dl, backmatch = FALSE) recs <- gbif_backmatch(recs, dl) ``` It takes the object `gbif_request()` returned, which carries the key table as an attribute. A `taxify_ids()` table or a `taxify()` result works too, so records imported by hand can still be linked back: ```{r} matched <- taxify(spp, backbone = "gbif") recs <- gbif_backmatch(recs, matched) ``` ### Why not just join on taxonKey GBIF returns the records of a key's descendants along with its own, and a record identified below species level carries its own key. Requesting *Pinus nigra* (5284809) returns records whose `taxonKey` is 5686674, *P. nigra* subsp. *salzmannii*; those rows carry the requested key in `speciesKey` and nowhere else. A join on `taxonKey` misses those rows. In a sample of 300 records from that request, 50 were identified to a subspecies, and a join on `taxonKey` linked the other 250. `gbif_backmatch()` linked all 300. For each record it reads the keys in order, `taxonKey`, then `acceptedTaxonKey`, `speciesKey`, `genusKey` and `familyKey`, and takes the first one that was requested. Some records carry none of the requested keys. Those rows are kept, with `NA` in `requested_key`, `input_name` and `accepted_name`, and a message reports how many there were. ## Citing the download A GBIF download gets a DOI once it is ready, and citing it is what lets someone else retrieve the same records. The download key travels with the records, so `cite()` reports the DOI next to the backbone the names were matched against: ```{r} recs <- gbif_fetch(dl) cite(recs) #> ── taxify citations ──────────────────────────────────────────────── #> [1] Colling G (2026). taxify: Offline Taxonomic Name Matching (version 0.6.0). #> [2] GBIF Secretariat (2024). GBIF Backbone Taxonomy. doi:10.15468/39omei #> [3] GBIF.org (2026-09-23) GBIF Occurrence Download https://doi.org/10.15468/dl.ckzs32 #> ──────────────────────────────────────────────────────────── ``` `cite()` reads the DOI from GBIF when it is called, because GBIF issues it only once the download has finished preparing. A download still running is reported as having no DOI yet. `cite(recs, file = "refs.bib")` writes the same entries as BibTeX, the download included. ## Results matched against another backbone Every match above named `backbone = "gbif"`. taxify's default chain starts at COL XR, so a plain `taxify(spp)` resolves most names against another backbone and the result carries no GBIF keys. Passing such a result works: the rows without a GBIF key are re-matched against GBIF, with a message, because GBIF only accepts its own keys. ```{r} taxify(spp) |> gbif_request(dry_run = TRUE) #> No rows matched by GBIF; re-matching the names against the GBIF backbone, whose keys a GBIF request needs. #> 11 GBIF key(s) from 4 matched name(s). About 4,090,085 occurrence record(s) across them. #> [1] 2878688 7911626 7586523 5285637 7718215 8116613 9079676 8251433 7393648 #> [10] 3117424 5361889 ``` A mixed result, where the one name only GBIF carries landed on `gbif` and the rest elsewhere, is handled the same way: the rows without a key are re-matched and the rest are kept, so every name in the list reaches the request. `gbif_request()` takes matching arguments only when it does the matching itself. With a `taxify()` result as input, they go in the `taxify()` call: ```{r} taxify(spp, backbone = "gbif", kingdom = "Plantae") |> gbif_request() ``` ## See also - `taxify_ids()` lists the keys and their occurrence counts. - `vignette("large-scale")` covers lists too large to hold in memory. - `vignette("quickstart")` introduces name matching.