GBIF Marks 25 Years of FAIR Biodiversity Data, Now Spanning 70 Countries and 3.8 Billion Records

As GBIF marks 25 years and 3.8 billion records across 70 countries, a new overview looks ahead – tackling data gaps to shape the next quarter-century.

Digital biodiversity data has come a long way. Twenty-five years ago, the Global Biodiversity Information Facility (GBIF) launched with 21 founding countries and a simple but ambitious goal: make the world’s biodiversity data freely and openly available to anyone, anywhere.

Today, GBIF’s network includes 70 participating countries, along with data providers across more than 120 countries, offering open access to over 3.8 billion biodiversity data records. This milestone is explored in a new overview article published in the Biodiversity Data Journal, which traces GBIF’s trajectory over the past quarter-century and the challenges that remain as it looks ahead.

GBIF Executive Secretary Dr. Joe Miller remarked:

As GBIF celebrates its 25th anniversary, I’m very proud to serve as Executive Secretary of the organisation and to have contributed to this milestone publication. Detailing the history of GBIF both as an infrastructure and a global network, the article pays tribute to all the people involved in the network throughout a quarter century—be they staff, participants, nodes, volunteers, data publishers or data users. It represents the past, present and future of GBIF. I believe the paper will help strengthen our network and its capacity to respond to biodiversity data needs both now and in the future.

GBIF was established in 2001, following a proposal from the Organisation for Economic Cooperation and Development’s Megascience Forum which concluded that “an international mechanism is needed to make biodiversity data and information accessible worldwide”.

The call came in response to a problem facing the international community after the 1992 Rio Earth Summit: the world’s biodiversity knowledge was scattered across a small number of institutions, concentrated in wealthy countries, with no shared way to bring it together in support of the newly signed Convention on Biological Diversity.

Now, a quarter-century later, GBIF has become the foundational infrastructure for biodiversity science and policy worldwide.

A Growing Global Network

Biodiversity data in GBIF
Annual biodiversity data made available through the Global Biodiversity Information Facility as of July 2026. Credit to GBIF.

GBIF’s work is carried out through a distributed network of 47 voting country participants and 42 organisational participants, each represented by a national or thematic “node” that mobilises data, builds local capacity, and connects biodiversity communities to GBIF’s infrastructure. More than 2,700 institutions, including museums, universities, government agencies, citizen science platforms, and a growing number of private-sector organisations, have published data through the network, with new publishers joining at a rate of more than two every three days, totalling nearly 3500 organisations involved.

key numerical metrics of the Global Biodiversity Information Facility
Infographic summary of key numerical metrics of the Global Biodiversity Information Facility as of August 2026. See dynamic metrics at GBIF.org and at https://www.gbif.org/analytics/global.

The Biodiversity Data Journal itself is one example of this network in action, as it operates its own GBIF-hosted data portal – one of more than 20 such portals across Pensoft-published journals – providing direct access to hundreds of datasets and hundreds of thousands of occurrence records drawn from the journal’s publications.

Data Supporting Science and Policy

GBIF data mediation
The universe of GBIF data mediation.

GBIF-mediated data now underpin an average of eight new peer-reviewed research papers per day, and have contributed to more than 15,000 publications to date, spanning fields from climate change and food security to public health and invasive species management. An independent 2023 assessment estimated that GBIF generates roughly €12 in societal benefit for every €1 invested, with researcher time savings alone valued at €35 million annually.

The organisation’s impact also extends deeply into international policy. GBIF-mediated data support multiple indicators under the Kunming-Montreal Global Biodiversity Framework, inform assessments by the International Union for Conservation of Nature (IUCN) and the Intergovernmental Science-Policy Platform on Biodiversity and Ecosystem Services (IPBES), and are increasingly used by the private sector to meet emerging nature-related disclosure requirements.

Adapting to New Kinds of Data

Over 25 years, GBIF has continually expanded to accommodate new sources of biodiversity data – from digitised natural history specimens and citizen science observations to DNA-based records and, most recently, structured survey and monitoring data designed to support large-scale biodiversity tracking. In 2026, GBIF adopted the Catalogue of Life as its taxonomic backbone, the result of a multi-year collaboration to build shared infrastructure for reconciling species names across datasets.

Persistent Gaps Remain

key events of GBIF

A timeline of key events in the history of the Global Biodiversity Information Facility. Credit to Miller et al., 2026.

Despite this growth, GBIF recognises that several significant challenges remain. The global loss of biodiversity continues to be described as a crisis in major international assessments, even as the data infrastructure has improved. And the data itself continues to reflect long-standing imbalances – regions with high biodiversity, particularly in the Global South, remain underrepresented, while well-studied, charismatic groups such as birds are overrepresented compared to less visible taxa; furthermore, many marine species and hard-to-identify cryptic habitats and organisms remain inadequately documented.

Closing these gaps was part of the original motivation for creating GBIF a quarter-century ago, and it remains valid today. GBIF’s own assessment is that the network cannot resolve these asymmetries alone, doing so will depend on broader shifts in global scientific funding, data culture, and capacity, alongside GBIF’s continued efforts to expand its Participant network and to diversify the types of data it can mobilise.

An additional, ongoing challenge is data heterogeneity and data quality, because GBIF indexes data, rather than directly vetting every record, it relies on its network of data publishers to maintain quality at source, and on users to report issues – a distributed model that keeps the network scalable but is not without friction.

Looking Ahead

capacity development at GBIF.
Capacity development across the Global Biodiversity Information Facility community.  Credit to Miller et al., 2026.

Guided by its 2023–2027 Strategic Framework, GBIF is prioritising continued growth of its Participant network, expanded support for survey and monitoring data, and deeper integration of DNA-derived biodiversity records – all aimed at closing longstanding geographic and taxonomic gaps in digital biodiversity data worldwide.

Original source:

Miller J, Mandeville C, Schigel D, Gamboa Martinez J, Nielsen AM, Sheldon S, Hahn A, Raymond M, Bundgaard-Jensen S, Bagard Laursen M, Copas K, Blissett M, Høfft M, Méndez Hernández F, Noesgaard D, Rodrigues A, Russell L, Stjernegaard Jeppesen T, Elkjær Ørum-Kristensen A, Grosjean M, Waller J, Podolskiy M, Goodson H, Novakovikj S, Frøslev T, Ingenloff K, Sørensen Nilsson A, Svenningsen C, van der Meijden D, Suen A, Hakan Uzun A, Vaskova M, Nielsen C, Marentes Herrera E, Schaldemose Reibke N, Robertson T (2026) The Global Biodiversity Information Facility at 25. Biodiversity Data Journal 14: e208528. https://doi.org/10.3897/BDJ.14.e208528

FAIR biodiversity data in Pensoft journals thanks to a routine data auditing workflow

Data audit workflow provided for data papers submitted to Pensoft journals.

To avoid publication of openly accessible, yet unusable datasets, fated to result in irreproducible and inoperable biological diversity research at some point down the road, Pensoft takes care for auditing data described in data paper manuscripts upon their submission to applicable journals in the publisher’s portfolio, including Biodiversity Data Journal, ZooKeys, PhytoKeys, MycoKeys and many others.

Once the dataset is clean and the paper is published, biodiversity data, such as taxa, occurrence records, observations, specimens and related information, become FAIR (findable, accessible, interoperable and reusable), so that they can be merged, reformatted and incorporated into novel and visionary projects, regardless of whether they are accessed by a human researcher or a data-mining computation.

As part of the pre-review technical evaluation of a data paper submitted to a Pensoft journal, the associated datasets are subjected to data audit meant to identify any issues that could make the data inoperable. This check is conducted regardless of whether the dataset are provided as supplementary material within the data paper manuscript or linked from the Global Biodiversity Information Facility (GBIF) or another external repository. The features that undergo the audit can be found in a data quality checklist made available from the website of each journal alongside key recommendations for submitting authors.

Once the check is complete, the submitting author receives an audit report providing improvement recommendations, similarly to the commentaries he/she would receive following the peer review stage of the data paper. In case there are major issues with the dataset, the data paper can be rejected prior to assignment to a subject editor, but resubmitted after the necessary corrections are applied. At this step, authors who have already published their data via an external repository are also reminded to correct those accordingly.

“It all started back in 2010, when we joined forces with GBIF on a quite advanced idea in the domain of biodiversity: a data paper workflow as a means to recognise both the scientific value of rich metadata and the efforts of the the data collectors and curators. Together we figured that those data could be published most efficiently as citable academic papers,” says Pensoft’s founder and Managing director Prof. Lyubomir Penev.
“From there, with the kind help and support of Dr Robert Mesibov, the concept evolved into a data audit workflow, meant to ‘proofread’ the data in those data papers the way a copy editor would go through the text,” he adds.
“The data auditing we do is not a check on whether a scientific name is properly spelled, or a bibliographic reference is correct, or a locality has the correct latitude and longitude”, explains Dr Mesibov. “Instead, we aim to ensure that there are no broken or duplicated records, disagreements between fields, misuses of the Darwin Core recommendations, or any of the many technical issues, such as character encoding errors, that can be an obstacle to data processing.”

At Pensoft, the publication of openly accessible, easy to access, find, re-use and archive data is seen as a crucial responsibility of researchers aiming to deliver high-quality and viable scientific output intended to stand the test of time and serve the public good.

CASE STUDY: Data audit for the “Vascular plants dataset of the COFC herbarium (University of Cordoba, Spain)”, a data paper in PhytoKeys

To explain how and why biodiversity data should be published in full compliance with the best (open) science practices, the team behind Pensoft and long-year collaborators published a guidelines paper, titled “Strategies and guidelines for scholarly publishing of biodiversity data” in the open science journal Research Ideas and Outcomes (RIO Journal).

Audit finds biodiversity data aggregators ‘lose and confuse’ data

In an effort to improve the quality of biodiversity records, the Atlas of Living Australia (ALA) and the Global Biodiversity Information Facility (GBIF) use automated data processing to check individual data items. The records are provided to the ALA and GBIF by museums, herbaria and other biodiversity data sources.

However, an independent analysis of such records reports that ALA and GBIF data processing also leads to data loss and unjustified changes in scientific names.

The study was carried out by Dr Robert Mesibov, an Australian millipede specialist who also works as a data auditor. Dr Mesibov checked around 800,000 records retrieved from the Australian Museum, Museums Victoria and the New Zealand Arthropod Collection. His results are published in the open access journal ZooKeys, and also archived in a public data repository.

“I was mainly interested in changes made by the aggregators to the genus and species names in the records,” said Dr Mesibov.

“I found that names in up to 1 in 5 records were changed, often because the aggregator couldn’t find the name in the look-up table it used.”

data_auditAnother worrying result concerned type specimens – the reference specimens upon which scientific names are based. On a number of occasions, the aggregators were found to have replaced the name of a type specimen with a name tied to an entirely different type specimen.

The biggest surprise, according to Dr Mesibov, was the major disagreement on names between aggregators.

“There was very little agreement,” he explained. “One aggregator would change a name and the other wouldn’t, or would change it in a different way.”

Furthermore, dates, names and locality information were sometimes lost from records, mainly due to programming errors in the software used by aggregators to check data items. In some data fields the loss reached 100%, with no original data items surviving the processing.

“The lesson from this audit is that biodiversity data aggregation isn’t harmless,” said Dr Mesibov. “It can lose and confuse perfectly good data.”

“Users of aggregated data should always download both original and processed data items, and should check for data loss or modification, and for replacement of names,” he concluded.

###

Original source:

Mesibov R (2018) An audit of some filtering effects in aggregated occurrence records. ZooKeys 751: 129-146. https://doi.org/10.3897/zookeys.751.24791