Counts by publication year
Welcome to the KB information portal of the Kompetenznetzwerk Bibliometrie (KB). This portal provides an interactive exploration of the OpenAlex database, which is integrated into the KB data infrastructure (Schmidt et al. 2025).
Bibliometric data is an important resource for understanding science and research. It helps reveal patterns of knowledge production, collaboration, and dissemination that would otherwise remain hidden. Policymakers use bibliometrics to monitor research trends, institutions use it for evaluation and strategy, and researchers can reflect on how their fields are represented and recognized. At the same time, bibliometric databases never offers a complete picture. Databases differ in coverage, structure, and focus, which influences the picture they show about research.
OpenAlex is one of the largest openly available bibliometric databases, covering scholarly works, authors, institutions, journals, publishers, and funders worldwide. Unlike commercial databases, it is free and transparent, making it a valuable foundation for open and reproducible analyses. Within KB, OpenAlex is used both as a research object and as a practical infrastructure component. The KB enriches and adapts the data with its expertise, especially in the context of the German science system, to make the information more accurate and useful.
This dashboard allows users to explore OpenAlex data across multiple perspectives. By making complex structures visible, it aims to support both specialists and non-specialists in interpreting trends, identifying biases, and understanding the diversity of scholarly communication.
The dashboard consists of six main pages, each offering a different perspective on the data:
Entities by Year This perspective shows trends in the number of items, journals, publishers, organizations, and authors over time - both in the entire database and among items affiliated with Germany. Growth in entities can reflect real structural changes in research, such as the emergence of new publishers, or the founding of new journals. However, it can also reveal changes in the database itself, as coverage and indexing practices evolve over time. For this reason, the dashboard also allows comparison across different snapshots of OpenAlex, with future snapshots to be added. This feature highlights how bibliometric databases are dynamic and continuously updated, which has important implications for longitudinal analyses.
Languages This perspective compares the number of items published in different languages. While English dominates international scholarly publishing, a vast amount of research is conducted and communicated in other languages. For example, research on local education policy in Germany, agricultural practices in Latin America, or public health studies in China may be published in German, Spanish, or Chinese rather than in English. Such work is highly relevant within its local or regional context but may remain invisible in global evaluations if only English-language publications are considered. This perspective underscores the importance of linguistic diversity for a more complete understanding of global scholarly communication.
Item Types Here, users can analyze the distribution of items across different publication types. Scientific knowledge is disseminated not only through journal articles but also through books, conference proceedings, datasets, and other outputs. Recognizing these forms is crucial: neglecting books disadvantages disciplines such as the humanities and social sciences, while ignoring datasets downplays open science practices, reproducibility, and important parts of scholarly work. This perspective highlights how different types of output contribute to the research ecosystem.
Research Fields This view presents the number of items categorized under OECD research fields, a widely used classification system that enables standardized comparison across databases. Such standardization is important because research fields differ in database coverage. Bibliometric databases have historically favored fields with strong journal cultures (such as medicine or physics), while disciplines with more localized or non-journal communication patterns (such as history, law, or education) remain underrepresented. By showing how OpenAlex covers different research fields, this perspective provides insight into disciplinary diversity of bibliometric data.
Research Fields in dynamics This page provides an opportunity to track a dynamics in different research field. This could be particularly interesting for the expert in these fields.
German Institutions Finally, this perspective allows users to compare the number of items affiliated with German institutions, using two different disambiguation procedures: the standard OpenAlex institutional ID and the KB I-Kodierung. Correctly attributing research outputs to institutions is a well-known challenge in global bibliometric systems. Inconsistent or incomplete affiliation data can lead to errors that misrepresent institutional output, for example by undercounting publications or wrongly attributing them to other organizations. Building on its expertise in the German science system, KB has developed improved methods for affiliation disambiguation (data is openly available). This ensures a more accurate and reliable representation of German institutions in bibliometric analyses.
A short description of what each indicator counts, and of the conventions that apply across all pages, is on the Methodology page.
This page explains where the figures in the portal originate from, how each indicator is calculated.
A snapshot is an instantaneous copy of a database that is taken and then left unchanged. Each page allows to select the desired snapshot to display, and more than one for comparison. Snapshots are labelled by month and year (for example April 2026). The underlying OpenAlex snapshot is dated within a few weeks of the snapshot label. Bibliometric databases are continuously updated. Items are added, corrected, reclassified, and removed. Therefore, two snapshots may report different figures for the same publication year. Any differences between snapshots provide information about the database itself, rather than the research output.
Publication year. All calculations are based on the publication years from 2010 onwards. As items are still being added to it, the most recent year is usually incomplete. Each snapshot also has its own final year. For example, the April 2025 snapshot ends in 2024, whereas the April 2026 snapshot ends in 2025.
Germany-affiliated and global. The Geographical scope control on each page switches between the two. Global counts all items in the database. Germany counts only items that are flagged in the KB infrastructure as having at least one author affiliation in Germany. The whole counting method is used: an item with one German co-author out of twenty, for example, is counted as one item under both settings.
Item set. Where the selection is offered, both counts come from the same pass over the same rows.
{Article, Data Paper} is included. Both capitalisations are matched, because the two snapshots store them differently.Percentages. Each percentage represents a proportion of the total for its respective series: one database, one snapshot, one sample and one publication set. Percentages are read down a column.
Distinct items in the database, grouped by publication year. This shows the size of the database in that year.
Distinct journal identifiers appear on items from each publication year, but only for sources categorised as journals. A journal is counted in every year in which it carries at least one indexed item. Therefore, a ten-year journal appears in all ten years, and the values cannot be added together.
Distinct publisher identifiers on the items of each publication year. Items for which no publisher is recorded do not contribute. This figure is influenced by the number of active publishers and the completeness with which each database records publisher information.
Distinct organizations named in the affiliations of the items in each publication year. The global and German figures are calculated differently. Globally, all affiliations on the items from that year are used. In Germany, however, the set of items is first narrowed down to those affiliated with Germany. Only the German addresses on these items are then matched to organisations, meaning that a German item with co-authors in three countries would only contribute its German organisations.
Distinct disambiguated author identifiers affiliated with the items of each publication year. Author identity is inferred: two records for the same person increase the count, while merging two people into one identifier decreases it. For Germany, the count is restricted to authors affiliated with a German affiliation on items affiliated with Germany.
Items grouped by the language recorded for them. Where a database records more than one language, the first is used. Language codes are then mapped onto language names en/eng to English, de/ger to German, and so on for around twenty languages. Anything outside this list is categorised as other, and items with no recorded language are categorised as No data. The counts are summed over the whole period, so the Languages page shows a single figure for each language from 2010 onwards. Only the fifteen most frequent languages are displayed on the page, but each percentage represents a proportion of the total number of items indexed by the database in the selected sample. This is why the visible bars add up to less than 100%.
Items grouped by document type. Where a database records more than one type, the first is used, so an item recorded as {Article, Data Paper} is counted only as an article. This differs from the publication-set filter above, which tests the entire list. Therefore, the two need not agree exactly. Raw type labels may appear in different spellings ( e.g. Chapter, book-chapter, Book-Chapter) and are harmonized to a single vocabulary before display. A few closely related types are merged: book sections and book parts are merged as Book chapter, report components as Report. Types not in the vocabulary retain their original name rather than being folded into a residual category. “Other” means the literal type other as recorded in the database. No data means no type was recorded. As with languages, the counts are summed over the whole period and the percentages represent the proportion of everything indexed in the selected sample. Very rare types (< 1,000 items) are not displayed.
As neither database directly records OECD fields of science, both are mapped onto them via a shared OECD lookup table, using their own subject vocabularies. An item can carry several subject categories and is counted in each OECD field to which it belongs. Therefore, field counts may exceed the number of items, and the percentages may sum to more than 100%. Items with no subject category are not displayed on this page, which is why the field totals may be lower than the item totals on the Entities page.
This page compares two methods of attributing the same OpenAlex items to German institutions throughout the entire period:
Both counts refer to distinct items. The dotted diagonal line indicates perfect agreement: points above the line are institutions to which the KB procedure attributes more items, while points below the line are institutions to which OpenAlex attributes more items. The distance from the diagonal reflects the disagreement between the two methods regarding the same underlying data.
An institution appears only if both procedures produced a figure for it, and only if its KB record can be matched to an OpenAlex institution record through a shared ROR identifier. Institutions that are not recognised by one procedure are therefore absent rather than plotted at zero, and the page understates the extent of the disagreement. A small number of institutions are excluded due to known record issues. Sector labels come from the KB sector classification. An institution belonging to more than one sector may appear more than once.