Start here

How OpenAlex works

OpenAlex is a map of the world’s research ecosystem: papers, people, institutions, journals, topics, and funders, linked to one another. This page covers how that map gets built and kept current.

The pipeline

Research outputs are the main artery of the system. When a researcher publishes an article, book, or dataset, information about it is registered with agencies like Crossref and DataCite, or deposited in institutional and national repositories like HAL. OpenAlex pulls records from these sources continuously.

Matching links records to known entities. Each incoming record gets matched against persistent-identifier systems: affiliation text to institutions in ROR, authors to ORCID records (and to each other — see Author disambiguation), journal titles to ISSN. This forms the foundation of the knowledge graph.

Enrichment adds the connections that make the map useful. We link outputs to other outputs by extracting reference metadata (citations), and we run text classifiers over titles and abstracts to understand what each work is about, linking it to topics, subfields, keywords, and even SDGs — see Aboutness.

Ingest sources

The core sources we pull records from include Crossref, DataCite, PubMed, HAL, DOAJ, ORCID, MAG, arXiv, Dergipark, OSTI, RePEc, Zenodo, the ISSN registry, university repositories such as UNC’s CDR and Michigan’s Deep Blue, thousands of other institutional repositories (full list), parsing of 60M+ open access PDFs, journal landing pages, direct publisher feeds — and corrections from users through community curation.

Update cadence

As new works are published (or new records of old works are minted), they are matched and added continuously — the database evolves hourly. How updates reach you depends on the access channel: the web interface and API serve the live dataset, while bulk data is released on a schedule.

Pipeline dashboard

The ingest pipeline — getting works, finding affiliations and citations, tagging aboutness — is complex and changes frequently. In the interest of openness, we share our internal monitoring dashboard publicly: view the pipeline dashboard. We can’t provide support for it, but it can be useful if you’re curious how the pipeline is running.

View as Markdown