Keywords
A keyword is a short phrase describing what a work is about — machine-learning, type-1-diabetes, citation. Keywords are one of OpenAlex’s aboutness signals: finer-grained than topics and a good fit for narrow, specific slices of the literature. There are about 65,000 keywords in the system, and each work can be tagged with up to five. A keyword’s OpenAlex ID is a readable slug rather than the usual letter-and-number code — machine-learning — so a keyword looks like https://openalex.org/keywords/machine-learning; fetch one at api.openalex.org/keywords/machine-learning.
How it’s made
Keywords are derived from topics: every topic carries a curated set of associated keywords, and a work’s keywords are drawn from the keyword sets of the topics it was assigned. Tagging a work runs in four steps:
- Gather candidates. Take the work’s assigned topics (up to three) and pull the keywords associated with each. This yields up to 30 candidate keywords.
- Score similarity. Score each candidate against the work’s title and abstract using embeddings from the BGE M3-Embedding model, a multilingual embedding model that captures semantic meaning — so a keyword can match even when its exact phrase never appears in the text, and across languages.
- Apply a threshold. Keep only candidates whose similarity score clears a threshold. This filters out keywords that belong to the work’s topic area but aren’t really relevant to the specific work.
- Keep the top five. From the keywords that pass, the five highest-scoring are assigned to the work.
Each keyword on a work therefore carries a score — its similarity to that work’s title and abstract. On the keyword object itself, works_count and cited_by_count roll those assignments up across the whole corpus.
The keyword-extraction pipeline is open source: openalex-keywords (v2) on GitHub. See Aboutness for how keywords compare to topics, SDGs, and text search.
Attributes
This is the canonical dictionary of every attribute on a keyword object. Attributes shared with other entities are documented once on Common attributes; keyword-specific notes are below.
id
String. The OpenAlex ID for this keyword. Unlike most entities, a keyword’s ID is a readable slug rather than a letter-and-number code, e.g. https://openalex.org/keywords/machine-learning. See Common attributes.
display_name
String. The keyword’s human-readable label, e.g. Machine learning. See Common attributes.
works_count
Integer. The number of works tagged with this keyword. See Common attributes.
cited_by_count
Integer. The total citations received by all works tagged with this keyword. See Common attributes.
works_api_url
String. A ready-made Works API URL returning every work tagged with this keyword, e.g. https://api.openalex.org/works?filter=keywords.id:keywords/machine-learning. A convenience link — it’s the same query you’d build with the keywords.id filter.
created_date
String. The date this keyword was added to OpenAlex (YYYY-MM-DD). See Common attributes.
updated_date
String. The ISO 8601 UTC timestamp of the last change to this keyword object. See Common attributes.
In the API
The Keywords endpoint is at api.openalex.org/keywords. Fetch a single keyword by its slug ID — /keywords/machine-learning — or a list, and filter, search, sort, group, and page over the fields above.
To find the works carrying a keyword, filter on the Works endpoint:
https://api.openalex.org/works?filter=keywords.id:machine-learning
For the full list of filterable, sortable, and groupable fields see the Keywords API reference; for all endpoints see the endpoints index.