Provenance and verification
Every record states how much of it has been confirmed and how recently it changed. The rules follow the provenance discipline of scholarly-corpus-builder: identifiers before plausibility, open sources before anything else, and no value upgraded because it merely looks right.
Verification states
| State | Rule | Records |
|---|---|---|
| Verified | Persistent identifiers (ISSN, DOI prefix) checked and publisher evidence reviewed under PQF. | 31 |
| Partially verified | ISSN resolved to an OpenAlex source record; journal-level evidence not reviewed by POSI. | 992 |
| Needs check | Harvested from an open registry and plausible, but not independently confirmed. Do not use for high-confidence claims. | 23,769 |
| Rejected | Could not be confirmed, or evidence contradicts the registry record. Retained only as an audit trail. | 0 |
The state describes the record's identity, not the journal's quality. A Discovered record marked needs check may describe an excellent journal; it only means POSI has not confirmed the record yet.
Freshness
- Current24,792
Record updated within 6 months of the data cutoff.
- Aging0
Record last updated 6-12 months before the data cutoff.
- Stale0
Record last updated more than 12 months before the data cutoff. Refresh due.
- Unknown0
No reliable update timestamp on the record.
Journal records use a 6 to 12 month window. Freshness is computed from each record's own update timestamp against the data cutoff; it is never assumed.
Source order
- 1
Persistent identifiers
ISSN, ISSN-L, DOI prefix, OpenAlex source id. Checked against the issuing registry before anything is called verified.
- 2
Open registry metadata
Crossref, OpenAlex, DOAJ and the ISSN Portal. Taken as-is and labelled with the registry it came from.
- 3
Publisher declarations
Frequency, peer-review model, license and country as stated by the journal. Recorded as declared, never as observed.
- 4
Computed values
Subjects, ratings and citation indicators from posi-engine under a named, versioned specification.
Each field's basis is listed in the record schema.
Duplicates and identity
Records are merged on identifiers in this order: ISSN-L, canonical ISSN pair, OpenAlex source id. Title similarity alone never merges two records. Every possible duplicate found during the identity migration was resolved with a live lookup: 166 groups merged, 5 confirmed distinct, none left ambiguous. The audit is linked from Datasets.
Unknown values stay unknown. An ISSN, year, publisher or license that was not found is stored as null and shown as "Not recorded", never filled with a best guess.
Small samples
Citation indicators computed from very few items are labelled on every record page: fewer than 5 items is illustrative, 5 to 19 is a limited sample. These labels travel with the number so a value from three articles is never read as if it came from three hundred.
Declared versus observed
What a journal says about itself (peer-review model, frequency, license) is kept apart from what POSI or a registry observed (registered articles, DOAJ listing, ISSN country). Record pages show them in separate sections so a declaration is never presented as a measurement.