Data Sources & Attribution
Every fact, figure, and entity on 4ort.xyz comes from open data sources we cite below. We blend CC0 datasets with US Federal public-domain APIs to build a knowledge graph for AI agents. No scraped private content, no proprietary feeds, no licensing surprises.
⭐ Sources that create entities (10 live)
These datasets mint rows in the graph. Every entity records which one created it, and that source's licence is returned with the entity through the API.
Wikidata CC0 1.0 Universal ↗
MusicBrainz CC0 1.0 Universal ↗
OpenAlex CC0 1.0 Universal ↗
GLEIF / Legal Entity Identifiers CC0 1.0 Universal ↗
USDA FoodData Central US Government Public Domain ↗
openFDA (NDC drugs + device UDI) US Government Public Domain ↗
NHTSA vPIC Vehicle Catalog US Government Public Domain ↗
GitHub (public repository metadata) Facts only (metadata, not code) ↗
Hugging Face Hub (model/dataset metadata) Facts only (metadata, not weights) ↗
SEC EDGAR registrants US Government Public Domain ↗
🔧 Sources that enrich existing entities (7)
These add columns, signals and imagery to entities that already exist.
SEC EDGAR (filings, Form 4, Exhibit 21, XBRL) US Government Public Domain ↗
Wikipedia Pageviews CC0 1.0 Universal ↗
Wikipedia Clickstream CC0 1.0 Universal ↗
Wikimedia sitelinks + aliases CC0 1.0 Universal ↗
Wikimedia Commons (public-domain only) Public Domain files only ↗
DomCop OpenPageRank Free tier, attribution required ↗
Common Crawl CC BY / Common Crawl ToU ↗
📊 Federal APIs (attribution disclaimers required)
US federal statistical series. Public domain, but each agency requires that we state the data is theirs and that they do not endorse us.
Federal Reserve Board (H.15) US Government Public Domain ↗
US Treasury Fiscal Data US Government Public Domain ↗
US Bureau of Labor Statistics US Government Public Domain ↗
US Bureau of Economic Analysis US Government Public Domain ↗
US Census Bureau US Government Public Domain ↗
US Energy Information Administration US Government Public Domain ↗
CIA World Factbook US Government Public Domain ↗
🛰️ News-derived layer (inferred, labelled, never mixed with CC0)
Entities and facts we derive from news coverage rather than import from an open dataset. They live in a separate inferred layer with their own trust labels, and every one carries the source URLs it was derived from. Nothing here is presented as Wikidata-grounded.
GDELT Project (GKG + Event database) Open data, attribution required ↗
Press-release wires (PR Newswire, Business Wire, GlobeNewswire) First-party corporate statements (facts, attributed) ↗
🔒 Licensed, internal use only (never redistributed)
We license these for internal ranking and prioritisation. They are not published on any page, not returned by any API, and are not part of the redistributable dataset. Listed here because complete provenance means listing what we use, not only what we publish.
DataForSEO Commercial licence — internal use only ↗
🔮 Coming soon (planned ingest, all CC0/PD)
Not yet in the graph. Listed so you can see where coverage is heading.
Open Library CC0 1.0 Universal ↗
ROR (Research Organization Registry) CC0 1.0 Universal ↗
ORCID public data file CC0 1.0 Universal ↗
PubMed / MEDLINE US Government Public Domain ↗
ClinicalTrials.gov US Government Public Domain ↗
USPTO PatentsView US Government Public Domain ↗
Natural Earth Open Data Commons PDDL ↗
🎯 What we DON'T use
- No scraped Wikipedia article text — Wikipedia prose is CC BY-SA (share-alike); we ingest only Wikidata's structured data and Wikimedia's CC0 statistics.
- No FRED — its terms restrict AI use; we take the underlying series from the Federal Reserve Board directly instead.
- No proprietary financial market feeds, and no private or surveillance data.
- No licensed commercial dataset is ever redistributed. Where we license data for internal ranking (see DataForSEO above), it stays internal by design.