Data Sources & Attribution

Every fact, figure, and entity on 4ort.xyz comes from open data sources we cite below. We blend CC0 datasets with US Federal public-domain APIs to build a knowledge graph for AI agents. No scraped private content, no proprietary feeds, no licensing surprises.

Architectural principle: Wikidata is our source-of-truth for entities. Federal APIs (BLS, BEA, Census, EIA, Treasury, FRB) supply economic indicators. SEC EDGAR provides corporate filings. Wikimedia provides imagery. Each layer is open or government public domain — every redistribution is legally clean.

⭐ Sources that create entities (10 live)

These datasets mint rows in the graph. Every entity records which one created it, and that source's licence is returned with the entity through the API.

Wikidata CC0 1.0 Universal

License: CC0 1.0 Universal · Ingest: Wikidata dump + EventStreams
Use: The spine of the graph: entities, claims, relationships, identifiers.

MusicBrainz CC0 1.0 Universal

License: CC0 1.0 Universal · Ingest: MusicBrainz core dump
Use: Artists, releases, labels and works — music metadata core data.

OpenAlex CC0 1.0 Universal

License: CC0 1.0 Universal · Ingest: OpenAlex snapshot
Use: Scholarly works, authors and institutions — the open citation graph.

GLEIF / Legal Entity Identifiers CC0 1.0 Universal

License: CC0 1.0 Universal · Ingest: GLEIF golden copy
Use: Legal entities worldwide: LEI, registered name, jurisdiction, legal form.

USDA FoodData Central US Government Public Domain

License: US Government Public Domain · Ingest: USDA Branded Foods
Use: Branded food products with nutrition panels and ingredients.

openFDA (NDC drugs + device UDI) US Government Public Domain

License: US Government Public Domain · Ingest: openFDA public datasets
Use: Drug products (NDC directory) and medical devices (UDI database).

NHTSA vPIC Vehicle Catalog US Government Public Domain

License: US Government Public Domain · Ingest: NHTSA vPIC
Use: Vehicle makes, models and specifications by model year.

GitHub (public repository metadata) Facts only (metadata, not code)

License: Facts only (metadata, not code) · Ingest: GitHub public API
Use: Repository facts — owner, language, stars, topics. No source code is ingested.

Hugging Face Hub (model/dataset metadata) Facts only (metadata, not weights)

License: Facts only (metadata, not weights) · Ingest: Hugging Face public API
Use: Model and dataset facts — author, task, library. No weights are ingested.

SEC EDGAR registrants US Government Public Domain

License: US Government Public Domain · Ingest: SEC EDGAR submissions
Use: US public-company registrants minted as entities from EDGAR filings data.

🔧 Sources that enrich existing entities (7)

These add columns, signals and imagery to entities that already exist.

SEC EDGAR (filings, Form 4, Exhibit 21, XBRL) US Government Public Domain

License: US Government Public Domain · Ingest: SEC EDGAR full-text + structured filings
Use: Filing history, insider transactions, subsidiary lists and XBRL financial facts.

Wikipedia Pageviews CC0 1.0 Universal

License: CC0 1.0 Universal · Ingest: Wikimedia dumps
Use: Real-world attention signal powering popularity, trending and rankings.

Wikipedia Clickstream CC0 1.0 Universal

License: CC0 1.0 Universal · Ingest: Wikimedia dumps
Use: Reader navigation edges between entities.

Wikimedia sitelinks + aliases CC0 1.0 Universal

License: CC0 1.0 Universal · Ingest: Wikidata dump
Use: Cross-language names and article links used for entity resolution.

Wikimedia Commons (public-domain only) Public Domain files only

License: Public Domain files only · Ingest: Commons dumps, PD-filtered
Use: Entity imagery, filtered to public-domain files.

DomCop OpenPageRank Free tier, attribution required

License: Free tier, attribution required · Ingest: OpenPageRank API · Attribution required
Use: Domain authority signal for source ranking and agent-accessibility scoring.

Common Crawl CC BY / Common Crawl ToU

License: CC BY / Common Crawl ToU · Ingest: Common Crawl index · Attribution required
Use: Web-page discovery for the agent-accessibility index.

📊 Federal APIs (attribution disclaimers required)

US federal statistical series. Public domain, but each agency requires that we state the data is theirs and that they do not endorse us.

Federal Reserve Board (H.15) US Government Public Domain

License: US Government Public Domain · Ingest: Federal statistical API · Attribution required
Use: Benchmark interest rates and yields.

US Treasury Fiscal Data US Government Public Domain

License: US Government Public Domain · Ingest: Federal statistical API · Attribution required
Use: Federal debt, revenue and outlay series.

US Bureau of Labor Statistics US Government Public Domain

License: US Government Public Domain · Ingest: Federal statistical API · Attribution required
Use: Employment, CPI and wage series.

US Bureau of Economic Analysis US Government Public Domain

License: US Government Public Domain · Ingest: Federal statistical API · Attribution required
Use: GDP and national accounts series.

US Census Bureau US Government Public Domain

License: US Government Public Domain · Ingest: Federal statistical API · Attribution required
Use: Population and business statistics.

US Energy Information Administration US Government Public Domain

License: US Government Public Domain · Ingest: Federal statistical API · Attribution required
Use: Energy production, price and consumption series.

CIA World Factbook US Government Public Domain

License: US Government Public Domain · Ingest: Federal statistical API · Attribution required
Use: Country reference indicators.

🛰️ News-derived layer (inferred, labelled, never mixed with CC0)

Entities and facts we derive from news coverage rather than import from an open dataset. They live in a separate inferred layer with their own trust labels, and every one carries the source URLs it was derived from. Nothing here is presented as Wikidata-grounded.

GDELT Project (GKG + Event database) Open data, attribution required

License: Open data, attribution required · Ingest: GDELT 15-minute global news feed · Attribution required
Use: News-derived entity discovery (the Living Edge frontier), country stability signals and recent-article matching. Entities minted from GDELT are stored in the inferred layer, never mixed into the CC0 Wikidata-grounded layer.
Requires >=2 independent publishers before an entity or fact is minted.

Press-release wires (PR Newswire, Business Wire, GlobeNewswire) First-party corporate statements (facts, attributed)

License: First-party corporate statements (facts, attributed) · Ingest: Public RSS feeds + syndicated release pages · Attribution required
Use: Product and company facts extracted from press releases. Facts sourced only to the subject's own release are labelled `vendor_claimed` and never presented as independently verified; a second independent publisher promotes them to `corroborated`.
Extraction is local (GLiNER/Relex), batch-only — never in a request path.

🔒 Licensed, internal use only (never redistributed)

We license these for internal ranking and prioritisation. They are not published on any page, not returned by any API, and are not part of the redistributable dataset. Listed here because complete provenance means listing what we use, not only what we publish.

DataForSEO Commercial licence — internal use only

License: Commercial licence — internal use only · Ingest: Licensed commercial API · Not redistributed
Use: Search-volume and CPC data licensed for INTERNAL ranking and prioritisation only. It is never published on any public page, never returned by any API response, and is not part of the redistributable dataset.

🔮 Coming soon (planned ingest, all CC0/PD)

Not yet in the graph. Listed so you can see where coverage is heading.

Open Library CC0 1.0 Universal

License: CC0 1.0 Universal · Ingest: Planned ingest
Use: 30M+ book records with ISBNs, editions and author bibliographies.

ROR (Research Organization Registry) CC0 1.0 Universal

License: CC0 1.0 Universal · Ingest: Planned ingest
Use: Canonical identifiers for 110k research institutions worldwide.

ORCID public data file CC0 1.0 Universal

License: CC0 1.0 Universal · Ingest: Planned ingest
Use: Researcher identifiers and affiliations.

PubMed / MEDLINE US Government Public Domain

License: US Government Public Domain · Ingest: Planned ingest
Use: Biomedical literature and clinical research records.

ClinicalTrials.gov US Government Public Domain

License: US Government Public Domain · Ingest: Planned ingest
Use: Clinical trial registrations, daily refreshed.

USPTO PatentsView US Government Public Domain

License: US Government Public Domain · Ingest: Planned ingest
Use: Patents linked to inventors and assignee entities.

Natural Earth Open Data Commons PDDL

License: Open Data Commons PDDL · Ingest: Planned ingest
Use: Public-domain geographic boundaries and place data.

🎯 What we DON'T use

Last updated 2026-05-16 · Questions? support@4ort.ai