Data
The real corpus-inventory metadata behind this showcase, downloadable with correct content-types. 12 files: the full 266-cluster series inventory, panel-capability counts, and the headline campaign totals.
Real corpus metadata — no constructed series. Roadmap Every file here is a genuine measurement of the extracted knowledge base (per-cluster document counts, year spans, extracted-table counts), served verbatim or as a faithful CSV serialization of a real JSON structure. There are no harmonized economic time-series files, because none are built yet — the Anu data-construction pipeline is scoped, not constructed. The Cluster A fiscal pilot (1856–1990) is first in line; see Methodology for the full plan and limits. The 4.4 GB knowledge base and ~33 GB of outputs are not shipped.
Download everything
One zip of all 12 corpus-inventory artifacts plus a provenance README that states exactly what is real and what is not yet built.
Robert-DB ingestion (enriched-table catalog)
| Artifact | Schema / notes | Format | Size | Download |
|---|---|---|---|---|
| Enriched-table catalog (full, ~32k tables) | table_uid, doc_id, title_en, title_raw, topic, geo_facet, geography, units, period_coverage, start_year, n_rows, n_cols, quality_grade, obs_status, transcription_status, verification_status, confidence_score, enrichment_tier, source_relpath, campaign — one row per agent-enriched (Tier-A/B) extracted table from the USSR Robert database. Honest per-table metadata (NOT observation values). | CSV | 21.2 MB | Download |
| Enriched-table catalog (Parquet) | Same columns as the catalog CSV, as Apache Parquet for analytic tooling (DuckDB / pandas / Arrow). | PARQUET | 3.3 MB | Download |
| Year-series concordance families (CSV) | cluster_id, label, description, kind, method, n_members, span, topic — the 175 curated concordance families (sets of linked tables recurring across years that form would-be time-series panels). The construction targets for the next phase. | CSV | 29.1 KB | Download |
Robert-DB ingestion (corpus structure)
| Artifact | Schema / notes | Format | Size | Download |
|---|---|---|---|---|
| Tables by topic (CSV) | topic, n_tables — enriched-table counts per the ratified taxonomy (15 named topic categories plus an (unclassified) bucket = 16 rows). Drives the Charts page. | CSV | 0.5 KB | Download |
| Tables by decade (CSV) | decade, n_tables — enriched-table counts by the start decade of each table's period coverage (parsed). Drives the Charts page. | CSV | 0.2 KB | Download |
| Tables by geography (CSV) | geo_facet, n_tables — enriched-table counts by geographic level (USSR total / union republic / oblast / RSFSR / economic region / imperial Russia / foreign). | CSV | 0.2 KB | Download |
| Tables by two-axis quality (CSV) | obs_status, transcription_status, n_tables — the Robert two-axis quality cross-tab (SDMX obs_status × transcription V/H/L/R/X). | CSV | 0.2 KB | Download |
Corpus inventory (real metadata)
| Artifact | Schema / notes | Format | Size | Download |
|---|---|---|---|---|
| Series-cluster inventory (full, 266 clusters) | series, n_docs, years, span, tables, body, sample — the complete 266-row inventory clustered off the validated HDARP knowledge base. One row per series cluster: normalized cluster slug, document count, distinct-year count, year span, extracted-table count, body-chunk count, and an example document title. Genuine corpus measurements — NOT a constructed time series. | CSV | 28.8 KB | Download |
| Panel-capable clusters (flat CSV) | series_cluster, example_document, n_docs, n_years, span, extracted_tables, body_chunks — the panel-capable clusters (repeated periodic publications across years), a flat CSV view of inventory.json (faithfully serialized). | CSV | 4.9 KB | Download |
| Richest-by-tables clusters (flat CSV) | series_cluster, example_document, n_docs, n_years, span, extracted_tables, body_chunks — clusters ranked by extracted-table count, a flat CSV view of inventory.json. | CSV | 2.4 KB | Download |
Project headlines (real)
| Artifact | Schema / notes | Format | Size | Download |
|---|---|---|---|---|
| Project + inventory totals (CSV) | metric, value, scope — every headline count with an explicit scope tag (full campaign / full KB / clustered subset), so the two real totals (98,748 full-KB tables vs 95,827 clustered) are never conflated. Serialized from headlines.json + inventory.json. | CSV | 0.5 KB | Download |
Anu construction scope (ROADMAP — not yet built)
Scope records, not constructed series. The rows in
these files define the Anu construction targets (cluster id, name, doc
count, span) — the plan, not built data. Every status value is
"scoped (not yet built)".
| Artifact | Schema / notes | Format | Size | Download |
|---|---|---|---|---|
| Flagship clusters A–E (scope CSV) | cluster_id, name, scoped_docs, span, status, note — the 5 flagship Anu construction targets. status is 'scoped (not yet built)' for ALL rows: these are construction-plan records, NOT constructed series. Serialized from headlines.json (anu.clusters). | CSV | 0.9 KB | Download |
Tip: the same inventory powers the interactive charts on the Explore page, and every download here is regenerated by the code shown on the Code page. Provenance and limits: Methodology.