Skip to content
Heterodata An Arcanum Research project Vlad
Vlad

Data

The real corpus-inventory metadata behind this showcase, downloadable with correct content-types. 12 files: the full 266-cluster series inventory, panel-capability counts, and the headline campaign totals.

Real corpus metadata — no constructed series. Roadmap Every file here is a genuine measurement of the extracted knowledge base (per-cluster document counts, year spans, extracted-table counts), served verbatim or as a faithful CSV serialization of a real JSON structure. There are no harmonized economic time-series files, because none are built yet — the Anu data-construction pipeline is scoped, not constructed. The Cluster A fiscal pilot (1856–1990) is first in line; see Methodology for the full plan and limits. The 4.4 GB knowledge base and ~33 GB of outputs are not shipped.

Download everything

One zip of all 12 corpus-inventory artifacts plus a provenance README that states exactly what is real and what is not yet built.

Download vlad_corpus_inventory.zip Site repro bundle (.zip)

Robert-DB ingestion (enriched-table catalog)

ArtifactSchema / notesFormatSizeDownload
Enriched-table catalog (full, ~32k tables) table_uid, doc_id, title_en, title_raw, topic, geo_facet, geography, units, period_coverage, start_year, n_rows, n_cols, quality_grade, obs_status, transcription_status, verification_status, confidence_score, enrichment_tier, source_relpath, campaign — one row per agent-enriched (Tier-A/B) extracted table from the USSR Robert database. Honest per-table metadata (NOT observation values). CSV 21.2 MB Download
Enriched-table catalog (Parquet) Same columns as the catalog CSV, as Apache Parquet for analytic tooling (DuckDB / pandas / Arrow). PARQUET 3.3 MB Download
Year-series concordance families (CSV) cluster_id, label, description, kind, method, n_members, span, topic — the 175 curated concordance families (sets of linked tables recurring across years that form would-be time-series panels). The construction targets for the next phase. CSV 29.1 KB Download

Robert-DB ingestion (corpus structure)

ArtifactSchema / notesFormatSizeDownload
Tables by topic (CSV) topic, n_tables — enriched-table counts per the ratified taxonomy (15 named topic categories plus an (unclassified) bucket = 16 rows). Drives the Charts page. CSV 0.5 KB Download
Tables by decade (CSV) decade, n_tables — enriched-table counts by the start decade of each table's period coverage (parsed). Drives the Charts page. CSV 0.2 KB Download
Tables by geography (CSV) geo_facet, n_tables — enriched-table counts by geographic level (USSR total / union republic / oblast / RSFSR / economic region / imperial Russia / foreign). CSV 0.2 KB Download
Tables by two-axis quality (CSV) obs_status, transcription_status, n_tables — the Robert two-axis quality cross-tab (SDMX obs_status × transcription V/H/L/R/X). CSV 0.2 KB Download

Corpus inventory (real metadata)

ArtifactSchema / notesFormatSizeDownload
Series-cluster inventory (full, 266 clusters) series, n_docs, years, span, tables, body, sample — the complete 266-row inventory clustered off the validated HDARP knowledge base. One row per series cluster: normalized cluster slug, document count, distinct-year count, year span, extracted-table count, body-chunk count, and an example document title. Genuine corpus measurements — NOT a constructed time series. CSV 28.8 KB Download
Panel-capable clusters (flat CSV) series_cluster, example_document, n_docs, n_years, span, extracted_tables, body_chunks — the panel-capable clusters (repeated periodic publications across years), a flat CSV view of inventory.json (faithfully serialized). CSV 4.9 KB Download
Richest-by-tables clusters (flat CSV) series_cluster, example_document, n_docs, n_years, span, extracted_tables, body_chunks — clusters ranked by extracted-table count, a flat CSV view of inventory.json. CSV 2.4 KB Download

Project headlines (real)

ArtifactSchema / notesFormatSizeDownload
Project + inventory totals (CSV) metric, value, scope — every headline count with an explicit scope tag (full campaign / full KB / clustered subset), so the two real totals (98,748 full-KB tables vs 95,827 clustered) are never conflated. Serialized from headlines.json + inventory.json. CSV 0.5 KB Download

Anu construction scope (ROADMAP — not yet built)

Scope records, not constructed series. The rows in these files define the Anu construction targets (cluster id, name, doc count, span) — the plan, not built data. Every status value is "scoped (not yet built)".

ArtifactSchema / notesFormatSizeDownload
Flagship clusters A–E (scope CSV) cluster_id, name, scoped_docs, span, status, note — the 5 flagship Anu construction targets. status is 'scoped (not yet built)' for ALL rows: these are construction-plan records, NOT constructed series. Serialized from headlines.json (anu.clusters). CSV 0.9 KB Download

Tip: the same inventory powers the interactive charts on the Explore page, and every download here is regenerated by the code shown on the Code page. Provenance and limits: Methodology.

A validated knowledge base of Imperial Russian & Soviet official statistics, extracted via HDARP. Russian source-document titles are translated into English (original preserved). Figures reconciled 2026-06-07. All numbers trace to canonical project artifacts — nothing on this site is fabricated; harmonized time-series panels are scoped but not yet constructed (see the Roadmap). Data is reconstructed for research and education; defer to the original archival sources.