Skip to content
Heterodata An Arcanum Research project Vlad
Vlad

Methodology & provenance

Every number, chart and download on this site is a real measurement of the extracted corpus, transcribed from a canonical project artifact — nothing is fabricated. This page is the Data Provenance Record (DPR): sources · construction · units · limits · last-updated · citations.

Offline Real corpus metadata Series: roadmap (not built)

Read this first — what is real, and what is not. The corpus inventory is REAL: 1,964 documents cataloged, 1,246 Opus-validated, grouped into 266 series clusters with genuine per-cluster document, year-span and extracted-table counts. The harmonized economic series are NOT built: the Anu data-construction pipeline that would turn panel-capable clusters into year × category × value panels is scoped, not constructed. What is now served is the Robert-database ingestion: a browsable table-level catalog of the extracted statistics (title, units, geography, period, taxonomy and quality per table). No harmonized year×value time series are shown or shipped anywhere on this site.

Data Provenance Record

Sources
Headline campaign counts — Projects/USSR/CAMPAIGN_STATUS.md (canonical reconciled status surface, "Last reconciled 2026-06-07"). The series-cluster inventory — ANU_SERIES_INVENTORY.csv (266 clusters, generated by integration/cluster_series.py off the validated knowledge base, summarized into inventory.json by build_data.py). The construction roadmap — the project's Anu pipeline scoping document. The underlying primary sources are budget schedules (rospis), ministry yearbooks, foreign-trade compendia, national-economy (narodnoe khozyaystvo) handbooks, census volumes and military-ministry reports, 1804–2008.
Translation & vetting of displayed labels
Everything the site displays is vetted, organized, and translated into English. Each cluster is labeled by an example source document; the Russian titles are rendered in English as the primary display string, with the Russian original preserved verbatim and shown in parentheses (and as a hover tooltip in Explore). The translations are a curated, hand-authored map (translations.json) consumed by build_data.py at build time, so the transformation is reproducible and auditable. During translation every displayed label was vetted; empty/garbled extraction artifacts are dropped and the drop is logged in the build report. (In the current build: all displayed Russian labels translated, none dropped.) Translations are aids to navigation — authoritative figures should defer to the original archival sources.
Construction (what IS built — the inventory)
The corpus was built with HDARP (Heterodox Document Archive Reconstruction Pipeline — an AI-agent document-extraction method): every chunk gets full 4-type extraction (body text + tables + equations + figures); pre-1918 Russian orthography (ъ ѣ і ѳ) is preserved verbatim; output is Opus-validated. That produced 98,748 extracted tables and 41,800 body-text chunks across a 4.4 GB knowledge base in 850 processing batches. The validated KB was then clustered into 266 series clusters (98 of them panel-capable — repeated periodic publications spanning ≥ 2 distinct years). The Explore charts and the Data downloads are computed deterministically from these measurements — see Code.
Robert-database ingestion (the table-level catalog you browse & chart)
Real The validated knowledge base was ingested by the Robert Database Framework v1.0 into a canonical per-project SQLite database (Projects/USSR/Technical/RobertDB/robertdb.sqlite). That run recovered honest, per-table metadata for tens of thousands of extracted tables and is what the Explore and Charts pages serve. The site ships a compact public slice (vlad_slice.sqlite, baked by build_slice.py) carrying one row per agent-enriched (Tier-A/B) table with: the table's English title + Russian original, units, geography, period coverage (+ parsed start year), a topic from the ratified 16-category USSR statistical taxonomy, a geographic-level facet, and the Robert two-axis quality model — the SDMX obs_status axis crossed with a transcription axis (V verified · H human-grade · L LLM-read · R needs repair · X failed). It also ships the 175 curated concordance families (sets of linked tables recurring across years — the would-be panels). Tier-C mechanical-only tables are excluded so every browsable row carries recovered metadata. No observation values are shipped — this is the honest catalog of what was extracted, not the cell data; the 4.4 GB knowledge base stays in the project repository. Every value is copied or aggregated verbatim from the Robert database; nothing is fabricated.
Construction (what is SCOPED — the harmonized series)
Roadmap The next phase applies the Anu data-construction framework to turn the knowledge base into harmonized, analyzable time-series panels (a single source of truth per series, with full provenance). Status: scoped (not yet built). 108 panel-capable series and 158 single-doc cross-sectional sources have been identified, but no panels are constructed yet. (The two panel counts measure different things: the 98 above is the conservative tally surfaced in Explore — clusters in the 266-cluster auto-inventory with a year span and ≥ 2 distinct years; the 108 here is the wider scoping count from the Anu pipeline plan, which also picks up panel candidates the automatic clustering does not group. Neither is a constructed series.) Build order (value × readiness):
  • Cluster A — Imperial → Soviet Fiscal Panel (PILOT) (~60 docs · 1856–1990): Annual state revenue-by-source + expenditure-by-ministry across two regimes; deep-validated 1856–1887 core.
  • Cluster B — Foreign Trade Panel (~58 docs · 1918–1990): Imports/exports by country & commodity; ~12k tables, biggest single Soviet panel.
  • Cluster C — National-Economy Yearbooks (narodnoe khozyaystvo) (~60 docs · 1922–1990): Core macro indicators by union republic-year; a Soviet regional panel.
  • Cluster D — Military Ministry Reports (~54 docs · 1859–1911): Annual war-ministry reports; 8,633 tables — largest imperial table trove.
  • Cluster E — Sector Panels (prison / agriculture / industry / census / transport) (1851–1988): Penal statistics, agriculture, industry, the 1897 census + revisions, MPS railway/transport bulletins.

The Cluster A fiscal pilot (1856–1990) goes first — a self-contained, deep-validated core that exercises every Anu skill and yields the currency / category / orthography harmonization machinery reused everywhere. See the standalone construction roadmap for the per-series recipe and phasing.

Units & coverage (of the inventory)
The inventory's units are counts and year spans, not economic quantities (no rubles, tonnes or indices are constructed yet). Each cluster row carries: n_docs (documents), years (distinct year count), span (first–last year), tables (extracted-table count) and body (body-chunk count). Temporal coverage of the corpus spans 1804–2008; the panel-capable clusters concentrate in the imperial fiscal/military era (1856–1911) and the Soviet statistical era (1918–1990).
Limits
  • No harmonized series exist. This is the defining limit: every "panel" referenced here is a candidate (a cluster of repeated publications), not a constructed time series. Nothing on this site is a year×value series.
  • Two real totals, two scopes. The full-KB totals (98,748 tables · 41,800 body chunks) and the clustered-subset totals (95,827 tables · 41,384 body chunks across 797 docs) measure different scopes — both real, both labeled. The clustered subset excludes documents that did not group into a periodic series.
  • Not shipped: the 4.4 GB knowledge base and ~33 GB of outputs. This is a compact metadata showcase (each shipped file < ~130 KB).
  • Status buckets are honest: of 1,964 cataloged docs, 289 were de-duplicated, 57 quarantined as unreadable/stub, 30 deprecated, 305 skipped and 3 documented as permanent losses (source re-acquisition pending).
Citations
Primary sources (English title — Russian original): General State Schedule of Revenues and Expenditures (Общая государственная роспись доходов и расходов, Imperial budget schedule, 1856–1908); State Budget of the USSR (Государственный бюджет СССР, 1918–1990); Most Loyal Report on the Activities of the War Ministry (Всеподданнейший отчёт о действиях военного министерства, 1859–1911); National Economy of the USSR (Народное хозяйство СССР, yearbooks, 1922–1990); Foreign Trade of the USSR (Внешняя торговля СССР, compendia, 1918–1990). Method: the HDARP extraction standard and the Anu data-construction framework, Arcanum Research.

Last updated — DPR 2026-06-21 · Robert-DB slice baked from robertdb.sqlite (Robert Database Framework; DB created 2026-06-11) · headlines reconciled 2026-06-07 · scope doc ANU_PIPELINE_SCOPE.md (2026-06-06).

Surface → source map

SurfaceHeadlineCanonical source artifact
Overview 1,964 docs · 1,246 verified CAMPAIGN_STATUS.md
Explore Enriched-table catalog + year-series families robertdb.sqlite → vlad_slice.sqlite (build_slice.py)
Charts Tables by topic / decade / geography / quality vlad_slice.sqlite agg_* (Robert DB)
Data Catalog + aggregates + 266-cluster inventory parquet/*.csv · inventory.json · headlines.json
Roadmap 5 flagship clusters (scoped, not built) Anu pipeline scope doc

Verification: every numeric field on this site is a direct value-comparison against the named JSON/CSV source artifact; the cluster inventory is regenerated from the canonical CSV by build_data.py.

A validated knowledge base of Imperial Russian & Soviet official statistics, extracted via HDARP. Russian source-document titles are translated into English (original preserved). Figures reconciled 2026-06-07. All numbers trace to canonical project artifacts — nothing on this site is fabricated; harmonized time-series panels are scoped but not yet constructed (see the Roadmap). Data is reconstructed for research and education; defer to the original archival sources.