Skip to content
Heterodata An Arcanum Research project Vlad
Vlad

Corpus Explorer

The Robert-database ingestion of the Imperial Russian & Soviet statistical corpus, browsable one row per enriched extracted table. Each row is the honest metadata the Robert Database Framework recovered: the source table's title (translated to English where that pass has run, with the Russian original always preserved), units, geography, period, a topic from the ratified taxonomy (15 named categories plus an (unclassified) bucket for tables not yet assigned a topic), and the two-axis quality grade. These are genuine per-table measurements of the corpus — no values are fabricated.

Enrichment depth is uneven, and the table shows it. Of the 55,749 rows, 25,677 have an English title, a taxonomy topic, a geography facet and a parsed period. The rest were metadata-enriched later and are still awaiting the translation and classification passes: they display their Russian title and their quality grade, and their topic / period cells are blank. Blank means not recovered yet — never zero, never a guess. Filter by any topic (or by year) to browse only the fully classified subset.

55,749
enriched tables (catalog)
25,677
topic-classified
27,936
with recovered units
16
topic buckets (15 named + unclassified)
1765–2013
period span (data years)
175
year-series families

A "table" here is a single extracted statistical table from a source publication. Tier-A/B (agent-enriched) tables are shown; mechanical-only (Tier-C) tables are excluded so every row carries honest recovered metadata. The period span above is the range of data years covered by the tables (each table's earliest period-coverage year) — which can predate or postdate a source's publication date, so it extends beyond the 1804–2008 publication span of the source documents. See the Charts for the corpus structure and Methodology for the quality model and provenance.

Browse the enriched-table catalog

Table (English — Russian original)TopicGeography PeriodUnitsSizeQuality

Year-series families

Curated concordances — sets of linked tables that recur across years and together form a would-be time-series panel. These are the highest-value targets for the next (harmonized-series) construction phase. Click a family to see its member tables across the years.

FamilyTopicSpanMember tablesKind

A validated knowledge base of Imperial Russian & Soviet official statistics, extracted via HDARP. Russian source-document titles are translated into English (original preserved). Figures reconciled 2026-06-07. All numbers trace to canonical project artifacts — nothing on this site is fabricated; harmonized time-series panels are scoped but not yet constructed (see the Roadmap). Data is reconstructed for research and education; defer to the original archival sources.