Vlad
The Soviet statistical corpus, read from the state's own statistical returns — a validated knowledge base of Imperial Russian and Soviet official statistics (1804–2008), extracted from primary sources.
The catalog is live and real — 55,749 enriched tables you can browse, filter and download, of which 25,677 carry a full English-translated, taxonomy-classified record. The Data button ships catalog metadata (title, units, geography, period and two-axis quality per table) — not observation values. Harmonized year×value economic series are scoped, not yet built (see the Roadmap); none are invented here.
Named for Vladimir Lenin — the great political economist — and Vladimir Voytinsky, the Soviet economist and statistician. Both worked from the same raw material this site collects: the official statistical returns of the Russian state. One turned them into an argument about how the economy worked; the other spent a career making sure the figures existed at all.
Vladimir Lenin
His The Development of Capitalism in Russia (1899) is an empirical book before it is a political one: hundreds of pages of zemstvo and factory statistics, industry by industry and province by province, marshalled to show that Russian agriculture and handicraft were already capitalist. Whatever one makes of the conclusions, the method was to read the state's own numbers more carefully than the state did.
Vladimir Voytinsky
Compiler of Die Welt in Zahlen (“The World in Figures”, seven volumes, 1925–1928), one of the first attempts to put the world economy — output, prices, labour, trade — into a single comparable statistical frame. He later worked on employment policy in Weimar Germany and, in exile, at the U.S. Social Security Administration. His discipline was the unglamorous one this corpus depends on: get the series right, and say where it came from.
This site serves the Robert-database ingestion of the corpus: a browsable, table-level catalog of the extracted Imperial Russian and Soviet statistics. Every figure is a real measurement of the knowledge base, and every Russian source-document title is translated into English (with the Russian original preserved). Browse the enriched-table catalog (units, geography, period, taxonomy and two-axis quality per table) and the corpus-structure charts. Harmonized year×value economic series are scoped but not yet constructed — see the Roadmap; none are invented here.
What this is
Vlad is a large-scale extraction of primary statistical publications spanning the Russian Empire and the USSR. Source PDFs — budget schedules (rospis), ministry yearbooks, foreign-trade compendia, national-economy (narodnoe khozyaystvo) statistical handbooks, census volumes and more — were turned into structured, machine-readable text and tables. The displayed labels are translated into English, with each Russian original kept alongside.
Method. HDARP (Heterodox Document Archive Reconstruction Pipeline) — AI-agent document extraction; full 4-type (body text + tables + equations + figures) per chunk; pre-1918 Russian orthography (ъ ѣ і ѳ) preserved verbatim; independently model-validated.
Status. Of 1,964 cataloged documents, 1,246 are VERIFIED, with 809 carrying substantial content (41,800 body chunks, 98,748 tables). 289 were de-duplicated, 57 quarantined as unreadable/stub sources, and 3 documented as permanent losses.
The corpus, clustered
The validated knowledge base groups into 266 series clusters (797 documents, 95,827 extracted tables). Of these, 98 are panel-capable — repeated periodic publications across multiple years that can become time-series panels.
The corpus then passed through the Robert Database Framework, which recovered honest per-table metadata (English/Russian title, units, geography, period, a topic from a ratified taxonomy of 15 named categories plus an (unclassified) bucket, and a two-axis quality grade) and curated year-series families of linked tables. The result is browsable in the Explorer, visualized in the Charts, and the construction plan for harmonized series is in the Roadmap.
Reconciling the two views: the full campaign catalogs 95,827 clustered tables across 266 series clusters; the enriched, per-table browsable catalog is the 55,749-table Tier-A/B subset that carries honest recovered metadata (the mechanical-only remainder is excluded). The two totals are the same corpus at different enrichment depths, never conflated. Depth varies within the catalog too: 25,677 of those tables have been through the English-translation and taxonomy-classification passes; the other 30,072 carry their Russian title, geography and quality grade but no English title, topic or period yet — those fields are left empty rather than guessed.