About Vlad
A validated knowledge base of Imperial Russian and Soviet official statistics, extracted from primary sources (1804–2008) and read from the state's own statistical returns.
Source: Projects/USSR/CAMPAIGN_STATUS.md (reconciled 2026-06-07).
The name
Named for Vladimir Lenin — the great political economist — and Vladimir Voytinsky, the Soviet economist and statistician. Lenin (1870–1924) built The Development of Capitalism in Russia (1899) out of the empire's own zemstvo and factory returns, industry by industry and province by province — an argument made of statistics. Voytinsky (1885–1960) compiled Die Welt in Zahlen ("The World in Figures", seven volumes, 1925–1928), an early attempt to put the world economy into one comparable statistical frame; he later worked on employment policy in Weimar Germany and, in exile, at the U.S. Social Security Administration. Between them they mark the two halves of the job this corpus serves: the argument, and the arithmetic it has to rest on.
What this is
Vlad is a large-scale extraction of primary statistical publications spanning the Russian Empire and the USSR (1804–2008). Source PDFs — budget schedules (rospis), ministry yearbooks, foreign-trade compendia, national-economy (narodnoe khozyaystvo) statistical handbooks, census volumes, military-ministry reports, prison and railway statistics — were turned into structured, machine-readable text and tables. Displayed source-document titles are translated into English, with the Russian original preserved alongside.
Method. HDARP (Heterodox Document Archive Reconstruction Pipeline) — AI-agent document extraction; full 4-type (body text + tables + equations + figures) per chunk; pre-1918 Russian orthography (ъ ѣ і ѳ) preserved verbatim; independently model-validated.
The honest-overview framing
This site is deliberately an honest metadata overview, not a data product. The corpus inventory is real and measurable — 1,964 documents cataloged, 1,246 Opus-validated, 98,748 extracted tables grouped into 266 series clusters. The harmonized economic series that would let you chart Soviet revenue or trade over time are scoped but not yet constructed. Rather than fabricate them, the site shows exactly what exists and labels the rest as roadmap. Nothing here is invented.
Why a Soviet statistical corpus?
Imperial Russian and Soviet official statistics are voluminous, internally rich, and notoriously hard to use: a century and a half of publications across two regimes, multiple currency reforms (silver / credit / gold ruble after the 1897 reform; Soviet ruble redenominations in 1922, 1947, 1961), a Julian→ Gregorian calendar break (Feb 1918), shifting administrative units (governorates → oblasts/krais → union republics), pre-1918 orthography, and category drift as ministries and revenue articles rename and split year to year. The project's wager is that careful extraction plus disciplined harmonization can turn this into a coherent, provenance-tracked panel — a 130-year fiscal continuity across two regimes being the flagship target. The extraction (this knowledge base) is done; the harmonization (the construction roadmap) is the next phase.
For research and education
This data is reconstructed from historical sources for research and educational use. Despite Opus validation, machine extraction of century-old multilingual scans can carry transcription error. For any authoritative or citable figure, defer to the original archival publications; the example-document titles and the Methodology page name the underlying sources.
The Arcanum Research ecosystem
Vlad is one site in a family of open research tools. Every site shares the same design system and links both anchors below.
heterodata.org
→The Arcanum Research hub — the front door to all the data tools.
heterodata.orgnickanderson.us
→Author / umbrella apex — the person behind the projects.
nickanderson.usvolcker.org
→A sibling HDARP→Anu project: comparative empirical testing of theories of banking over 161 years of US data.
volcker.orgOffline Real corpus metadata Series: roadmap (not built)