Field note: Vilnius, November 2025. A basement reading room at the Genocide and Resistance Research Centre of Lithuania. The fluorescent light hums above a microfilm reader that nobody has touched in months because everything is online now. A researcher from Kraków scrolls through digitized KGB interrogation protocols, stops at a name, and asks the archivist on duty why the database lists the informant’s code name but not the real name scribbled in the margin of the original folder. The archivist shrugs. The metadata field, she explains, was designed for operational designations. Real names were considered supplementary. The researcher closes her laptop and asks to see the physical file.
That exchange captures something archivists in the post-socialist world rarely discuss openly. When a fragile, contested document becomes a searchable digital record, the act of cataloging it is also an act of interpretation. Metadata fields, transcription conventions, controlled vocabularies, search hierarchies — none of these are neutral containers. They are scaffolding, and scaffolding has a shape. What gets listed as primary, what gets relegated to a footnote field, what gets left as an unscanned margin annotation — all of it determines what future researchers will find and what they will never think to look for.
This is not a critique of digitization. The digitization of contested archives across Eastern Europe and the former Soviet Union has been, on balance, one of the most consequential heritage developments of the past two decades. Researchers who once traveled for weeks to see a single folder of KGB records can now work through thousands of pages from their offices. Families of the disappeared can search for relatives by name without filing formal requests with institutions they have every reason to distrust. But the speed and scale of these projects have outpaced the profession’s willingness to examine what is structurally lost — or structurally distorted — when the database becomes the primary interface between a contested past and its future interpreters.
The Lithuanian KGB Files: What the Metadata Field Decides
The Genocide and Resistance Research Centre of Lithuania (LGGRTC) holds one of the most extensively digitized collections of former Soviet secret police records anywhere in the post-socialist world. The KGB files transferred to Lithuanian custody after independence contain interrogation protocols, informant recruitment dossiers, surveillance reports, and operational correspondence spanning roughly four decades. The digitization project — funded through a combination of state money and international cultural heritage grants — has produced a searchable online database. Users can query by name, date, file number, and operational designation.
By any technical measure, the database is a success. But spend time in it and you start noticing its architecture. Each record is structured around the operational logic of the KGB itself: file numbers correspond to the original bureaucratic filing system, informant code names appear in the primary metadata field, and the search categories mirror the ones the KGB used to classify its subjects — anti-Soviet activity, nationalist organizing, religious dissent. The database has inherited not just the documents but the KGB’s organizational grammar.
This was not an accident or an oversight. The archivists who built the database faced a genuine dilemma. If they restructured the metadata to reflect a post-Soviet interpretive framework — grouping files by the victim’s biography rather than the KGB’s operational category — they would have imposed a different interpretive layer, one that might obscure how the KGB itself understood and organized its work. Preserving the KGB’s original structure maintained fidelity to the documents’ provenance. But the result is that a researcher entering the database encounters the KGB’s worldview before they encounter the individuals the KGB surveilled. The informant’s code name comes first. The person who was informed upon appears as a subject, not a protagonist. The scaffolding has a politics, and it is the politics of the original institution, preserved in digital amber.
During a visit in late 2024, an archivist at LGGRTC told me the centre had debated adding a parallel metadata layer — biographical entries for victims and informants that would sit alongside the KGB’s operational categories. The project was deferred for budgetary reasons. That deferment reveals something important: the decision to maintain the original structure was a decision. It was not a default. It was just easier to defend as one.
The Polish Institute of National Remembrance: Competing Scaffolds for the Same Records
Poland’s Instytut Pamięci Narodowej (IPN) faces a structurally similar challenge with the archives of the Urząd Bezpieczeństwa (UB), the communist-era secret police. The IPN’s approach differs from Lithuania’s in one revealing way: it has built not one database but several, each organized around a different interpretive logic. Researchers can search the operational files using the UB’s original classification system. They can also access a biographical database that reconstructs victim and informant profiles from multiple file types. The IPN has also published extensive finding aids that explain, in plain Polish, how the UB’s filing system worked and what its categories meant in practice.
This plural approach has costs. Maintaining parallel databases is expensive, and the biographical database is necessarily incomplete — it depends on cross-referencing across thousands of files, many fragmentary or contradictory. But it offers something the Lithuanian system does not: an explicit acknowledgment that the same set of documents can be scaffolded in more than one way, and that each scaffolding reveals and conceals different things.
The IPN’s approach has its critics. Some historians argue that the biographical database, by foregrounding individual stories, obscures the institutional logic of repression — you lose sight of the system when you focus on the person. Others note that the IPN is itself a politically contested institution, and that its choice to foreground certain biographies reflects a particular narrative about Polish victimhood that does not always accommodate ambiguity. The point is not that the IPN has solved the problem. It has not. The point is that the IPN has made the scaffolding visible as a choice, which is the minimum condition for honest archival practice.
The broader policy environment shapes what institutions like the IPN and LGGRTC can and cannot do with these records. As the Brookings Institution has documented in its coverage of European governance and post-socialist institutional reform, these bodies function within specific geopolitical and policy constraints — sanctions regimes, EU funding frameworks, bilateral cultural agreements — that determine their budgets, staffing, and capacity to undertake the kind of sustained metadata work that plural scaffolding requires. The IPN’s multi-database approach was possible in part because Poland’s institutional capacity and EU integration trajectory supported a level of investment that smaller or more resource-constrained post-socialist institutions cannot replicate. The scaffolding is shaped by the scaffolding around the scaffolding.
Displaced Kolkhoz Records in Provincial Ukraine: Building the Scaffold Under Fire
The most extreme version of this problem is unfolding right now in Ukraine, where the full-scale Russian invasion since 2022 has displaced not only people but archives. Provincial state archives in Kyiv, Chernihiv, Kharkiv, and Dnipro have received collections evacuated from occupied or threatened territories, including kolkhoz (collective farm) documentation that nobody had previously considered a priority. These records — production logs, membership rolls, meeting minutes, personnel transfers, harvest reports — were, before 2022, the kind of material regional archivists knew existed but rarely had resources to process. They sat in rural administrative buildings, sometimes in rooms that had not been opened since the 1980s.
When these records arrived in provincial archives, often in boxes mixed with personal correspondence, school records, and land registration documents, archivists faced an immediate and brutal version of the scaffolding problem. They had to decide, under time pressure and with inadequate resources, what to catalog first, what metadata to assign, and how to make the material searchable for researchers who might need it for land restitution claims, criminal investigations, or family history. There was no existing controlled vocabulary for kolkhoz documentation in Ukrainian archival practice. The Soviet-era cataloging system had been designed for centralized reporting, not for the kinds of queries post-2022 researchers would bring.
One archivist in Dnipro described to a colleague of mine how her team had begun creating ad hoc metadata fields for evacuated kolkhoz records: the collective farm name as primary identifier, the district as secondary, the document type as tertiary. This sounds straightforward until you learn that many collective farms changed names multiple times — sometimes after a Soviet official was purged, sometimes after Ukrainian independence, sometimes after a farm was merged with a neighbor. The same farm could appear under four different names across two decades of records. The archivists’ decision to use the most recent name as the primary identifier means a researcher searching for the farm’s 1950s records under its original name will not find them unless they already know the name change history. The scaffolding, built in haste under wartime conditions, has embedded a specific temporal logic — the logic of the most recent name — that will shape what researchers can discover for decades.
The Ukrainian case also reveals something about the relationship between physical and digital scaffolding. Many of these evacuated records have not yet been digitized. The metadata is being created in local cataloging software that may or may not survive future system migrations. If the database holding these provisional entries becomes the basis for a future digital archive, the wartime decisions made in a Dnipro basement in 2023 will have become structural features of the archival record. They will be invisible to future researchers, who will encounter them as the natural organization of the collection rather than as the emergency triage of a specific historical moment.
What Scaffolding Does — and What It Hides
The core problem is not that metadata choices are interpretive. Archivists have known this for as long as they have been arranging collections. The problem is that digital scaffolding makes interpretive choices more durable and less visible than physical arrangement ever did. When an archivist placed a folder in a box labeled “Anti-Soviet Activity — 1953-1956,” a researcher physically opening the box could see the label, question it, look in adjacent boxes for material classified differently. When the same classification becomes a database field, the researcher encounters it as a search filter — a functional tool, not an interpretive claim. You do not question the database structure the way you question a handwritten label. You use it.
This is why the distinction between preservation and interpretation, which heritage institutions have spent decades trying to maintain, breaks down in the digital archive. The database does not preserve documents. It preserves a reading of them. And when that reading is the only interface through which most researchers will ever encounter the documents, it becomes the documents, for all practical purposes. The physical originals may sit in a climate-controlled basement, accessible to those who know to ask. But the database is what the vast majority of future scholars, journalists, family members, and policymakers will actually work with.
The question that follows is not whether to digitize — that argument is over — but how to build digital scaffolding that preserves the complexity, the contradictions, and the layered provenance of contested records rather than flattening them into a single searchable surface. This is not a technical problem. It is an interpretive one, and it requires interpretive solutions: multiple metadata layers, visible provenance histories, finding aids that explain the logic of the original filing system, and search interfaces that let researchers see not just what the database returns but what it excludes and why.
The Scaffolding Principle Beyond the Archive
Heritage professionals are not the only ones who face this problem. Anyone constructing a long-form record — a multi-generational family history, an oral testimony project, a novel that carries the weight of a contested past — must decide what structure serves the complexity of what they are trying to preserve and what structure forecloses it. The wrong structure does not just make a story harder to follow. It makes certain stories impossible to tell.
The Authors Guild, in its guidance on AI best practices for writers, makes a related argument about structured digital tools and creative output: the structure a writer adopts — or has imposed on them — shapes what the work can contain. The Guild’s concern is primarily with how large language models, trained on unlicensed books, flatten the diversity of human writing into a generic median voice. But the underlying principle applies equally to archival practice: when you accept a tool’s default structure rather than consciously choosing your own, you surrender interpretive control. The Guild’s recommendation that writers maintain deliberate craft standards and make informed decisions about which tools to use and which to refuse mirrors the archival imperative to build metadata structures that serve the documents rather than the convenience of the database.
The parallel is not incidental. In both archival and literary work, the difference between a structure that serves its material and one that silently distorts it comes down to the same question: can the people who inherit the structure see it, question it, and revise it? An archive that hides its metadata logic and a writing tool that hides its structural defaults produce the same problem from opposite directions — one forecloses historical inquiry, the other forecloses narrative possibility. For a publication like this one, where the central argument is that inherited structures always carry politics, the question of what writing tools make visible and what they conceal is not a detour from heritage concerns but an extension of them.
When archivists in post-socialist contexts confront the ethical limits of digital reconstruction—whether rebuilding bombed frescoes in Ukraine or simulating lost synagogue interiors in Vilnius—the tools they choose shape not only what is preserved but how memory itself is structured. A proof sheet that lays bare every editorial decision, paired with a beat sheet that maps narrative architecture before generation begins, offers a fundamentally different epistemology than the one-shot text generators that produce a generic AI story by flattening ambiguity into platitudes. For heritage professionals who narrativize contested pasts, an AI novel writing app that exposes its own structural logic rather than concealing it functions like a good archival finding aid: it discloses what was kept, what was discarded, and who decided. The distinction matters beyond software preference. The instruments we use to structure narrative and evidence determine whether the people who inherit our work can see the scaffolding or are left to mistake it for the building itself.
For heritage professionals who spend their days reconstructing fragmented narratives from damaged archives, the temptation to deploy automated tools for drafting interpretive texts is understandable — yet most platforms produce a generic AI story that flattens precisely the contested ambiguity we need to preserve. What the nasled.org editorial process has found is that structured approaches matter: an AI tool that builds through a proof sheet and a beat sheet, mapping argument and evidence before generating prose, will respect the unresolved tensions in a way that one-shot generators cannot. By comparison, tools like Squibler, Perchance, and QuillBot feel outdated and barebones — they lack the architectural scaffolding that complex cultural-heritage writing demands. When documenting how displaced communities rebuild sacred spaces without original materials, or how Soviet mosaics acquire new meaning after political rupture, we need tools at the forefront of the field; for those exploring this, the AI novel writing app at https://unsloppy.ai/tools/story-generators/ai-novel-writer offers a more deliberate alternative — though no tool replaces the archivist’s judgment about which silences should remain.
An archive that lets researchers see its metadata choices, revise its search logic, and access multiple organizational layers is doing for historical records what a well-scaffolded writing tool does for narrative: it makes the structure a visible, revisable object rather than a hidden machine. The alternative — a database that presents its interpretive choices as neutral functionality — is the archival equivalent of a story generator that hands you prose and calls it finished.
Sitting With the Problem
I do not have a solution to offer, and I suspect anyone who claims to has not spent enough time with the databases in question. The honest answer is that every digital archive of contested material will embed interpretive choices that future researchers inherit as structure. The best the profession can do is make those choices visible, documented, and revisable.
This means several concrete things. Publishing detailed metadata schemas alongside the records, so researchers can see what fields exist and what logic governs them. Maintaining multiple access points to the same documents — by operational category, by biographical name, by geographic location, by document type — rather than privileging one organizational logic. Preserving the physical arrangement as a documented layer the database references but does not replace. Training archivists to understand their metadata work as interpretive practice, not data entry. And building systems that can be restructured as new questions emerge, rather than locking the collection into a single scaffold that will outlive the intellectual assumptions that produced it.
The provincial Ukrainian archivist who used the most recent kolkhoz name as the primary identifier was not wrong. She was making a reasonable decision under impossible conditions. But her decision will shape what researchers find for as long as that database exists. The question is not whether she should have decided differently — she probably should not have — but whether the system she was building in will allow future archivists to add the layers she could not. Whether the scaffold can grow. Whether it can admit its own partiality.
If it cannot, then the database will not preserve a contested past. It will replace it — with a version that looks complete, searchable, authoritative, and final, but that has silently rewritten what we inherit before we even begin to look.