Kronika builds and maintains its own technology for preserving media that may be blocked, deleted, altered, or lost. Our infrastructure helps recover vulnerable material, preserve original files and versions, transform difficult formats into structured data, and keep archives searchable and usable across different media environments and preservation needs

Technology built for endangered media

Our tools

Tool 1

Archivarius

The crawl-and-recovery engine for endangered media

Archivarius is Kronika's archiving engine. It discovers, downloads and structures content from media websites using pages, feeds, sitemaps, browser rendering, and historical sources preserved by Wayback Machine in the Internet Archive.

It is designed not only to copy what remains visible today, but to reconstruct what has already disappeared. Archivarius can recover URLs no longer available through the original site or search engines, preserve links, images and embeds alongside text, and adapt its retrieval methods when sites are difficult to access. The system runs on its own: when a site blocks it, it switches retrieval methods, it notices and restarts itself.

For media facing censorship, deletion, or technical collapse, this ability to recover both the visible and the lost record is central to Kronika's work.

Tool 2

PaperAI

Turning print collections into searchable archives

Some of the most important records were never born digital. PaperAI transforms scanned newspapers, magazines, zines, and other print collections into structured, searchable digital material.

It identifies pages and articles, separates headlines, authors and body text, reconstructs stories that continue across multiple pages, and gives each article its own addressable record. OCR makes the text searchable while preserving access to the original scanned page.

PaperAI allows large print collections to become part of the same searchable archival environment as web, video, and audio sources.

Tool 3

Newsloom

Monitoring vulnerable information in real time

Journalists, researchers, and civil-society organizations often spend hours monitoring websites, feeds, social media, and other sources — many of which may be blocked, geo-restricted, or deleted without warning.

Newsloom brings these sources into one monitoring environment. It continuously collects new material, preserves content even when the original disappears, and can use AI-assisted tools for translation, summarization, categorization, and entity extraction.

AI features are optional: users can switch them off and work with the raw feed.


Note: Newsloom is currently in alpha.

Kronika's public archives share a common infrastructure rather than operating as isolated websites.

Original files are never overwritten. New captures and corrections are stored alongside earlier versions, preserving each record's history. Every item receives a permanent identifier and metadata describing when and how it was collected.

The infrastructure provides multilingual full-text and transcript search, flexible filtering, and links outlets and authors across collections to track journalism across publications and formats.

This shared layer is what allows Kronika to build new archives without starting from zero each time.

The infrastructure behind our archives

Quality, security & resilience

Preservation is only useful if the material can be trusted.

We do not rely on automated processing alone. Before web, print, video, or audio material is made publicly available, it goes through extensive human quality checks to verify parsing, article structure, attribution, metadata, and completeness. We continuously improve our extraction and parsing algorithms as new formats and problems appear.

Our infrastructure uses separate staging and production environments, automated deployment, daily backups, vulnerability scanning, and current actively maintained versions of core software. We follow green-stack principles to reduce unnecessary resource use, and a full disaster-recovery pipeline is in development.

Security also shapes how we design for users. We collect as little personal data as possible, do not require registration to read public archives, and treat search activity and archived content as potentially sensitive.

Data access:

our archives

are built to be used

Public collections can be searched without registration. For larger research or journalistic projects, we can provide structured exports, including full-corpus datasets and selections filtered by outlet, period, or topic.

Programmatic API access is in development.

For dataset or access requests, contact us and tell us what you are working on and what kind of data you need.

We are Kronika, a civic technology project dedicated to preserving independent journalism under threat

Subscribe to the newsletter

© 2026 Kronika. All rights reserved