Technology built for endangered media

Technology built for endangered media

Kronika builds and maintains its own technology for preserving media that may be blocked, deleted, altered, or lost. Our infrastructure helps recover vulnerable material, preserve original files and versions, transform difficult formats into structured data, and keep archives searchable and usable across different media environments and preservation needs.

Kronika builds and maintains its own technology for preserving media that may be blocked, deleted, altered, or lost. Our infrastructure helps recover vulnerable material, preserve original files and versions, transform difficult formats into structured data, and keep archives searchable and usable across different media environments and preservation needs.

Menu

Our tools

Our tools

Our tools

Tool 1

Archivarius

The crawl-and-recovery engine for endangered media

Archivarius is Kronika's archiving engine. It discovers, downloads and structures content from media websites using pages, feeds, sitemaps, browser rendering, and historical sources preserved by Wayback Machine in the Internet Archive.

It is designed not only to copy what remains visible today, but to reconstruct what has already disappeared. Archivarius can recover URLs no longer available through the original site or search engines, preserve links, images and embeds alongside text, and adapt its retrieval methods when sites are difficult to access. The system runs on its own: when a site blocks it, it switches retrieval methods, it notices and restarts itself.

For media facing censorship, deletion, or technical collapse, this ability to recover both the visible and the lost record is central to Kronika's work.

Tool 2

PaperAI

Turning print collections into searchable archives

Some of the most important records were never born digital. PaperAI transforms scanned newspapers, magazines, zines, and other print collections into structured, searchable digital material.

It identifies pages and articles, separates headlines, authors and body text, reconstructs stories that continue across multiple pages, and gives each article its own addressable record. OCR makes the text searchable while preserving access to the original scanned page.

PaperAI allows large print collections to become part of the same searchable archival environment as web, video, and audio sources.

Tool 3

Newsloom

Monitoring vulnerable information in real time

Journalists, researchers, and civil-society organizations often spend hours monitoring websites, feeds, social media, and other sources — many of which may be blocked, geo-restricted, or deleted without warning.

Newsloom brings these sources into one monitoring environment. It continuously collects new material, preserves content even when the original disappears, and can use AI-assisted tools for translation, summarization, categorization, and entity extraction.

AI features are optional: users can switch them off and work with the raw feed.


Note: Newsloom is currently in alpha.

Tool 1

Archivarius

The crawl-and-recovery engine for endangered media

Archivarius is Kronika's archiving engine. It discovers, downloads and structures content from media websites using pages, feeds, sitemaps, browser rendering, and historical sources preserved by Wayback Machine in the Internet Archive.

It is designed not only to copy what remains visible today, but to reconstruct what has already disappeared. Archivarius can recover URLs no longer available through the original site or search engines, preserve links, images and embeds alongside text, and adapt its retrieval methods when sites are difficult to access. The system runs on its own: when a site blocks it, it switches retrieval methods, it notices and restarts itself.

For media facing censorship, deletion, or technical collapse, this ability to recover both the visible and the lost record is central to Kronika's work.

Tool 2

PaperAI

Turning print collections into searchable archives

Some of the most important records were never born digital. PaperAI transforms scanned newspapers, magazines, zines, and other print collections into structured, searchable digital material.

It identifies pages and articles, separates headlines, authors and body text, reconstructs stories that continue across multiple pages, and gives each article its own addressable record. OCR makes the text searchable while preserving access to the original scanned page.

PaperAI allows large print collections to become part of the same searchable archival environment as web, video, and audio sources.

Tool 3

Newsloom

Monitoring vulnerable information in real time

Journalists, researchers, and civil-society organizations often spend hours monitoring websites, feeds, social media, and other sources — many of which may be blocked, geo-restricted, or deleted without warning.

Newsloom brings these sources into one monitoring environment. It continuously collects new material, preserves content even when the original disappears, and can use AI-assisted tools for translation, summarization, categorization, and entity extraction.

AI features are optional: users can switch them off and work with the raw feed.


Note: Newsloom is currently in alpha.

Tool 1

Archivarius

The crawl-and-recovery engine for endangered media

Archivarius is Kronika's archiving engine. It discovers, downloads and structures content from media websites using pages, feeds, sitemaps, browser rendering, and historical sources preserved by Wayback Machine in the Internet Archive.

It is designed not only to copy what remains visible today, but to reconstruct what has already disappeared. Archivarius can recover URLs no longer available through the original site or search engines, preserve links, images and embeds alongside text, and adapt its retrieval methods when sites are difficult to access. The system runs on its own: when a site blocks it, it switches retrieval methods, it notices and restarts itself.

For media facing censorship, deletion, or technical collapse, this ability to recover both the visible and the lost record is central to Kronika's work.

Tool 2

PaperAI

Turning print collections into searchable archives

Some of the most important records were never born digital. PaperAI transforms scanned newspapers, magazines, zines, and other print collections into structured, searchable digital material.

It identifies pages and articles, separates headlines, authors and body text, reconstructs stories that continue across multiple pages, and gives each article its own addressable record. OCR makes the text searchable while preserving access to the original scanned page.

PaperAI allows large print collections to become part of the same searchable archival environment as web, video, and audio sources.

Tool 3

Newsloom

Monitoring vulnerable information in real time

Journalists, researchers, and civil-society organizations often spend hours monitoring websites, feeds, social media, and other sources — many of which may be blocked, geo-restricted, or deleted without warning.

Newsloom brings these sources into one monitoring environment. It continuously collects new material, preserves content even when the original disappears, and can use AI-assisted tools for translation, summarization, categorization, and entity extraction.

AI features are optional: users can switch them off and work with the raw feed.


Note: Newsloom is currently in alpha.

This shared layer is what allows Kronika to build new archives without starting from zero each time.

This shared layer is what allows Kronika to build new archives without starting from zero each time.

The infrastructure provides multilingual full-text and transcript search, flexible filtering, and links outlets and authors across collections to track journalism across publications and formats.

The infrastructure provides multilingual full-text and transcript search, flexible filtering, and links outlets and authors across collections to track journalism across publications and formats.

Original files are never overwritten. New captures and corrections are stored alongside earlier versions, preserving each record's history. Every item receives a permanent identifier and metadata describing when and how it was collected.

Original files are never overwritten. New captures and corrections are stored alongside earlier versions, preserving each record's history. Every item receives a permanent identifier and metadata describing when and how it was collected.

Kronika's public archives share a common infrastructure rather than operating as isolated websites.

Kronika's public archives share a common infrastructure rather than operating as isolated websites.

The infrastructure behind our archives

The infrastructure behind our archives

The infrastructure behind our archives

Quality, security & resilience

Quality, security & resilience

Preservation is only useful

if the material can be trusted.

Preservation is only useful

if the material can be trusted.

We do not rely on automated processing alone. Before web, print, video, or audio material is made publicly available, it goes through extensive human quality checks to verify parsing, article structure, attribution, metadata, and completeness. We continuously improve our extraction and parsing algorithms as new formats and problems appear.

We do not rely on automated processing alone. Before web, print, video, or audio material is made publicly available, it goes through extensive human quality checks to verify parsing, article structure, attribution, metadata, and completeness. We continuously improve our extraction and parsing algorithms as new formats and problems appear.

Our infrastructure uses separate staging and production environments, automated deployment, daily backups, vulnerability scanning, and current actively maintained versions of core software. We follow green-stack principles to reduce unnecessary resource use, and a full disaster-recovery pipeline is in development.

Our infrastructure uses separate staging and production environments, automated deployment, daily backups, vulnerability scanning, and current actively maintained versions of core software. We follow green-stack principles to reduce unnecessary resource use, and a full disaster-recovery pipeline is in development.

Security also shapes how we design for users. We collect as little personal data as possible, do not require registration to read public archives, and treat search activity and archived content as potentially sensitive.

Security also shapes how we design for users. We collect as little personal data as possible, do not require registration to read public archives, and treat search activity and archived content as potentially sensitive.

Data access:

our archives

are built to be used

Data access: our archives

are built to be used

Data access: our archives are built to be used

Public collections can be searched without registration. For larger research or journalistic projects, we can provide structured exports, including full-corpus datasets and selections filtered by outlet, period, or topic.

Public collections can be searched without registration. For larger research or journalistic projects, we can provide structured exports, including full-corpus datasets and selections filtered by outlet, period, or topic.

Programmatic API access is in development.

Programmatic API access is in development.

For dataset or access requests, contact us and tell us what you are working on and what kind of data you need.

For dataset or access requests, contact us and tell us what you are working on and what kind of data you need.

We are Kronika, a civic technology project dedicated to preserving independent journalism under threat

Subscribe to the newsletter

© 2026 Kronika. All rights reserved

We are Kronika, a civic technology project dedicated to preserving independent journalism under threat

Subscribe to the newsletter

© 2026 Kronika. All rights reserved

We are Kronika, a civic technology project dedicated to preserving independent journalism under threat

Subscribe to the newsletter

© 2026 Kronika. All rights reserved