

Technology built for endangered media
Technology built for endangered media
Kronika builds and maintains its own technology for preserving media that may be blocked, deleted, altered, or lost. Our infrastructure helps recover vulnerable material, preserve original files and versions, transform difficult formats into structured data, and keep archives searchable and usable across different media environments and preservation needs.
Kronika builds and maintains its own technology for preserving media that may be blocked, deleted, altered, or lost. Our infrastructure helps recover vulnerable material, preserve original files and versions, transform difficult formats into structured data, and keep archives searchable and usable across different media environments and preservation needs.
Our tools
Our tools
Our tools
Tool 1
Archivarius
The crawl-and-recovery engine for endangered media
Archivarius is Kronika's archiving engine. It discovers, downloads and structures content from media websites using pages, feeds, sitemaps, browser rendering, and historical sources preserved by Wayback Machine in the Internet Archive.
It is designed not only to copy what remains visible today, but to reconstruct what has already disappeared. Archivarius can recover URLs no longer available through the original site or search engines, preserve links, images and embeds alongside text, and adapt its retrieval methods when sites are difficult to access. The system runs on its own: when a site blocks it, it switches retrieval methods, it notices and restarts itself.
For media facing censorship, deletion, or technical collapse, this ability to recover both the visible and the lost record is central to Kronika's work.
Tool 2
PaperAI
Turning print collections into searchable archives
Some of the most important records were never born digital. PaperAI transforms scanned newspapers, magazines, zines, and other print collections into structured, searchable digital material.
It identifies pages and articles, separates headlines, authors and body text, reconstructs stories that continue across multiple pages, and gives each article its own addressable record. OCR makes the text searchable while preserving access to the original scanned page.
PaperAI allows large print collections to become part of the same searchable archival environment as web, video, and audio sources.
Tool 3
Newsloom
Monitoring vulnerable information in real time
Journalists, researchers, and civil-society organizations often spend hours monitoring websites, feeds, social media, and other sources — many of which may be blocked, geo-restricted, or deleted without warning.
Newsloom brings these sources into one monitoring environment. It continuously collects new material, preserves content even when the original disappears, and can use AI-assisted tools for translation, summarization, categorization, and entity extraction.
AI features are optional: users can switch them off and work with the raw feed.
Note: Newsloom is currently in alpha.
Tool 1
Archivarius
The crawl-and-recovery engine for endangered media
Archivarius is Kronika's archiving engine. It discovers, downloads and structures content from media websites using pages, feeds, sitemaps, browser rendering, and historical sources preserved by Wayback Machine in the Internet Archive.
It is designed not only to copy what remains visible today, but to reconstruct what has already disappeared. Archivarius can recover URLs no longer available through the original site or search engines, preserve links, images and embeds alongside text, and adapt its retrieval methods when sites are difficult to access. The system runs on its own: when a site blocks it, it switches retrieval methods, it notices and restarts itself.
For media facing censorship, deletion, or technical collapse, this ability to recover both the visible and the lost record is central to Kronika's work.
Tool 2
PaperAI
Turning print collections into searchable archives
Some of the most important records were never born digital. PaperAI transforms scanned newspapers, magazines, zines, and other print collections into structured, searchable digital material.
It identifies pages and articles, separates headlines, authors and body text, reconstructs stories that continue across multiple pages, and gives each article its own addressable record. OCR makes the text searchable while preserving access to the original scanned page.
PaperAI allows large print collections to become part of the same searchable archival environment as web, video, and audio sources.
Tool 3
Newsloom
Monitoring vulnerable information in real time
Journalists, researchers, and civil-society organizations often spend hours monitoring websites, feeds, social media, and other sources — many of which may be blocked, geo-restricted, or deleted without warning.
Newsloom brings these sources into one monitoring environment. It continuously collects new material, preserves content even when the original disappears, and can use AI-assisted tools for translation, summarization, categorization, and entity extraction.
AI features are optional: users can switch them off and work with the raw feed.
Note: Newsloom is currently in alpha.
Tool 1
Archivarius
The crawl-and-recovery engine for endangered media
Archivarius is Kronika's archiving engine. It discovers, downloads and structures content from media websites using pages, feeds, sitemaps, browser rendering, and historical sources preserved by Wayback Machine in the Internet Archive.
It is designed not only to copy what remains visible today, but to reconstruct what has already disappeared. Archivarius can recover URLs no longer available through the original site or search engines, preserve links, images and embeds alongside text, and adapt its retrieval methods when sites are difficult to access. The system runs on its own: when a site blocks it, it switches retrieval methods, it notices and restarts itself.
For media facing censorship, deletion, or technical collapse, this ability to recover both the visible and the lost record is central to Kronika's work.
Tool 2
PaperAI
Turning print collections into searchable archives
Some of the most important records were never born digital. PaperAI transforms scanned newspapers, magazines, zines, and other print collections into structured, searchable digital material.
It identifies pages and articles, separates headlines, authors and body text, reconstructs stories that continue across multiple pages, and gives each article its own addressable record. OCR makes the text searchable while preserving access to the original scanned page.
PaperAI allows large print collections to become part of the same searchable archival environment as web, video, and audio sources.
Tool 3
Newsloom
Monitoring vulnerable information in real time
Journalists, researchers, and civil-society organizations often spend hours monitoring websites, feeds, social media, and other sources — many of which may be blocked, geo-restricted, or deleted without warning.
Newsloom brings these sources into one monitoring environment. It continuously collects new material, preserves content even when the original disappears, and can use AI-assisted tools for translation, summarization, categorization, and entity extraction.
AI features are optional: users can switch them off and work with the raw feed.
Note: Newsloom is currently in alpha.



This shared layer is what allows Kronika to build new archives without starting from zero each time.
This shared layer is what allows Kronika to build new archives without starting from zero each time.
The infrastructure provides multilingual full-text and transcript search, flexible filtering, and links outlets and authors across collections to track journalism across publications and formats.
The infrastructure provides multilingual full-text and transcript search, flexible filtering, and links outlets and authors across collections to track journalism across publications and formats.
Original files are never overwritten. New captures and corrections are stored alongside earlier versions, preserving each record's history. Every item receives a permanent identifier and metadata describing when and how it was collected.
Original files are never overwritten. New captures and corrections are stored alongside earlier versions, preserving each record's history. Every item receives a permanent identifier and metadata describing when and how it was collected.
Kronika's public archives share a common infrastructure rather than operating as isolated websites.
Kronika's public archives share a common infrastructure rather than operating as isolated websites.
The infrastructure behind our archives
The infrastructure behind our archives
The infrastructure behind our archives
Quality, security & resilience
Quality, security & resilience
Preservation is only useful
if the material can be trusted.
Preservation is only useful
if the material can be trusted.
We do not rely on automated processing alone. Before web, print, video, or audio material is made publicly available, it goes through extensive human quality checks to verify parsing, article structure, attribution, metadata, and completeness. We continuously improve our extraction and parsing algorithms as new formats and problems appear.
We do not rely on automated processing alone. Before web, print, video, or audio material is made publicly available, it goes through extensive human quality checks to verify parsing, article structure, attribution, metadata, and completeness. We continuously improve our extraction and parsing algorithms as new formats and problems appear.
Our infrastructure uses separate staging and production environments, automated deployment, daily backups, vulnerability scanning, and current actively maintained versions of core software. We follow green-stack principles to reduce unnecessary resource use, and a full disaster-recovery pipeline is in development.
Our infrastructure uses separate staging and production environments, automated deployment, daily backups, vulnerability scanning, and current actively maintained versions of core software. We follow green-stack principles to reduce unnecessary resource use, and a full disaster-recovery pipeline is in development.
Security also shapes how we design for users. We collect as little personal data as possible, do not require registration to read public archives, and treat search activity and archived content as potentially sensitive.
Security also shapes how we design for users. We collect as little personal data as possible, do not require registration to read public archives, and treat search activity and archived content as potentially sensitive.
Data access:
our archives
are built to be used
Data access: our archives
are built to be used
Data access: our archives are built to be used
Public collections can be searched without registration. For larger research or journalistic projects, we can provide structured exports, including full-corpus datasets and selections filtered by outlet, period, or topic.
Public collections can be searched without registration. For larger research or journalistic projects, we can provide structured exports, including full-corpus datasets and selections filtered by outlet, period, or topic.
Programmatic API access is in development.
Programmatic API access is in development.
For dataset or access requests, contact us and tell us what you are working on and what kind of data you need.
For dataset or access requests, contact us and tell us what you are working on and what kind of data you need.
We are Kronika, a civic technology project dedicated to preserving independent journalism under threat
Subscribe to the newsletter

© 2026 Kronika. All rights reserved
We are Kronika, a civic technology project dedicated to preserving independent journalism under threat
Subscribe to the newsletter

© 2026 Kronika. All rights reserved
We are Kronika, a civic technology project dedicated to preserving independent journalism under threat
Subscribe to the newsletter

© 2026 Kronika. All rights reserved