#

warc

Here are 91 public repositories matching this topic...

ArchiveBox

ArchiveBox / ArchiveBox

🗃 Open source self-hosted web archiving. Takes URLs/browser history/bookmarks/Pocket/Pinboard/etc., saves HTML, JS, PDFs, media, and more...

Updated Dec 5, 2022
Python

internetarchive / heritrix3

Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project.

java warc heritrix webcrawling

Updated Nov 16, 2022
Java

conifer

Rhizome-Conifer / conifer

Collect and revisit web pages.

python docker archives warc web-archiving wayback webrecorder pywb

Updated Dec 6, 2022
Python

ArchiveTeam / grab-site

The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns

crawler spider archiving crawl warc

Updated Dec 5, 2022
Python

ipwb

oduwsdl / ipwb

InterPlanetary Wayback: A distributed and persistent archive replay system using IPFS

python docker service-worker ipfs memento warc web-archiving wayback hacktoberfest memento-rfc

Updated Dec 6, 2022
Python

webrecorder / webrecorder-player

Sponsor

Webrecorder Player for Desktop (OSX/Windows/Linux). (Built with Electron + Webrecorder)

electron warc web-archiving webrecorder pywb

Updated Sep 17, 2020
JavaScript

Florents-Tselai / WarcDB

WarcDB: Web crawl data as SQLite databases.

cli database sqlite crawling warc web-archiving web-data

Updated Nov 15, 2022
Python

wail

machawk1 / wail

🐋 Web Archiving Integration Layer: One-Click User Instigated Preservation

python gui warc web-archiving pyinstaller wayback heritrix openwayback

Updated Sep 7, 2022
Roff

webrecorder / replayweb.page

Sponsor

Serverless Web Archive Replay directly in the browser

service-worker warc web-archiving wayback-machine web-archive replay-web-page web-replay

Updated Dec 8, 2022
JavaScript

webrecorder / warcio

Sponsor

Streaming WARC/ARC library for fast web archive IO

python warc web-archiving web-archives pywb

Updated Jun 26, 2022
Python

bitextor

bitextor / bitextor

Bitextor generates translation memories from multilingual websites

Updated Dec 7, 2022
Python

commoncrawl / news-crawl

News crawling with Storm-crawler - stores content as WARC

crawler news web-crawler apache-storm warc commoncrawl common-crawl

Updated Nov 16, 2022
Java

warcreate

machawk1 / warcreate

Chrome extension to "Create WARC files from any webpage"

chrome-extension warc web-archiving

Updated May 31, 2022
JavaScript

cocrawler / cocrawler

CoCrawler is a versatile web crawler built using modern tools and concurrency.

screenshot crawler concurrency async-python python3 aiohttp warc aiohttp-client pluggable-modules

Updated Apr 29, 2022
Python

helgeho / ArchiveSpark

An Apache Spark framework for easy data processing, extraction as well as derivation for web archives and archival collections, developed at Internet Archive.

spark internet-archive warc web-archiving webarchive archivespark spark-framework

Updated Oct 8, 2021
Scala

cocrawler / cdx_toolkit

A toolkit for CDX indices such as Common Crawl and the Internet Archive's Wayback Machine

python warc web-archiving cdx web-archives commoncrawl cdx-api

Updated Mar 28, 2022
Python

N0taN3rd / wail

🐋 One-Click User Instigated Preservation

electron warc web-archiving high-fidelity-preservation browser-based-presrevation

Updated Feb 3, 2019
JavaScript

maxcountryman / warc-parquet

Sponsor

🗄️ A simple CLI for converting WARC to Parquet.

crawling parquet warc web-archiving duckdb

Updated Sep 2, 2022
Rust

N0taN3rd / node-warc

Parse And Create Web ARChive (WARC) files with node.js

warc web-archiving webarchive web-archives webarchiving warc-files chrome-remote-interface pupeteer

Updated Dec 2, 2022
JavaScript

CGamesPlay / chronicler

Offline-first web browser

electron browser warc

Updated Jan 14, 2019
JavaScript

Improve this page

Add a description, image, and links to the warc topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the warc topic, visit your repo's landing page and select "manage topics."