#

warc

Here are 89 public repositories matching this topic...

ArchiveBox

ArchiveBox / ArchiveBox

🗃 Open source self-hosted web archiving. Takes URLs/browser history/bookmarks/Pocket/Pinboard/etc., saves HTML, JS, PDFs, media, and more...

Updated Jun 10, 2022
Python

internetarchive / heritrix3

Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project.

java warc heritrix webcrawling

Updated Jun 22, 2022
Java

conifer

Rhizome-Conifer / conifer

Collect and revisit web pages.

python docker archives warc web-archiving wayback webrecorder pywb

Updated Jun 22, 2022
Python

ArchiveTeam / grab-site

The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns

crawler spider archiving crawl warc

Updated May 7, 2022
Python

ipwb

oduwsdl / ipwb

InterPlanetary Wayback: A distributed and persistent archive replay system using IPFS

python docker service-worker ipfs memento warc web-archiving wayback memento-rfc

Updated Jun 2, 2022
Python

webrecorder / webrecorder-player

Sponsor

Webrecorder Player for Desktop (OSX/Windows/Linux). (Built with Electron + Webrecorder)

electron warc web-archiving webrecorder pywb

Updated Sep 17, 2020
JavaScript

wail

machawk1 / wail

🐋 Web Archiving Integration Layer: One-Click User Instigated Preservation

python gui warc web-archiving pyinstaller wayback heritrix openwayback

Updated May 27, 2022
Roff

Florents-Tselai / WarcDB

WarcDB: Web crawl data as SQLite databases.

cli database sqlite crawling warc web-archiving web-data

Updated Jun 21, 2022
Python

webrecorder / warcio

Sponsor

Streaming WARC/ARC library for fast web archive IO

python warc web-archiving web-archives pywb

Updated May 31, 2022
Python

webrecorder / replayweb.page

Sponsor

Serverless Web Archive Replay directly in the browser

service-worker warc web-archiving wayback-machine web-archive replay-web-page web-replay

Updated Jun 15, 2022
JavaScript

bitextor

bitextor / bitextor

Bitextor generates translation memories from multilingual websites

crawler dictionaries tokenizer machine-translation wget apertium neural-machine-translation warc tmx statistical-machine-translation corpus-generator httrack sentence-segmentation corpus-tools creepy corpus-processing hunalign parallel-corpora document-aligner bicleaner

Updated Jun 21, 2022
Python

warcreate

machawk1 / warcreate

Chrome extension to "Create WARC files from any webpage"

chrome-extension warc web-archiving

Updated May 31, 2022
JavaScript

cocrawler / cocrawler

CoCrawler is a versatile web crawler built using modern tools and concurrency.

screenshot crawler concurrency async-python python3 aiohttp warc aiohttp-client pluggable-modules

Updated Apr 29, 2022
Python

commoncrawl / news-crawl

News crawling with Storm-crawler - stores content as WARC

crawler news web-crawler apache-storm warc commoncrawl common-crawl

Updated Mar 31, 2022
Java

helgeho / ArchiveSpark

An Apache Spark framework for easy data processing, extraction as well as derivation for web archives and archival collections, developed at Internet Archive.

spark internet-archive warc web-archiving webarchive archivespark spark-framework

Updated Oct 8, 2021
Scala

N0taN3rd / wail

🐋 One-Click User Instigated Preservation

electron warc web-archiving high-fidelity-preservation browser-based-presrevation

Updated Feb 3, 2019
JavaScript

cocrawler / cdx_toolkit

A toolkit for CDX indices such as Common Crawl and the Internet Archive's Wayback Machine

python warc web-archiving cdx web-archives commoncrawl cdx-api

Updated Mar 28, 2022
Python

CGamesPlay / chronicler

Offline-first web browser

electron browser warc

Updated Jan 14, 2019
JavaScript

N0taN3rd / node-warc

Parse And Create Web ARChive (WARC) files with node.js

warc web-archiving webarchive web-archives webarchiving warc-files chrome-remote-interface pupeteer

Updated Apr 28, 2022
JavaScript

ArchiveTeam / wget-lua

Wget-AT is a modern Wget with Lua hooks, Zstandard (+dictionary) WARC compression and URL-agnostic deduplication.

crawler scraper downloader spider lua ftp scraping crawling archiving wget crawl zstd crawlers warc webarchiving archiveteam wget-lua

Updated Jun 8, 2022
C

archivesunleashed / warclight

A Rails engine supporting the discovery of web archives.

ruby rails rails-engine solr discovery blacklight warc webarchives webarchive-discovery

Updated Jun 22, 2022
Ruby

centic9 / CommonCrawlDocumentDownload

Sponsor

A small tool which uses the CommonCrawl URL Index to download documents with certain file types or mime-types. This is used for mass-testing of frameworks like Apache POI and Apache Tika

java mime-types warc cdx-files commoncrawl

Updated Jun 7, 2022
Java

PromyLOPh / crocoite

Web archiving using Google Chrome

devtools archiving chrome-browser warc

Updated Dec 30, 2019
Python

pirate / internet-archiving-talk

Sponsor

🎭 An introduction to the Internet Archiving ecosystem, tooling, and some of the ethical dilemmas that the community faces.

slideshow wget talks warc censorship web-archiving ethics internet-archiving archivebox

Updated Oct 19, 2020
JavaScript

datatogether / warc

Golang WARC (Web ARChive) Library

golang package archiving warc iipc

Updated Aug 6, 2019
Go

jedireza / warc

⚙️ A Rust library for reading and writing WARC files

rust rust-library warc

Updated Mar 25, 2022
Rust

hrbrmstr / warc

Sponsor

📇 Tools to Work with the Web Archive Ecosystem in R

r rstats warc warc-files r-cyber warc-ecosystem

Updated Aug 20, 2017
R

Mixnode / mixnode-warcreader-php

Read Web ARChive (WARC) files in PHP.

php warc webarchive

Updated Mar 10, 2017
PHP

chatnoir-eu / chatnoir-resiliparse

A robust web archive analytics toolkit

python web cpp cython bigdata extraction warc webarchive htmlparser

Updated Jun 10, 2022
Cython

webrecorder / cdxj-indexer

Sponsor

CDXJ Indexing of WARC/ARCs

warc web-archiving

Updated Jun 21, 2022
Python

Improve this page

Add a description, image, and links to the warc topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the warc topic, visit your repo's landing page and select "manage topics."