Wayback Restorator

Configuration

Capture selection, crawl scope, requests, and storage settings.

Configure CLI restorations with options:

wayback-restorator restore 'https://web.archive.org/web/20260925084419/https://example.com/' \
  --concurrency 4 --max-pages 100

In the dashboard, use the corresponding form fields. The dashboard saves those settings with each queued job.

Environment variables

Every CLI option has an environment variable named after it: --max-pages becomes MAX_PAGES, and the capture URL is WAYBACK_URL. Options passed on the command line take precedence. Without a command, wayback-restorator restores the capture configured in the environment:

WAYBACK_URL='https://web.archive.org/web/20260925084419/https://example.com/' \
  CONCURRENCY=4 wayback-restorator

Capture selection

CLI optionDashboard fieldDefaultValues / meaning
--capture-modeCapture modereplayreplay or strict.
--max-page-capture-drift-daysPage drift (days)120Non-negative number of days between the requested capture and a page capture.
--max-linked-page-capture-drift-daysLinked page drift (days)365Older linked-page fallback window in replay mode; at least the page drift value.
--max-asset-capture-drift-daysAsset drift (days)3650Non-negative asset capture window used in strict mode.

Replay mode

Replay is the default. Pages use the page drift window. For a discovered linked page, the engine can also use an older capture within the linked-page window. Assets are selected from the available capture candidates across dates.

The report records linked pages recovered through the older-capture fallback, including their actual timestamps.

Strict mode

Strict mode applies the page drift window to pages and the asset drift window to assets. Select it when you want captures within specific date windows:

wayback-restorator restore 'https://web.archive.org/web/20260925084419/https://example.com/' \
  --capture-mode strict --max-page-capture-drift-days 30 --max-asset-capture-drift-days 90

Crawl scope

CLI optionDashboard fieldDefaultValues / meaning
--max-pagesMax pagesUnlimitedPositive integer; includes the starting page. 0 means unlimited.
--max-depthMax depthUnlimitedPositive integer; the starting page has depth 0. 0 means unlimited.
--max-filesMax filesUnlimitedPositive integer; counts resources reaching completed, unavailable, or failed states.
--cross-origin-assetsInclude cross-origin assetsOffInclude assets from other origins through Wayback. Turn off with --no-cross-origin-assets.

The dashboard uses empty fields for unlimited values. Page and depth settings control page discovery; assets referenced by selected pages are still discovered. File bounds count all resources, including pages and assets, across the saved job.

A job with a restored entry page stopped by a file bound with resources still pending is bounded. It writes a report; continuing it to completion creates the ZIP. A page or depth bound defines the selected crawl scope, which can produce a completed ZIP.

Requests

CLI optionDashboard fieldDefaultValues / meaning
--concurrencyConcurrent requests6Integer from 1 to 32.
--request-max-attemptsRequest attempts3Positive integer; total attempts for a transient request error.
--retry-base-delay-secondsRetry base (seconds)1Non-negative number; initial retry delay.
--retry-max-delay-secondsRetry max (seconds)30Non-negative number, at least the base delay.

Retries use an exponential delay capped by the maximum retry delay for transient network errors and retryable Wayback responses. Capture windows and retry delays must be below 100000.

Proxy

Set the PROXY_URL environment variable to route requests from the CLI and the dashboard through an HTTP or HTTPS proxy:

PROXY_URL=http://proxy.example:3128 \
  wayback-restorator restore 'https://web.archive.org/web/20260925084419/https://example.com/'

Credentials can be included in the URL, for example http://user:password@proxy.example:3128. A value without a scheme uses http://. An empty value uses a direct connection. The proxy has no command-line option, so credentials stay out of your shell history.

Storage

CLI optionEnvironment variableDefaultContents
--work-dirWORK_DIRdata/workJob state and downloaded resources.
--output-dirOUTPUT_DIRdata/outputReports and ZIPs.
--job-idJOB_IDDerived from the capture URLJob directory name: 1–180 letters, digits, dots, underscores, tildes, or hyphens.

Relative paths start from the directory where you run the command. Both restore and web accept the directory options; the dashboard generates its own job IDs. Keep the working files to continue saved jobs.

The Docker image uses /app/data/work and /app/data/output. Mount a host directory or volume at /app/data to keep them.

Dashboard server

CLI optionEnvironment variableDefaultPurpose
--hostWEB_HOST127.0.0.1Address the dashboard listens on.
--portWEB_PORT8080Port the dashboard listens on.

Keep the default host unless the dashboard runs in a container. It has no authentication.

Output plugins

--plugins-json takes a JSON array of plugin configurations. Its default is []. In the dashboard, paste the array into Output plugins JSON.

wayback-restorator restore 'https://web.archive.org/web/20260925084419/https://example.com/' \
  --plugins-json '[{"id":"clean","plugin":"strip_external_urls","options":{}}]'

See Plugins for available plugins, options, and custom transformations.

On this page