Description
Website Recovery Importer takes a recovery package (a WXR file plus the site’s images) and imports it into your WordPress site. It is designed for the packages produced by Tesora from the Internet Archive’s Wayback Machine, but it will import any package that follows the same simple layout (a manifest.json, one or more wxr/part-*.xml files, and a media/ folder).
Important: this plugin does not crawl or scrape the Wayback Machine itself. It imports a package that was prepared beforehand. It just does the WordPress side — reliably and safely.
What it does:
- Bulk import in batches via AJAX, so large sites don’t hit PHP time or memory limits.
- Everything as DRAFT by default — you review before anything goes live. It never auto-publishes.
- Sanitizes third-party HTML with an explicit
wp_ksesallow-list — scripts, iframes and inline event handlers are stripped, regardless of your role. - Places the images and keeps the content pointing at them.
- Rewrites internal links from the old domain to your site, for every host variant (http/https, with/without www).
- Page vs post respected from the package, with a switch to force everything to Pages or Posts.
- Idempotent — re-importing the same package updates instead of duplicating (stable per-page identity).
- Two ways in — upload the package as a ZIP, or drop it by FTP into
wp-content/uploads/tesora-import/.
The content comes from third parties (archived pages). Review copyright, trademarks and content before publishing. You use it at your own responsibility — which is exactly why everything lands as a draft.
Screenshots





Installation
- In WordPress go to Plugins Add New Upload Plugin and choose the plugin ZIP, then Activate. (Or copy the
website-recovery-importerfolder intowp-content/plugins/.) - Go to Tools Website Recovery Importer.
- Either upload a package ZIP, or drop the package folder by FTP into
wp-content/uploads/tesora-import/<domain>/and it will be detected. - Choose the type behaviour (respect the package / force Pages / force Posts) and press Import. Everything is created as a draft.
FAQ
-
Does it download pages from the Wayback Machine?
-
No. It imports a package that was already prepared (for example by Tesora). This plugin only handles the WordPress import side.
-
Will it publish my drafts automatically?
-
No. Everything is imported as a draft by default. You decide what to publish.
-
Is the imported HTML safe?
-
The HTML is third-party content, so it is always passed through a strict
wp_ksesallow-list before being saved — scripts, iframes and event handlers are removed even for administrators. -
What if I import the same package twice?
-
It won’t duplicate. Each page has a stable identity, so a second import updates the existing draft instead of creating a copy (unless you’ve edited it, in which case your edits are kept).
Reviews
There are no reviews for this plugin.
Contributors & Developers
“Website Recovery Importer” is open source software. The following people have contributed to this plugin.
ContributorsTranslate “Website Recovery Importer” into your language.
Interested in development?
Browse the code, check out the SVN repository, or subscribe to the development log by RSS.
Changelog
1.2.5
- Security: the nonce and capability check is now written out in full inside each AJAX handler, instead of living in a helper method. Same behaviour, but static analysers and reviewers can see it. Removed the file-wide PHPCS suppression that was hiding it, and narrowed the two remaining suppressions to the exact lines, with the reason spelled out.
1.2.4
- Compatibility: marked as tested up to WordPress 7.1.
1.2.3
- Compatibility: marked as tested up to WordPress 7.0.
1.2.2
- Code quality: passes the official Plugin Check. Documented the direct filesystem use in the streaming ZIP extractor (WP_Filesystem has no streaming API and would prompt for FTP credentials mid-import), prefixed the uninstaller’s global variables, and corrected the “Tested up to” header.
- The admin screen now shows the plugin version in a small footer.
1.2.1
- Fix: the legal notice no longer crowds the header. It now shows as a full-width band under the title (added the standard WordPress
wp-header-endmarker so admin notices are placed below the header instead of inside it).
1.2.0
- New: language switch (English / Spanish) and a light/dark theme toggle on the import screen.
- New: package preview before importing — it shows the domain and how many pages and images were detected, so you confirm before anything is created.
- New: a Stop button during a running import, with resume from where it left off.
1.1.0
- Renamed to “Website Recovery Importer”.
- Hardening: the WXR reader now rejects DOCTYPE / external entities (XXE), the ZIP extractor caps the decompressed size (zip-bomb guard) with real-byte accounting, media files are validated as real images before being copied, the job lock is now atomic, and a failing (“poison”) item is retried only a bounded number of times.
- Import errors are now reported incrementally instead of only at the end.
- More robust internal-link rewriting (word-boundary aware) and a domain-independent stable identity for safer idempotent re-imports.
1.0.1
- Extraction now uses an allow-list (images + manifest.json + WXR + .txt). Other bundled assets such as .js/.css from the archived site are skipped instead of rejecting the whole package, which fixes an import error on packages that carried the dead site’s scripts or styles. Dangerous file types (php, svg, html, js) are still never written to disk.
1.0.0
- Initial release: batch import (ZIP or FTP), draft-by-default, HTML sanitizing, media placement, internal-link rewriting, idempotent re-import, page/post switch.
