Chat with AI and run tools instantly.
Let AI agents search Internet Archive books and PDFs, explore catalog metadata, inspect Wayback snapshots, and retrieve archived HTML.
Try this workflow
Search books by title
Search Internet Archive for books titled Plato published as texts, return the top matches with identifier, creator, and date.
Opens ChatGPT on the web or desktop and asks it to use the WebMCP tools available here.
You’ll sign in to MCPBundles when your client connects.
https://mcp.mcpbundles.com/bundle/internet-archiveSign in once, then chat here with saved access — one connection to many servers, with a history of what your AI ran.
Chat with AI and run tools instantly.
Browse all toolsLive probe refreshed Sep 14, 2026 · Endpoint host mcp.mcpbundles.com
Built for
Web Archiving, Research & Academia, Journalism & Media, Legal Discovery
Search books by title
Uses catalog metadata search — not Wayback URL lookup.
Search Internet Archive for books titled Plato published as texts, return the top matches with identifier, creator, and date.
Find phrase in scanned texts
OCR search inside volumes — distinct from title-only metadata search.
Full-text search Internet Archive for 'restorative justice' in digitized books and return highlighted excerpts with page numbers and item identifiers.
Find archived snapshots
Turns a URL into a concrete historical record with replay links.
Check whether this URL has Wayback Machine captures, then return the closest snapshot date, archive link, HTTP status, and content type.
Retrieve archived HTML
Moves from Wayback discovery to content replay for historical analysis.
Find a reliable archived snapshot of this page from early 2023, retrieve the raw HTML, and summarize the visible page title and key text.
What can agents do with Internet Archive here?
Agents can search catalog metadata, run OCR full-text search in books/PDFs, fetch item metadata and files, check Wayback snapshots, batch-check URLs, search CDX metadata, build capture timelines, and retrieve archived HTML.
Does Internet Archive need sign-in or a paid account?
No. The public catalog, full-text, Wayback, and metadata workflows here work without an Archive.org sign-in or paid account.
When should I use full-text search vs catalog metadata search?
Use catalog metadata when matching titles, creators, collections, or identifiers. Use full-text search when you need words that appear inside scanned books or PDFs, with highlights and page numbers.
Domain knowledge for Internet Archive — workflow patterns, data models, and gotchas for your AI agent.
Two distinct search surfaces plus Wayback web archaeology. No auth required.
| Goal | Surface |
|---|---|
| Was a web URL captured? Replay old HTML? | Wayback — availability, CDX, timemap, get content |
| Find items by title, creator, collection, mediatype | Catalog metadata — Lucene queries on item records |
| Walk an entire collection (deep paging) | Catalog scrape — cursor batches (min 100 per call) |
| Find a phrase inside scanned books/PDFs | Full-text (OCR) — highlighted excerpts + page numbers |
| Files and description for one identifier | Item metadata lookup |
| Page locations for text in one known book | Search inside item (when the Books API responds) |
Catalog metadata matches titles/subjects/collections — not words on page 282. Full-text search matches OCR inside volumes. Wayback matches live-web URLs over time. Do not substitute one for another.
Archive Catalog Scrape
Scroll through large Internet Archive catalog result sets using the Scrape API. Use when metadata search pagination is not enough to walk an entire collection. Returns an opaque cursor for the next batch. Searches item metadata only — not OCR/full-text inside books.
Archive Fulltext Search
Search OCR text inside Internet Archive books, PDFs, and other digitized items. Returns highlighted excerpts and page numbers when available. This is not Wayback web snapshot search and not catalog title-only metadata search — it matches words appearing inside scanned volumes.
Archive Get Item
Retrieve metadata and optional file inventory for one Internet Archive item. Use after catalog or full-text search when you need description, collections, mediatype, and downloadable file names for a specific identifier.
Archive Metadata Search
Search Internet Archive catalog metadata — titles, creators, collections, mediatypes, dates, and download counts. Uses the Advanced Search index (same backend as archive.org advanced search). This searches item records, not OCR text inside books. For phrase matches inside scanned books/PDFs, use OCR full-text search instead. For historical web page snapshots, use the Wayback tools instead.
Archive Search Inside
Find where a phrase occurs inside one Internet Archive book or PDF. Returns page numbers and surrounding text when the Books Search Inside API responds. Prefer corpus-wide full-text search to discover items; use this when you already know the identifier and need in-item page locations.
Get Wayback Available
Check if a URL has been archived by the Wayback Machine. Returns the closest available snapshot with its timestamp and archive URL. Pass a timestamp to find the snapshot closest to a specific date. For multiple URLs at once, use the batch availability check instead. This endpoint can return empty results for some well-known URLs — CDX index search is a more reliable alternative.
Post Wayback Available
Check availability of multiple URLs in a single batch request. Pass an array of {url, tag?, timestamp?, closest?} objects. Each result is tagged for easy collation. Use this instead of repeating single-URL availability checks. For thorough archive metadata beyond simple availability, use CDX index search.
Wayback Cdx Search
Search the Internet Archive's Wayback Machine CDX index for detailed archive metadata. The CDX index contains detailed information about every archived URL, including timestamps, HTTP status codes, content types, and content digests. This tool provides raw access to the Wayback Machine's index data, useful for historical analysis, content discovery, and research. Results are returned in the requested format with full technical metadata about archived web pages.
Wayback Get Content
Retrieve the actual archived HTML content of a web page from the Wayback Machine. Returns the archived version of a web page's content, either as processed HTML (with Wayback Machine interface) or raw original content. This allows you to access the actual historical content of web pages, useful for content analysis, historical research, and retrieving pages that may no longer exist on the live web. Use with timestamps from CDX search or timemap tools for specific archived versions.
Wayback Timemap
Retrieve a complete timeline (timemap) of all archived snapshots for a URL. Returns all available archived versions of a web page with their timestamps, using the Memento protocol standard. This provides a complete view of how a page has changed over time, useful for historical research, content analysis, and tracking website evolution. The timemap shows when pages were archived and provides direct links to view each snapshot.
Agents can search catalog metadata, run OCR full-text search in books/PDFs, fetch item metadata and files, check Wayback snapshots, batch-check URLs, search CDX metadata, build capture timelines, and retrieve archived HTML.
No. The public catalog, full-text, Wayback, and metadata workflows here work without an Archive.org sign-in or paid account.
Use catalog metadata when matching titles, creators, collections, or identifiers. Use full-text search when you need words that appear inside scanned books or PDFs, with highlights and page numbers.
Use CDX search when availability returns an empty result, when you need many captures, or when you need metadata such as status codes, MIME types, digests, and timestamp ranges.
Add the MCPBundles server URL to your MCP client configuration (Claude Desktop, Cursor, VS Code, etc.). The URL format is: https://mcp.mcpbundles.com/bundle/internet-archive. Authentication is handled automatically.
Internet Archive provides 10 tools that can be called by AI agents, along with a SKILL.md that gives your AI agent domain knowledge about when and how to use them.
Internet Archive uses open data APIs — no authentication required.
Connect Internet Archive to any MCP client in minutes
Other MCP servers in this category from the directory index
Exa offers fast and intelligent web search capabilities, along with web crawling features to gather ...
Discover and book the best experiences worldwide! Immerse yourself in concerts, art, events, top att...
SuggestAPI exposes commerce tools through MCP and WebMCP on its Agent Gateway
Recall-first web search API for comprehensive event retrieval — finds every relevant event across th...
Real-time web search API and MCP for AI agents and RAG apps.
Unpaywall API providing open access status and full-text availability for scholarly articles.
5 tools