Chat with AI and run tools instantly.
Let AI agents search Internet Archive books and PDFs, explore catalog metadata, inspect Wayback snapshots, and retrieve archived HTML.
Try this workflow
Search books by title
Search Internet Archive for books titled Plato published as texts, return the top matches with identifier, creator, and date.
Chat with AI and run tools instantly.
Browse all toolsBuilt for
Web Archiving, Research & Academia, Journalism & Media, Legal Discovery
Search books by title
Uses catalog metadata search — not Wayback URL lookup.
Search Internet Archive for books titled Plato published as texts, return the top matches with identifier, creator, and date.
Find phrase in scanned texts
OCR search inside volumes — distinct from title-only metadata search.
Full-text search Internet Archive for 'restorative justice' in digitized books and return highlighted excerpts with page numbers and item identifiers.
Find archived snapshots
Turns a URL into a concrete historical record with replay links.
Check whether this URL has Wayback Machine captures, then return the closest snapshot date, archive link, HTTP status, and content type.
Retrieve archived HTML
Moves from Wayback discovery to content replay for historical analysis.
Find a reliable archived snapshot of this page from early 2023, retrieve the raw HTML, and summarize the visible page title and key text.
What can agents do with Internet Archive here?
Agents can search catalog metadata, run OCR full-text search in books/PDFs, fetch item metadata and files, check Wayback snapshots, batch-check URLs, search CDX metadata, build capture timelines, and retrieve archived HTML.
Does Internet Archive need sign-in or a paid account?
No. The public catalog, full-text, Wayback, and metadata workflows here work without an Archive.org sign-in or paid account.
When should I use full-text search vs catalog metadata search?
Use catalog metadata when matching titles, creators, collections, or identifiers. Use full-text search when you need words that appear inside scanned books or PDFs, with highlights and page numbers.
Domain knowledge for Internet Archive — workflow patterns, data models, and gotchas for your AI agent.
Two distinct search surfaces plus Wayback web archaeology. No auth required.
| Goal | Surface |
|---|---|
| Was a web URL captured? Replay old HTML? | Wayback — availability, CDX, timemap, get content |
| Find items by title, creator, collection, mediatype | Catalog metadata — Lucene queries on item records |
| Walk an entire collection (deep paging) | Catalog scrape — cursor batches (min 100 per call) |
| Find a phrase inside scanned books/PDFs | Full-text (OCR) — highlighted excerpts + page numbers |
| Files and description for one identifier | Item metadata lookup |
| Page locations for text in one known book | Search inside item (when the Books API responds) |
Catalog metadata matches titles/subjects/collections — not words on page 282. Full-text search matches OCR inside volumes. Wayback matches live-web URLs over time. Do not substitute one for another.
Scroll through large Internet Archive catalog result sets using the Scrape API. Use when metadata search pagination is not enough to walk an entire co...
Search OCR text inside Internet Archive books, PDFs, and other digitized items. Returns highlighted excerpts and page numbers when available. This is ...
Retrieve metadata and optional file inventory for one Internet Archive item. Use after catalog or full-text search when you need description, collecti...
Search Internet Archive catalog metadata — titles, creators, collections, mediatypes, dates, and download counts. Uses the Advanced Search index (same...
Find where a phrase occurs inside one Internet Archive book or PDF. Returns page numbers and surrounding text when the Books Search Inside API respond...
Check if a URL has been archived by the Wayback Machine. Returns the closest available snapshot with its timestamp and archive URL. Pass a timestamp t...
Check availability of multiple URLs in a single batch request. Pass an array of {url, tag?, timestamp?, closest?} objects. Each result is tagged for e...
Search the Internet Archive's Wayback Machine CDX index for detailed archive metadata. The CDX index contains detailed information about every archiv...
Retrieve the actual archived HTML content of a web page from the Wayback Machine. Returns the archived version of a web page's content, either as pro...
Retrieve a complete timeline (timemap) of all archived snapshots for a URL. Returns all available archived versions of a web page with their timestam...
Agents can search catalog metadata, run OCR full-text search in books/PDFs, fetch item metadata and files, check Wayback snapshots, batch-check URLs, search CDX metadata, build capture timelines, and retrieve archived HTML.
No. The public catalog, full-text, Wayback, and metadata workflows here work without an Archive.org sign-in or paid account.
Use catalog metadata when matching titles, creators, collections, or identifiers. Use full-text search when you need words that appear inside scanned books or PDFs, with highlights and page numbers.
Use CDX search when availability returns an empty result, when you need many captures, or when you need metadata such as status codes, MIME types, digests, and timestamp ranges.
Add the MCPBundles server URL to your MCP client configuration (Claude Desktop, Cursor, VS Code, etc.). The URL format is: https://mcp.mcpbundles.com/bundle/internet-archive. Authentication is handled automatically.
Internet Archive provides 10 tools that can be called by AI agents, along with a SKILL.md that gives your AI agent domain knowledge about when and how to use them.
Internet Archive uses open data APIs — no authentication required.
Connect Internet Archive to any MCP client in minutes
https://mcp.mcpbundles.com/bundle/internet-archiveThe link prefills the Add custom connector dialog — you still review the values and click Add, then Connect to complete OAuth.
Internet Archive and paste the MCP URL into Remote MCP server URL.Custom connectors at claude.ai require a paid Claude plan (Pro, Max, Team, or Enterprise).
More search integrations you might like
World Bank API providing access to global economic indicators, development data, and country statist...
Zenodo open research repository hosted by CERN. Search and retrieve research outputs including paper...
Crunchbase is the leading platform for company intelligence. Search and look up organizations, peopl...
NCBI E-utilities API providing access to PubMed and other biomedical databases with search and retri...
OpenAlex is an open catalog of the global research system. Search works, retrieve work metadata, and...
Semantic Scholar Academic Graph, Recommendations, and Datasets APIs — search papers, explore citatio...