Chorelet
Apify Actor

Internet Archive Scraper

Search 50 million archive.org items and get a direct download link for every file.

Pay per result on Apify. Runs that return nothing cost nothing; Apify's free plan includes $5 of usage a month, and Apify Bronze, Silver and Gold subscribers get 10%, 20% and 30% off.

Query the Internet Archive by text, collection, media type, creator, subject, language and the year of the work, and get one row per item: title, creator, description, dates, collections, subjects, language, licence, downloads, size, rating and the item page. A plain multi-word search runs as a phrase, and the archive's own Lucene syntax goes through untouched.

Turn on «a row per file» and every item brings its files: name, format, size, duration, MD5 and a direct download URL. Derivatives — thumbnails, torrents, checksums — are skipped by default, and a format filter understands both the archive's names and plain extensions, so «mp4» finds files labelled 512Kb MPEG4.

What you get

  • • Collection, media type, creator, subject and year filters
  • • Direct download URL for every file
  • • Derivatives skipped, formats matched by extension
  • • Licence and download count on every row
  • • Monitored daily

Typical uses

  • • Building a public-domain media library
  • • Datasets for speech, OCR and film research
  • • Finding what is safe to reuse by licence
  • • Bulk download lists for a collection

Output

identifier, title, creator, description, mediaType, detailsUrlThe item
date, publishedAt, year, collections, subjects, language, licenseUrlHow it is catalogued
downloads, itemSizeBytes, rating, reviews, formatsHow popular and how big
fileName, fileFormat, fileSizeBytes, durationSeconds, fileUrl, md5File rows

Run it from code

Same Actor, same output, from your own scripts (any language) or from n8n, Make and Zapier.

curl -X POST "https://api.apify.com/v2/acts/chorelet~internet-archive-scraper/run-sync-get-dataset-items?token=YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"collections":["librivoxaudio"],"mediaTypes":["audio"],"includeFiles":true,"fileFormats":["mp3"],"maxResultsPerQuery":200}'