Skip to content

Instantly share code, notes, and snippets.

View electro0nes's full-sized avatar
🏠
Working from home

Moein Erfanian electro0nes

🏠
Working from home
View GitHub Profile
@Smerity
Smerity / get_all_urls.py
Created June 23, 2015 01:05
Collect all URLs for NYTimes in the Common Crawl URL Index
import requests
show_pages = 'http://index.commoncrawl.org/CC-MAIN-2015-18-index?url={query}&output=json&showNumPages=true'
get_page = 'http://index.commoncrawl.org/CC-MAIN-2015-18-index?url={query}&output=json&page={page}'
query = 'nytimes.com/*'
show = requests.get(show_pages.format(query=query))
pages = show.json()['pages']
results = set()