Skip to content

Instantly share code, notes, and snippets.

@blizzarddreams
Created February 22, 2015 03:32
Show Gist options
  • Select an option

  • Save blizzarddreams/6318b02bbacf0a4db4e6 to your computer and use it in GitHub Desktop.

Select an option

Save blizzarddreams/6318b02bbacf0a4db4e6 to your computer and use it in GitHub Desktop.
import scrapy
from scrapy.selector import Selector
class PkmnSpider(scrapy.Spider):
name = "pkmn"
allowed_domains = ["http://www.serebii.net"]
start_urls = [
"http://www.serebii.net/pokedex-xy/647.shtml",
]
def parse(self, response):
sel = Selector(response)
name = sel.xpath('/html/body/table[2]/tbody/tr[2]/td[2]/font/div[2]/div/table[1]/tbody/tr/td[1]/table/tbody/tr/td[2]/font/b').extract()
print name
@aaearon

aaearon commented Feb 22, 2015

Copy link
Copy Markdown

When Chrome/firefox(maybe, not sure) render tables, they add a 'tbody' element that isn't actually there. So when Scrapy scrapes, the 'tbody' tag in the xpath throws the whole thing off. Remove 'tbody's from your xpaths.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment