Created
April 24, 2010 14:03
-
-
Save simonw/377676 to your computer and use it in GitHub Desktop.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| "Extracts the text from every slide in a keynote presentation" | |
| # Limitation: doesn't currently deal with markup nested inside a p tag, e.g. | |
| # the following: | |
| # <ns2:p ns2:list-level="1" ns2:style="SFWPParagraphStyle-990">Pen knife photo: | |
| # herzogbr <ns2:link | |
| # href="http://www.flickr.com/photos/80516279@N00/2274372747/"><ns2:span | |
| # ns2:style="SFWPCharacterStyle-158" | |
| # >http://www.flickr.com/photos/80516279@N00/2274372747/</ns2:span></ns2:link> | |
| # </ns2:p> | |
| import zipfile | |
| from xml.etree import ElementTree as ET | |
| def extract_keynote_slide_text(fp): | |
| z = zipfile.ZipFile(fp, 'r') | |
| root = ET.fromstring(z.read('index.apxl')) | |
| slides = root.find( | |
| '{http://developer.apple.com/namespaces/keynote2}slide-list' | |
| ) | |
| slide_texts = [] | |
| for slide in slides: | |
| slide_text = [] | |
| ps = slide.findall('.//{http://developer.apple.com/namespaces/sf}p') | |
| for p in ps: | |
| text = p.text | |
| if text is None or not text.strip(): | |
| continue | |
| slide_text.append(text) | |
| slide_texts.append(slide_text) | |
| return slide_texts |
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Maybe a TextIterator helps with the nested markup, something like:
" ".join([t.strip() for t in ET.ElementTextIterator(ps)])