Skip to content

Instantly share code, notes, and snippets.

@simonw
Created April 24, 2010 14:03
Show Gist options
  • Select an option

  • Save simonw/377676 to your computer and use it in GitHub Desktop.

Select an option

Save simonw/377676 to your computer and use it in GitHub Desktop.
"Extracts the text from every slide in a keynote presentation"
# Limitation: doesn't currently deal with markup nested inside a p tag, e.g.
# the following:
# <ns2:p ns2:list-level="1" ns2:style="SFWPParagraphStyle-990">Pen knife photo:
# herzogbr <ns2:link
# href="http://www.flickr.com/photos/80516279@N00/2274372747/"><ns2:span
# ns2:style="SFWPCharacterStyle-158"
# >http://www.flickr.com/photos/80516279@N00/2274372747/</ns2:span></ns2:link>
# </ns2:p>
import zipfile
from xml.etree import ElementTree as ET
def extract_keynote_slide_text(fp):
z = zipfile.ZipFile(fp, 'r')
root = ET.fromstring(z.read('index.apxl'))
slides = root.find(
'{http://developer.apple.com/namespaces/keynote2}slide-list'
)
slide_texts = []
for slide in slides:
slide_text = []
ps = slide.findall('.//{http://developer.apple.com/namespaces/sf}p')
for p in ps:
text = p.text
if text is None or not text.strip():
continue
slide_text.append(text)
slide_texts.append(slide_text)
return slide_texts
@stefanw

stefanw commented Apr 24, 2010

Copy link
Copy Markdown

Maybe a TextIterator helps with the nested markup, something like:
" ".join([t.strip() for t in ET.ElementTextIterator(ps)])

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment