Skip to content

Instantly share code, notes, and snippets.

@simonw
Created April 11, 2010 20:01
Show Gist options
  • Select an option

  • Save simonw/363022 to your computer and use it in GitHub Desktop.

Select an option

Save simonw/363022 to your computer and use it in GitHub Desktop.
"""
A ridiculously simple term extractor - extracts sequences of one or more Title Case
Words, e.g. "Gordon Brown is the Prime Minister" would extract "Gordon Brown" and
"Prime Minister".
"""
import re
# Stopwords from http://snowball.tartarus.org/algorithms/english/stop.txt
stop_words = set(map(lambda s: s.strip(),
"""
i me my myself we us our ours ourselves you your yours yourself
yourselves he him his himself she her hers herself it its itself
they them their theirs themselves what which who whom this that these
those am is are was were be been being have has had having do does did
doing would should could ought a an the and but if or because as until
while of at by for with about against between into through during before
after above below to from up down in out on off over under again further
then once here there when where why how all any both each few more most
other some such no nor not only own same so than too very
""".split()
))
word_re = re.compile(r'[^\s,\.?!]+')
punctuation_re = re.compile(r'[,\.?!]')
term_re = re.compile(r'(?:[A-Z][^\s]+(?: |$))+')
def extract_terms(s):
sentences = [l.strip() for l in punctuation_re.split(s) if l.strip()]
terms = []
for sentence in sentences:
extracted = term_re.findall(sentence)
for term in extracted:
term = term.strip(' "\':')
if term.endswith("'s"):
term = term[:-2]
if term.lower() not in stop_words:
terms.append(term)
return terms
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment