Skip to content

Instantly share code, notes, and snippets.

@dmesquita
Created October 21, 2019 11:50
Show Gist options
  • Select an option

  • Save dmesquita/f7695e4c8db0e3c1b9788165edebf90c to your computer and use it in GitHub Desktop.

Select an option

Save dmesquita/f7695e4c8db0e3c1b9788165edebf90c to your computer and use it in GitHub Desktop.
from sklearn.datasets import fetch_20newsgroups
from sklearn.model_selection import RandomizedSearchCV
from classify_documents_model.pipeline import text_clf
categories = ['talk.politics.guns',
'talk.politics.mideast',
'talk.politics.misc']
twenty_train = fetch_20newsgroups(subset='train',
categories=categories, shuffle=True, random_state=42)
parameters = {
'tokens__lemmatization': (True, False),
'tokens__remove_stopwords': (True, False)
}
X = twenty_train.data[:400]
y = twenty_train.target[:400]
clf = RandomizedSearchCV(text_clf, parameters, cv=5, iid=False,
scoring="precision_macro",
n_iter=3, n_jobs=-1, verbose=3)
clf.fit(X, y)
print("Best params: {}".format(clf.best_params_))
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment