Skip to content

Instantly share code, notes, and snippets.

View simonw's full-sized avatar

Simon Willison simonw

View GitHub Profile
@simonw
simonw / sessions.json
Created January 21, 2019 23:52
SRCCON sessions from 2018 (just in case they get over-written for 2019) - from https://schedule.srccon.org/sessions.json
[
{
"day": "Thursday",
"description": "Get your badges and get some food (plus plenty of coffee), as you gear up for the first day of SRCCON!",
"everyone": "y",
"facilitators": "",
"facilitators_twitter": "",
"id": "thursday-breakfast",
"length": "",
"notepad": "",
csvs-to-sqlite https://candidates.democracyclub.org.uk/media/candidates-all.csv \
--table=candidates \
-c election \
-f name \
-f party_name \
-f post_label \
democracyclub.db
datasette publish heroku democracyclub.db \
--name="democracyclub-datasette" \
@simonw
simonw / README.md
Last active March 10, 2019 01:47
How I created dams.now.sh

How I created dams.now.sh

Try it out at https://dams.now.sh/ - see this Twitter thread for background.

I started by grabbing the URLs to every downloadable Excel spreadsheet.

I navigated to the "Downloads (Public)" link starting from https://nid-test.sec.usace.army.mil/ - then I ran this JavaScript in my browser's console to extract all of the URLs as a JSON blob.

console.log(JSON.stringify(

Array.from(

@simonw
simonw / readable_diff.py
Created March 13, 2019 05:03
The earlier version of what is now https://github.com/simonw/csv-diff
import csv
from dictdiffer import diff
def load_trees(filepath):
fp = csv.reader(open(filepath))
headings = next(fp)
rows = [dict(zip(headings, line)) for line in fp]
return {r["TreeID"]: r for r in rows}
@simonw
simonw / fetch_metadata_for_doc_ids.py
Created April 3, 2019 23:49
Fetch metadata from Google Drive API for a list of doc_ids (because their batch API is extremely difficult to figure out)
def fetch_metadata_for_doc_ids(doc_ids, oauth_token):
boundary = 'batch_boundary'
headers = {
'Authorization': 'Bearer {}'.format(oauth_token),
'Content-Type': 'multipart/mixed; boundary=%s' % boundary,
}
body = ''
for doc_id in doc_ids:
req = 'GET https://www.googleapis.com/drive/v3/files/{}?fields=*'.format(doc_id)
body += '--%s\n' % boundary
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
# For a sample Starlette app
from starlette.applications import Starlette
from starlette.responses import JSONResponse
import sys
import sqlite3
application = Starlette()
@application.route("/")
@simonw
simonw / CSV conf CSV schedule.ipynb
Created May 9, 2019 18:45
Code for scraping the CSVConf schedule. This is pretty messy - I wrote most of it on a plane with no internet connection so I had to get it working against the offline data I had accidentalyl cached.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
@simonw
simonw / example-Locations.xml
Last active July 21, 2019 06:19
Convert Locations.kml (pulled from an iPhone backup) to SQLite
<?xml version="1.0" encoding="utf-8"?>
<kml xmlns="http://www.opengis.net/kml/2.2">
<Document>
<Placemark>
<TimeStamp>
<when>2015-12-18T19:12:32</when>
</TimeStamp>
<name>2015-12-18 19:12:32 Source: WhatsApp</name>
<Point>
<coordinates>-0.120970480144024,51.510383605957</coordinates>
@simonw
simonw / installing-mysqldb.md
Created June 21, 2019 04:17
Installing MySQLdb mysqlclient with Homebrew and pip in OS X Mojave

Installing MySQLdb mysqlclient with Homebrew and pip in OS X Mojave

First I used Homebrew to install the MySQL server along with mysql_config, needed to compile the Python extension:

brew install mysql

I had to then use this incantation to get pip install mysqlclient to work:

LIBRARY_PATH=$LIBRARY_PATH:/usr/local/opt/openssl/lib/ pip install mysqlclient