Skip to content

Instantly share code, notes, and snippets.

View ResidentMario's full-sized avatar

Aleksey Bilogur ResidentMario

View GitHub Profile
@ResidentMario
ResidentMario / coding_challenges.md
Created September 5, 2019 21:59
Some philosophy

Some thoughts on a part of coding challenges that I find difficult

...and what it means for my life more broadly.

After a couple of days of struggling with introductory Leetcode exercises, it's pretty clear that there is something fundamental to coding challenges that doesn't come to me naturally.

Doing well on a coding challenge means implementing a method that solves for all possible cases of the given problem with the given constraints as written. This means that all "normal cases" and all "edge cases" must be discovered and considered simultaneously—in advance of the actual algorithm design.

This differs from my typical workflow when implementing a new algorithm. My typical workflow is iterative. I define a narrowly scoped "normal case" version of the problem, and write an algorithm that solves that narrow problem. I then relax the constraints of the problem one step at a time, modifying the algorithm as needed to account for the additional cases.

@ResidentMario
ResidentMario / README.md
Last active July 27, 2019 13:40
Building your own MTA train arrivals dataset: a how-to

Note: this Gist has been superceded by the tutorial section of the gtfs_tripify documentation.

Interested in New York City transit? Want to learn more reasons why your particular train commute is good or bad? This gist will show you how to roll your own daily MTA train arrival dataset using Python. The result can then be used to explore questions about train service that schedule data alone couldn't answer.

Building a daily roll-up

To begin, visit the MTA GTFS-RT Archive at http://web.mta.info/developers/data/archives.html:

@ResidentMario
ResidentMario / train.py
Last active April 28, 2019 03:14
A simple demo model used in the fahr quickstart documentation.
import numpy as np
from keras.models import Sequential
from keras.layers import Dense, Activation
from keras.optimizers import SGD
# generate dummy data
def sample_threeclass(n, ratio=0.8):
np.random.seed(42)
y_0 = np.random.randint(2, size=(n, 1))
switch = (np.random.random(size=(n, 1)) <= ratio)
# define the model
# init a new model with 192x192 CNN layers
# padding='same' will downsample to 96x96
# # which is the expected input size for pretrain
model = Sequential()
model.add(Conv2D(64, kernel_size=(3, 3), input_shape=(192, 192, 3), activation='relu', padding='same'))
model.add(Conv2D(64, kernel_size=(3, 3), activation='relu', padding='same'))
model.add(MaxPooling2D(pool_size=(2, 2)))
model = Sequential()
model.add(Conv2D(64, kernel_size=(3, 3), input_shape=(96, 96, 3), activation='relu', padding='same'))
model.add(Conv2D(64, kernel_size=(3, 3), input_shape=(96, 96, 3), activation='relu', padding='same'))
model.add(MaxPooling2D(pool_size=(2, 2)))
# load the pretrained model
prior = load_model('resnet48_128/model-48.h5')
# add all but the first two layers of VGG16 to the new model
# strip the input layer out, this is now 96x96
#
# Step 1: download and format the data.
#
# I use the Quilt T4 Python API to do this. `t4` adds a lot of features on top
# of Amazon S3 buckets that make them more useful to data scientists. But you
# can also use either raw `boto3` or the `aws` CLI package, if so inclined.
#
import numpy as np
import pandas as pd
# fit the model
model.fit_generator(
train_generator,
steps_per_epoch=len(train_generator.filenames) // batch_size,
epochs=20,
validation_data=validation_generator,
validation_steps=len(train_generator.filenames) // batch_size,
callbacks=[
EarlyStopping(patience=3, restore_best_weights=True),
ReduceLROnPlateau(patience=2)
prior = keras.applications.VGG16(
include_top=False,
weights='imagenet',
input_shape=(48, 48, 3)
)
model = Sequential()
model.add(prior)
model.add(Flatten())
model.add(Dense(256, activation='relu', name='Dense_Intermediate'))
import t4
import pandas as pd
# download the data
img_dir = 'images_cropped/'
metadata_filepath = 'X_meta.csv'
open_fruits = t4.Package.browse('quilt/open_fruit', 's3://quilt-example')
open_fruits['training_data/X_meta.csv'].fetch(metadata_filepath)
open_fruits['images_cropped'].fetch(img_dir)