Skip to content

Instantly share code, notes, and snippets.

@fengchangfight
Last active September 15, 2017 14:31
Show Gist options
  • Select an option

  • Save fengchangfight/521ff9ff87400fe21f7ab7e9a8012062 to your computer and use it in GitHub Desktop.

Select an option

Save fengchangfight/521ff9ff87400fe21f7ab7e9a8012062 to your computer and use it in GitHub Desktop.
hadoop assignment2
able,991
about,11
burger,15
actor,22
Jan-01 able,5
Feb-02 about,3
Mar-03 about,8
Apr-04 able,13
Feb-22 actor,3
Feb-23 burger,5
Mar-08 burger,2
Dec-15 able,100
#!/usr/bin/env python
import sys
# --------------------------------------------------------------------------
#This mapper code will input a <date word, value> input file, and move date into
# the value field for output
#
# Note, this program is written in a simple style and does not full advantage of Python
# data structures,but I believe it is more readable
#
# Note, there is NO error checking of the input, it is assumed to be correct
# meaning no extra spaces, missing inputs or counts,etc..
#
# See # see https://docs.python.org/2/tutorial/index.html for details and python tutorials
#
# --------------------------------------------------------------------------
for line in sys.stdin:
line = line.strip() #strip out carriage return
key_value = line.split(",") #split line, into key and value, returns a list
key_in = key_value[0].split(" ") #key is first item in list
value_in = key_value[1] #value is 2nd item
#print key_in
if len(key_in)>=2: #if this entry has <date word> in key
date = key_in[0] #now get date from key field
word = key_in[1]
value_out = date+" "+value_in #concatenate date, blank, and value_in
print( '%s\t%s' % (word, value_out) ) #print a string, tab, and string
else: #key is only <word> so just pass it through
print( '%s\t%s' % (key_in[0], value_in) ) #print a string tab and string
#Note that Hadoop expects a tab to separate key value
#but this program assumes the input file has a ',' separating key value
#!/usr/bin/env python
import sys
# --------------------------------------------------------------------------
#This reducer code will input a <word, value> input file, and join words together
# Note the input will come as a group of lines with same word (ie the key)
# As it reads words it will hold on to the value field
#
# It will keep track of current word and previous word, if word changes
# then it will perform the 'join' on the set of held values by merely printing out
# the word and values. In other words, there is no need to explicitly match keys b/c
# Hadoop has already put them sequentially in the input
#
# At the end it will perform the last join
#
#
# Note, there is NO error checking of the input, it is assumed to be correct, meaning
# it has word with correct and matching entries, no extra spaces, etc.
#
# see https://docs.python.org/2/tutorial/index.html for python tutorials
#
# San Diego Supercomputer Center copyright
# --------------------------------------------------------------------------
prev_word = " " #initialize previous word to blank string
months = ['Jan','Feb','Mar','Apr','Jun','Jul','Aug','Sep','Nov','Dec']
dates_to_output = [] #an empty list to hold dates for a given word
day_cnts_to_output = [] #an empty list of day counts for a given word
# see https://docs.python.org/2/tutorial/datastructures.html for list details
line_cnt = 0 #count input lines
for line in sys.stdin:
line = line.strip() #strip out carriage return
key_value = line.split('\t') #split line, into key and value, returns a list
line_cnt = line_cnt+1
#note: for simple debugging use print statements, ie:
curr_word = key_value[0] #key is first item in list, indexed by 0
value_in = key_value[1] #value is 2nd item
#-----------------------------------------------------
# Check if its a new word and not the first line
# (b/c for the first line the previous word is not applicable)
# if so then print out list of dates and counts
#----------------------------------------------------
if curr_word != prev_word:
# -----------------------
#now write out the join result, but not for the first line input
# -----------------------
if line_cnt>1:
for i in range(len(dates_to_output)): #loop thru dates, indexes start at 0
print('{0} {1} {2} {3}'.format(dates_to_output[i],prev_word,day_cnts_to_output[i],curr_word_total_cnt))
#now reset lists
dates_to_output =[]
day_cnts_to_output=[]
prev_word =curr_word #set up previous word for the next set of input lines
# ---------------------------------------------------------------
#whether or not the join result was written out,
# now process the curr word
#determine if its from file <word, total-count> or < word, date day-count>
# and build up list of dates, day counts, and the 1 total count
# ---------------------------------------------------------------
if (value_in[0:3] in months):
date_day =value_in.split() #split the value field into a date and day-cnt
#add date to lists of the value fields we are building
dates_to_output.append(date_day[0])
day_cnts_to_output.append(date_day[1])
else:
curr_word_total_cnt = value_in #if the value field was just the total count then its
#the first (and only) item in this list
# ---------------------------------------------------------------
#now write out the LAST join result
# ---------------------------------------------------------------
for i in range(len(dates_to_output)): #loop thru dates, indexes start at 0
print('{0} {1} {2} {3}'.format(dates_to_output[i],prev_word,day_cnts_to_output[i],curr_word_total_cnt))
@fengchangfight

Copy link
Copy Markdown
Author

cat join1_File*.txt | ./join1_mapper.py | sort | ./join1_reducer.py

@fengchangfight

Copy link
Copy Markdown
Author

hadoop jar /usr/lib/hadoop-mapreduce/hadoop-streaming.jar
-input /user/cloudera/input
-output /user/cloudera/output_join \
-mapper /home/cloudera/join1_mapper.py \
-reducer /home/cloudera/join1_reducer.py

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment