Skip to content

Instantly share code, notes, and snippets.

@jrunning
Last active August 29, 2015 14:08
Show Gist options
  • Select an option

  • Save jrunning/56b76e7580bbe2e9a07b to your computer and use it in GitHub Desktop.

Select an option

Save jrunning/56b76e7580bbe2e9a07b to your computer and use it in GitHub Desktop.

Okay, let's work through this iteratively, starting with your example:

str = "2:README:19:string:Hello world!spec.rb:20:string:describe RBFS1:rbfs:4:0:0:"
entries = {} # No entries yet!

The very first thing we need to know is how many files there are, and we know we know that's the number before the first ::

num_entries, rest = str.split(':', 2)
num_entries = Integer(num_entries)
# num_entries is now 2
# rest is now "README:19:string:Hello world!spec.rb:20:string:describe RBFS1:rbfs:4:0:0:"

The second argument to split says "I only want 2 pieces," so it stops splitting after the first :.) We use Integer(n) instead of n.to_i because it's stricter. (to_i will convert "10xyz" to 10; Integer will raise an error, which is what we want here.)

Now we know we have two files. We don't know anything else yet, but what's left of our string is this:

README:19:string:Hello world!spec.rb:20:string:describe RBFS1:rbfs:4:0:0:

The next thing we can get is the name and length of the first file.

name, len, rest = rest.split(':', 3)
len = Integer(len.to_i)
# name = "README"
# len  = 19
# rest = "string:Hello world!spec.rb:20:string:describe RBFS1:rbfs:4:0:0:"

Cool, now we have the name and length of the first file, so we can get its content:

content = rest.slice!(0, len)
# content = "string:Hello world!"
# rest = "spec.rb:20:string:describe RBFS1:rbfs:4:0:0:"
entries[name] = content
# entries = { "README" => "string:Hello world!" }

We used rest.slice! which modifies removes len characters from the front of the string and returns them, so content is just what we want (string:Hello world!) and rest is everything that was after it. Then we added it to entries Hash. One file down, one to go!

For the second file, we do the exact same thing:

name, len, rest = rest.split(':', 3)
len = Integer(len)
# name = "spec.rb"
# len = 20
# rest = "string:describe RBFS1:rbfs:4:0:0:"

content = rest.slice!(0, len)
# content = "string:describe RBFS"
# rest  = "1:rbfs:4:0:0:"
entries[name] = content
# entries = { "README" => "string:Hello world!",
#             "spec.rb" => "string:describe RBFS" }

Since we do the exact same thing twice, obviously we should do this in a loop! But before we write that, we need to get organized. So far we have two discrete steps: First, get the number of files. Second, get those files' contents. We also know we'll need to get the number of directories and the directories. We'll take a guess at how this'll look:

def parse(serialized)
  files, rest = parse_files(serialized)
  # `files` will be a Hash of file names and their contents and `rest` will be
  # the part of the string we haven't serialized yet
  directories, rest = parse_directories(rest)
  # `directories` will be a Hash of directory names and their contents

  files.merge(directories)
end

def parse_files(serialized)
  # Get the number of files from the beginning of the string
  num_entries, rest = str.split(':', 2)
  num_entries = Integer(num_entries)
  entries = {}

  # `rest` now starts with the first file (e.g. "README:19:...")
  num_entries.times do
    name, len, rest = rest.split(':', 3) # get the file name and length
    len = Integer(len)
    
    content = rest.slice!(0, len) # get the file contents from the beginning of the string
    entries[name] = content # add it to the hash
  end

  [ entries, rest ]
end

def parse_directories(serialized)
  # TBD...
end

That parse_files method is a bit long for my taste, though, so how about we split it up?

def parse_files(serialized)
  # Get the number of files from the beginning of the string
  num_entries, rest = str.split(':', 2)
  num_entries = Integer(num_entries)
  entries = {}

  # `rest` now starts with the first file (e.g. "README:19:...")
  num_entries.times do
    name, content, rest = parse_file(rest)
    entries[name] = content # add it to the hash
  end

  [ entries, rest ]
end

def parse_file(serialized)
  name, len, rest = serialized.split(':', 3) # get the name and length of the file
  len = Integer(len)

  content = rest.slice!(0, len) # use the length to get its contents
  [ name, content, rest ]
end

Clean!

Now, I'm going to give you a big spoiler: Since the serialization format is reasonably well-designed, we don't actually need a parse_directories method, because it would do exactly the same thing as parse_files. The only difference is that after this line:

name, content, rest = parse_file(rest)

...we want to do something different if we're parsing directories instead of files. In particular, we want to call parse(content), which will do all of this over again on the directory's contents. Since it's pulling double-duty now, we should probably change it's name to something more general like parse_entries, and we also need to give it another argument to tell it when to do that recursion.

def parse(serialized)
files, rest = parse_entries(serialized)
directories, = parse_entries(rest, true)
files.merge(directories)
end
def parse_entries(serialized, directories=false)
num_entries, rest = serialized.split(':', 2)
num_entries = Integer(num_entries)
entries = {}
num_entries.times do
name, content, rest = parse_entry(rest)
entries[name] = directories ? parse(content) : content
end
[ entries, rest ]
end
def parse_entry(serialized)
name, len, rest = serialized.split(':', 3)
len = Integer(len)
content = rest.slice!(0, len)
[ name, content, rest ]
end
p parse("2:README:19:string:Hello world!spec.rb:20:string:describe RBFS1:rbfs:4:0:0:")
# => { "README" => "string:Hello world!",
# "spec.rb" => "string:describe RBFS",
# "rbfs" => {}
# }
p parse("0:1:directory1:40:0:1:directory2:22:1:README:9:number:420:")
# => { "directory1" => {
# "directory2" => {
# "README" => "number:42"
# }
# }
# }
p parse("1:file0.txt:10:string:Foo2:directory1:51:2:file11.txt:11:string:Barrfile999.txt:8:number:90:directory2:30:1:file22.txt:12:string:Bazzz0:")
# => { "file0.txt" => "string:Foo",
# "directory1" => {
# "file11.txt" => "string:Barr",
# "file999.txt" => "number:9"
# },
# "directory2" => {
# "file22.txt" => "string:Bazzz"
# }
# }
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment