Skip to content

Instantly share code, notes, and snippets.

@pablodz
Created December 5, 2020 08:20
Show Gist options
  • Select an option

  • Save pablodz/6bb540eb94ad4d5fe70517e8bd742468 to your computer and use it in GitHub Desktop.

Select an option

Save pablodz/6bb540eb94ad4d5fe70517e8bd742468 to your computer and use it in GitHub Desktop.
#!/bin/bash
# NAME: pdflinkextractor
# AUTHOR: Glutanimate (http://askubuntu.com/users/81372/), 2013
# LICENSE: GNU GPL v2
# DEPENDENCIES: wget lynx
# DESCRIPTION: extracts PDF links from websites and dumps them to the stdout and as a textfile
# only works for links pointing to files with the ".pdf" extension
#
# USAGE: pdflinkextractor "www.website.com"
WEBSITE="$1"
echo "Getting link list..."
lynx -cache=0 -dump -listonly "$WEBSITE" | grep ".*\.pdf$" | awk '{print $2}' | tee pdflinks.txt
# OPTIONAL
#
# DOWNLOAD PDF FILES
#
echo "Downloading..."
xargs -n 1 curl -O < pdflinks.txt
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment