In this tutorial, I will teach you how to download PDF files from URLs using Python programming language. The complete script to download pdfs from website is given below.
We will make use of Beautiful Soup 4 and Requests libraries to build the functionality of downloading PDF files from URLs.
pip install requests
pip install bs4
# Import libraries import requests from bs4 import BeautifulSoup # URL from which pdfs to be downloaded url = "https://nanonets.com/blog/deep-learning-ocr/" # Requests URL and get response object response = requests.get(url) # Parse text obtained soup = BeautifulSoup(response.text, 'html.parser') # Find all hyperlinks present on webpage links = soup.find_all('a') i = 0 # From all links check for pdf link and # if present download file for link in links: if ('.pdf' in link.get('href', [])): i += 1 print("Downloading file: ", i) # Get response object for link response = requests.get(link.get('href')) # Write content in pdf file pdf = open("pdf"+str(i)+".pdf", 'wb') pdf.write(response.content) pdf.close() print("File ", i, " downloaded") print("All PDF files downloaded")
python code.py
We evaluated the performance of Llama 3.1 vs GPT-4 models on over 150 benchmark datasets…
The manufacturing industry is undergoing a significant transformation with the advent of Industrial IoT Solutions.…
If you're reading this, you must have heard the buzz about ChatGPT and its incredible…
How to Use ChatGPT in Cybersecurity If you're a cybersecurity geek, you've probably heard about…
Introduction In the dynamic world of cryptocurrencies, staying informed about the latest market trends is…
The Events Calendar Widgets for Elementor has become easiest solution for managing events on WordPress…