Python Urllib Module
urllib is a package that collects several modules for working with URLs.
urllib.requestfor opening and reading URLsurllib.errorcontaining the exceptions raised byurllib.requesturllib.parsefor parsing URLsurllib.robotparserfor parsingrobots.txtfiles
Importing the Module
import urllib.request
import urllib.parse
Fetching URLs
The urllib.request module defines functions and classes which help in opening URLs (mostly HTTP).
import urllib.request
with urllib.request.urlopen('http://www.python.org/') as response:
html = response.read()
print(html[:100])
Parsing URLs
The urllib.parse module defines a standard interface to break Uniform Resource Locator (URL) strings up in components.
from urllib.parse import urlparse
o = urlparse('http://www.cwi.nl:80/%7Eguido/Python.html')
print(o)
# Output: ParseResult(scheme='http', netloc='www.cwi.nl:80', path='/%7Eguido/Python.html', params='', query='', fragment='')
print(o.scheme)
print(o.port)
print(o.geturl())
Encoding Parameters
urllib.parse.urlencode() is useful for creating query strings.
import urllib.parse
import urllib.request
url = 'https://www.example.com'
values = {'name': 'Michael Foord',
'location': 'Northampton',
'language': 'Python' }
data = urllib.parse.urlencode(values)
data = data.encode('ascii') # data should be bytes
req = urllib.request.Request(url, data)
with urllib.request.urlopen(req) as response:
the_page = response.read()
Handling Errors
urllib.error contains exceptions raised by urllib.request.
import urllib.request
import urllib.error
try:
urllib.request.urlopen('http://www.python.org/fish.html')
except urllib.error.HTTPError as e:
print('The server couldn\'t fulfill the request.')
print('Error code: ', e.code)
except urllib.error.URLError as e:
print('We failed to reach a server.')
print('Reason: ', e.reason)