Python: concurrent.futures
I had been wanting to try some new features of Python 3 for a while now. I recently found some free time and tried concurrent.futures. Concurrent.futures allows us to write code that runs in parallel in Python without resorting to creating threads or forking processes.

Photo by Alosh Bennett
I had been wanting to try some new features of Python 3 for a while now.
I recently found some free time and tried concurrent.futures.
Concurrent.futures allows us to write code that runs in parallel in
Python without resorting to creating threads or forking processes.
We will use the following problem to write some code — write a script that downloads images over the internet to a local directory.
Our test setting will use the built-in http server that comes bundled with Python. This will make sure that the numbers we see are not influenced by network latency and we also don’t bombard some random server with 1000's of requests.
$ cd Pictures
$ python -m SimpleHTTPServer
Serving HTTP on 0.0.0.0 port 8000 ...
127.0.0.1 - - [03/Aug/2014 19:37:37] "GET /shakuni.png HTTP/1.1" 200 -We will run the script on my Intel 4 Core laptop running Ubuntu 14.04 with Python 3.4.
Processor 4x Intel(R) Core(TM) i7-3517U CPU @ 1.90GHz
Memory 3905MB (2645MB used)
Operating System Ubuntu 14.04.1 LTSThe first step is to establish some base benchmark to see how much time a sequential program would require. This is what the script looks like -
import requests
import time
URL = "http://127.0.0.1:8000/shakuni.png"
def download_ntimes(n):
for i in range(n):
requests.get(URL)
def measure(n):
start = time.clock()
download_ntimes(n)
end = time.clock()
print("Time to download a image %s times: %s CPU units" % (n, end-start))
measure(10000)
measure(20000)
measure(30000)
measure(40000)The times for this simple sequential download of images are -
$ python3.4 picit-seq.py
Time to download a image 10000 times: 79.61527 CPU units
Time to download a image 20000 times: 160.204535 CPU units
Time to download a image 30000 times: 240.62441900000002 CPU units
Time to download a image 40000 times: 317.55276999999995 CPU unitsNote that we have added a sanity check to verify that there is a linear increase in the time taken. So if 10k requests takes 80 seconds, 20k requests will require 160 seconds and so on.
Now let us look at a version that uses concurrent.futures. It is surprisingly easy to change the sequential code to run parallelly -
import requests
import time
import concurrent.futures
URL = "http://127.0.0.1:8000/shakuni.png"
def download_ntimes(n):
with concurrent.futures.ProcessPoolExecutor() as executor:
for i in range(n):
executor.map(requests.get, [URL])
def measure(n):
start = time.clock()
download_ntimes(n)
end = time.clock()
print("Time to download a image %s times: %s CPU units" % (n, end-start))
measure(10000)
measure(20000)
measure(30000)
measure(40000)Only the download_ntimes method has changed to download images asynchronously
using ProcessPoolExecutor.
From the documentation -
The ProcessPoolExecutor class uses a pool of processes to execute calls asynchronously. ProcessPoolExecutor uses the multiprocessing module, which allows it to side-step the global interpreter lock
Note also that we don’t specify how many workers to use. That is because
if max_workers is None or not given, it will default to the number of processors on the machine.
The download times now are -
$ python3.4 picit.py
Time to download a image 10000 times: 31.90206 CPU units
Time to download a image 20000 times: 57.468886 CPU units
Time to download a image 30000 times: 78.56723 CPU units
Time to download a image 40000 times: 98.28945499999998 CPU unitsWe see that the download times are down by a factor of 3 (approximately). Theoretically, we should have seen a 4x speed up because of the 4 cores my machine has. The difference is probably due to book keeping that concurrent.futures has to do.
So the question is — how many of us have the luxury of using Python 3 in real life 😉. But on the other hand, if you are writing standalone scripts today and not using Python 3, you are missing out on the fun!
Happy to hear your thoughts — disagreements first, applause ok.
