Problem solution · Python

Web Crawler Multithreaded

Web Crawler Multithreaded: a Python solution using breadth-first search. Learn the idea, check the complexity, and read the full code, with credit to Kamyu LeetCode Solutions.

Technique
Breadth-first search
Source
Kamyu LeetCode Solutions
Length
138 lines
Start with the idea.

Try the problem first. If you get stuck, read the approach below, then write your own solution. The full code is at the bottom.

Approach

Breadth-first search

For Web Crawler Multithreaded, the implementation explores reachable states in layers, which is the standard shape for unweighted shortest paths and minimum-step transitions.

  1. Model each valid configuration as a state and each legal move as an edge.
  2. Seed the queue with the starting state and mark it immediately.
  3. Expand each state once, recording distance or reachability for unseen neighbours.

Code notes

  • 138 lines of Python from the credited upstream file web-crawler-multithreaded.py.
  • The implementation visibly relies on sequence storage, ordered lookup, work queue.
  • No explicit loop blocks detected.

Complexity

Verify that each state and transition is processed only a bounded number of times; that determines the traversal cost.

Check the problem constraints before deciding whether this complexity will pass.

Source

Code and credit

This code comes from Kamyu LeetCode Solutions by kamyu104 and is used under the MIT licence.

Full codeWeb Crawler Multithreaded · PythonPython
Use this to learn the idea, then write your own version.
# Time:  O(|V| + |E|)# Space: O(|V|) import threadingimport Queue  # """# This is HtmlParser's API interface.# You should not implement it, or speculate about its implementation# """class HtmlParser(object):   def getUrls(self, url):       """       :type url: str       :rtype List[str]       """       pass  class Solution(object):    NUMBER_OF_WORKERS = 8        def __init__(self):        self.__cv = threading.Condition()        self.__q = Queue.Queue()     def crawl(self, startUrl, htmlParser):        """        :type startUrl: str        :type htmlParser: HtmlParser        :rtype: List[str]        """        SCHEME = "http://"        def hostname(url):            pos = url.find('/', len(SCHEME))            if pos == -1:                return url            return url[:pos]         def worker(htmlParser, lookup):            while True:                from_url = self.__q.get()                if from_url is None:                    break                name = hostname(from_url)                for to_url in htmlParser.getUrls(from_url):                    if name != hostname(to_url):                        continue                    with self.__cv:                        if to_url not in lookup:                           lookup.add(to_url)                           self.__q.put(to_url)                self.__q.task_done()         workers = []        self.__q = Queue.Queue()        self.__q.put(startUrl)        lookup = set([startUrl])        for i in xrange(self.NUMBER_OF_WORKERS):            t = threading.Thread(target=worker, args=(htmlParser, lookup))            t.start()            workers.append(t)        self.__q.join()        for t in workers:            self.__q.put(None)        for t in workers:            t.join()        return list(lookup)  # Time:  O(|V| + |E|)# Space: O(|V|)import threadingimport collections  class Solution2(object):    NUMBER_OF_WORKERS = 8        def __init__(self):        self.__cv = threading.Condition()        self.__q = collections.deque()        self.__working_count = 0     def crawl(self, startUrl, htmlParser):        """        :type startUrl: str        :type htmlParser: HtmlParser        :rtype: List[str]        """        SCHEME = "http://"        def hostname(url):            pos = url.find('/', len(SCHEME))            if pos == -1:                return url            return url[:pos]         def worker(htmlParser, lookup):            while True:                with self.__cv:                    while not self.__q:                        self.__cv.wait()                    from_url = self.__q.popleft()                    if from_url is None:                        break                    self.__working_count += 1                name = hostname(from_url)                for to_url in htmlParser.getUrls(from_url):                    if name != hostname(to_url):                        continue                    with self.__cv:                        if to_url not in lookup:                           lookup.add(to_url)                           self.__q.append(to_url)                           self.__cv.notifyAll()                with self.__cv:                    self.__working_count -= 1                    if not self.__q and not self.__working_count:                        self.__cv.notifyAll()         workers = []        self.__q = collections.deque([startUrl])        lookup = set([startUrl])        for i in xrange(self.NUMBER_OF_WORKERS):            t = threading.Thread(target=worker, args=(htmlParser, lookup))            t.start()            workers.append(t)        with self.__cv:            while self.__q or self.__working_count:                self.__cv.wait()            for i in xrange(self.NUMBER_OF_WORKERS):                self.__q.append(None)            self.__cv.notifyAll()        for t in workers:            t.join()        return list(lookup) 

Did this explanation save you time? I'm a Grade 11 student building this free library to make difficult algorithms easier to understand.

Buy me a coffee ↗