Problem solution · Python

Similar String Groups

Similar String Groups: a Python solution using disjoint set union. Learn the idea, check the complexity, and read the full code, with credit to Kamyu LeetCode Solutions.

Technique
Disjoint set union
Source
Kamyu LeetCode Solutions
Length
65 lines
Start with the idea.

Try the problem first. If you get stuck, read the approach below, then write your own solution. The full code is at the bottom.

Approach

Disjoint set union

For Similar String Groups, the implementation maintains connected components and merges them as relationships are processed.

  1. Give each element a component representative.
  2. Merge representatives when a connection is accepted.
  3. Answer connectivity or component queries from the compressed representatives.

Code notes

  • 65 lines of Python from the credited upstream file similar-string-groups.py.
  • The implementation visibly relies on sequence storage, hash lookup, ordered lookup.
  • 1 loop block detected.

Complexity

Account for every find and union operation; with path compression and ranked merging, the amortized cost is nearly constant per operation.

Check the problem constraints before deciding whether this complexity will pass.

Source

Code and credit

This code comes from Kamyu LeetCode Solutions by kamyu104 and is used under the MIT licence.

Full codeSimilar String Groups · PythonPython
Use this to learn the idea, then write your own version.
# Time:  O(n^2 * l) ~ O(n * l^4)# Space: O(n) ~ O(n * l^3) import collectionsimport itertools  class UnionFind(object):    def __init__(self, n):        self.set = range(n)        self.__size = n     def find_set(self, x):        if self.set[x] != x:            self.set[x] = self.find_set(self.set[x])  # path compression.        return self.set[x]     def union_set(self, x, y):        x_root, y_root = map(self.find_set, (x, y))        if x_root == y_root:            return False        self.set[min(x_root, y_root)] = max(x_root, y_root)        self.__size -= 1        return True     def size(self):        return self.__size  class Solution(object):    def numSimilarGroups(self, A):        def isSimilar(a, b):            diff = 0            for x, y in itertools.izip(a, b):                if x != y:                    diff += 1                    if diff > 2:                        return False            return diff == 2         N, L = len(A), len(A[0])        union_find = UnionFind(N)        if N < L*L:            for (i1, word1), (i2, word2) in \                    itertools.combinations(enumerate(A), 2):                if isSimilar(word1, word2):                    union_find.union_set(i1, i2)        else:            buckets = collections.defaultdict(list)            lookup = set()            for i in xrange(len(A)):                word = list(A[i])                if A[i] not in lookup:                    buckets[A[i]].append(i)                    lookup.add(A[i])                for j1, j2 in itertools.combinations(xrange(L), 2):                    word[j1], word[j2] = word[j2], word[j1]                    buckets["".join(word)].append(i)                    word[j1], word[j2] = word[j2], word[j1]            for word in A:  # Time:  O(n * l^4)                for i1, i2 in itertools.combinations(buckets[word], 2):                    union_find.union_set(i1, i2)        return union_find.size()  

Did this explanation save you time? I'm a Grade 11 student building this free library to make difficult algorithms easier to understand.

Buy me a coffee ↗