Web & product analytics · Quick take
The right algorithm isn’t always the newest
How do you group similar website journeys? I used Levenshtein edit distance to compare paths by their differences, choosing an approach that fit the question.
When is a familiar algorithm the better choice?
I think a well-defined problem and a well-understood algorithm can get you further than chasing the latest method. You need to know what a useful answer looks like before deciding how sophisticated the approach needs to be.
I used clustering based on Levenshtein edit distance for website pathing rather than an HDBSCAN approach. Levenshtein gives you a concrete way to compare sequences: how many insertions, deletions, or substitutions separate them? It makes the meaning of similarity something you can explain and examine.[1]
The useful discipline is to define the question, decide which differences in the data matter, and choose a method whose assumptions fit. Then check whether the result helps with the decision. That means examining what the groups have in common, where they break down, and why.
A note on the methods
Levenshtein is a distance measure; HDBSCAN is a clustering algorithm. HDBSCAN can itself use precomputed Levenshtein distances. The choice here was how to compare website paths and use those distances for clustering.[1][2]
It’s easy to chase the state of the art and overcomplicate the problem. Sometimes the right solution isn’t especially exciting. It just fits.
Sources & further reading
- Levenshtein distance
NIST Dictionary of Algorithms and Data Structures · Accessed
Technical definition of edit distance and its insertion, deletion, and substitution operations.
- Basic Usage of HDBSCAN: Distance matrices
HDBSCAN documentation · Accessed
Explains precomputed distance matrices, including Levenshtein distances, as input to HDBSCAN.