Group similar documents by embedding or TF-IDF cosine similarity.
Requires uv.
Install the docs-clustering-cli command from this repo:
# Minimal deps supporting `--method tfidf` only:
uv tool install .
# Optional heavy stack for the `--method st` embedding method:
uv tool install ".[st]"or directly from an upstream git remote:
uv tool install git+https://github.com/redhat-performance/docs-clustering
uv tool install "git+https://github.com/redhat-performance/docs-clustering[st]"For an ad-hoc run without installing:
uvx --from "git+https://github.com/redhat-performance/docs-clustering[st]" docs-clustering-cli --helpdocs-clustering-cli --data-dir data --method tfidf --out report.json
docs-clustering-cli --data-dir data --method st --model sentence-transformers/all-MiniLM-L6-v2 --out report.json
docs-clustering-cli --data-json docs.json --method multiset--data-json accepts a JSON file mapping document IDs to their text:
{
"123": "some text of first document",
"456": "another document"
}Prints per-file rankings, ranked similar pairs, and clusters of document picked
by similarity threshold. With --out, the exact same data is dumped as a
structured JSON file, which is meant to simplify integration with other tools
that consume the output to categorize errors:
{
"method": "tfidf",
"model": "-",
"threshold": 0.3,
"rankings": {
"fo1": [["fo2", 0.7041]],
"un1": [["un2", 0.8161], ["un3s", 0.7041]]
},
"pairs": [["un1", "un2", 0.8161], ["fo1", "fo2", 0.7041]],
"clusters": [["fo1", "fo2"], ["un1", "un2", "un3s"]]
}rankings: per-document list of[other-id, similarity], most similar first (respects--top-k)pairs: global[id, other-id, similarity]list, highest similarity firstclusters: the grouping your categorizer should consume
With no --out, nothing is written to disk; the report only goes to stdout.
| Flag | Default | Meaning |
|---|---|---|
--data-dir |
data |
Directory to scan for *.log files (mutually exclusive with --data-json) |
--data-json |
- | JSON file mapping document IDs to text (mutually exclusive with --data-dir) |
--method |
tfdif |
tfidf, st, setjacc, or multiset |
--model |
all-MiniLM-L6-v2 | Sentence-transformer model name (st only) |
--threshold |
0.6 (st) / 0.3 (tfidf) | Minimum similarity for clustering |
--top-k |
all | Limit per-file ranking rows |
--out |
(none) | JSON report file; omit to only print to stdout (no file written) |
tfidf— TF-IDF cosine similarity (no ML deps). Words are weighted by term frequency × inverse document frequency, so corpus-common words are down-weighted and rare/distinctive ones up-weighted; similarity is the cosine between the weighted term vectors. Pure lexical.st— Sentence-BERT (SBERT) embeddings + cosine similarity (uses the optionalstextra dependencies). Each document is embedded into a dense vector by a sentence-transformers model (all-MiniLM-L6-v2by default) and similarity is the cosine of the vectors. Captures semantics (synonyms, paraphrase), but only sees the first ~256 tokens (depends on a model) of a document.setjacc— Jaccard similarity coefficient (binary). Each document is a bag of unique words; similarity is the sum of the per-word minimum counts over the sum of the maximum counts. Word counts are reduced to present/absent per word. Suited to documents where repetition carries no meaning.multiset— Multiset Jaccard similarity (Ruzicka coefficient). Likesetjaccbut counts are weights per word. Suited to documents where repetition carries no meaning. Count-aware: repeating a term N times vs once lowers similarity, so repetition differences (e.g. 1 error row vs 4 identical ones) are visible.
The table below compares every method / tested model on
tests/data/errors-example.json, a file with 5 documents:
un1,un2,un3sdescribe the same underlying error (failed to pull therhacs-roxctlimage, Pod creation failed).un3sis significantly shorter: the repeatedUnknown errorblock appears once instead of four times.fo1,fo2are also the same issue (repo fork blocked, account blocked), but a different one from theun*group.
Ideally any method clusters {un1, un2, un3s} and {fo1, fo2} into two
separate groups and nothing else. Metrics per method/model:
margin= mean within-cluster similarity − mean cross-cluster similarity (bigger is a clearer separation)gap= lowest within-cluster − highest cross-cluster similarity (positive means a threshold exists that clusters perfectly; bigger is more headroom)default thr= does the output produce the correct two clusters with the tool's default--threshold(0.3 for lexical methods, 0.6 forst)
| Method (/ Model) | Intra mean | Inter mean | Margin | Min intra | Max inter | Gap | Default thr |
|---|---|---|---|---|---|---|---|
| tfidf | 0.8161 | 0.0444 | 0.7717 | 0.7041 | 0.0457 | 0.6584 | correct (0.3) |
| setjacc | 0.8130 | 0.0408 | 0.7722 | 0.7157 | 0.0444 | 0.6713 | correct (0.3) |
| multiset | 0.5635 | 0.0248 | 0.5387 | 0.2606 | 0.0357 | 0.2249 | correct (0.3) / wrong (0.6) |
| st / all-mpnet-base-v2 (438 MB) | 0.8940 | 0.3456 | 0.5484 | 0.7779 | 0.3681 | 0.4098 | correct (0.6) |
| st / paraphrase-MiniLM-L6-v2 (91 MB) | 0.9348 | 0.4974 | 0.4374 | 0.8925 | 0.5164 | 0.3761 | correct (0.6) |
| st / all-MiniLM-L6-v2 (91 MB) | 0.8134 | 0.3160 | 0.4974 | 0.7461 | 0.3870 | 0.3591 | correct (0.6) |
| st / bge-base-en-v1.5 (438 MB) | 0.9636 | 0.6263 | 0.3373 | 0.9362 | 0.6434 | 0.2928 | wrong — merges (0.6) |
| st / bge-small-en-v1.5 (134 MB) | 0.9625 | 0.7185 | 0.2440 | 0.9379 | 0.7430 | 0.1949 | wrong — merges (0.6) |
| st / gte-small (67 MB) | 0.9803 | 0.8545 | 0.1258 | 0.9661 | 0.8655 | 0.1006 | wrong — merges (0.6) |
Notes:
- Lexical methods (
tfidf,setjacc) are best here because the two issues share almost no vocabulary, so cross-cluster similarity is near zero. - Among sentence-transformer models
all-mpnet-base-v2is the best: high within-cluster similarity and a comfortable margin below the 0.6 default threshold.bge-*andgte-smallare score-inflated — everything looks similar — so the default threshold merges the two issues (they need a much higher threshold, ≥ 0.8). multisetis count-aware: it seesun3s(one repeated error block) as different fromun1/un2(four blocks), fragmenting theun*cluster at--threshold 0.6. At the lexical default 0.3 it still groups all three correctly.
Reproduce e.g. with:
docs-clustering-cli --data-json tests/data/errors-example.json --method st --model sentence-transformers/all-mpnet-base-v2normalize() only does generic cleanup: URLs, common date/time formats,
long hex strings, and runs of 5+ digits. Everything else is your job. A
cookbook for preparing documents before feeding them to this tool:
- Replace unique IDs with placeholders. Component names, namespaces,
build/run slugs, hostnames, ticket IDs →
<COMPONENT>,<RUN>, etc. Otherwise every method clusters by "same component" instead of "same issue". - Keep the error, drop the location. Keep messages, reasons, status codes. Generalize or remove where/which instance it happened, unless that's what you want to group by.
- Don't dedupe repeated blocks if the count matters. "Failed once" vs.
"retried 4 times" should look different. Either keep the repeats and use
--method multiset(count-aware), or encode the count as a token, e.g.retries=4. - Strip your own tooling headers/preambles (log wrappers, custom prefixes) before writing the JSON/log files — this tool won't recognize them.
make bootstrap # one time venv setup and pre-commit installation
make check-all # run linters
make test # run tests