Repository navigation
C/C++ repository collapses to 5 leaf nodes during dependency analysis. Silent fallback to whole-repository mode #75
Description
Activity
- changed the title
[-]C/C++ repository collapses to 3 leaf nodes during dependency analysis. Silent fallback to whole-repository mode[/-][+]C/C++ repository collapses to 5 leaf nodes during dependency analysis. Silent fallback to whole-repository mode[/+]on Jul 15, 2026 Thanks for the detailed analysis. The leaf-node selection logic only considered class, interface, and struct components, and included functions only when a repository contained none of those. In umoria, the small number of structs therefore caused all 786 free functions to be excluded. The resulting node set was tiny enough to trigger the “fits within threshold” check.
One clarification about the intended design: whole-repository mode does not embed only the files containing those five components and stop there. Leaf nodes serve as entry points, and the documentation agent can explore outward from them, expand the module tree, follow the dependency graph, and inspect the broader codebase using its code-reading tools.
That said, you are right that using only a few structs as entry points for a 30k-LOC codebase meant that coverage of the remaining code depended entirely on agent exploration, with no clear signal that most of the code had been excluded from the initial view.
This has been fixed in 840e1f4.
Functions are now included as leaf nodes when they represent a significant part of the repository’s architecture. This applies when:
- the repository has no OOP components, as before;
- it has only a small number of classes or structs compared with a much larger number of functions, as in umoria; or
- OOP components account for less than 20% of the component mix.
For a repository structured like umoria, this now produces hundreds of function entry points, allowing LLM-based clustering to run normally instead of being skipped. OOP-heavy repositories, such as typical Java and Python projects, are unaffected.
Reacted by StefanZThank you for fixing this. I can confirm that it works fine now of the given repo.
Reacted by Anh Nguyen HoangI have to walk back my earlier confirmation — sorry about that. When I said it works fine, I had only checked that functions are now considered at all; I did not look at the resulting node count. Coming back to this repository with a closer look, the underlying problem is still there.
The condition for including functions triggers correctly:
Including function components as leaf candidates: 5 class/interface/struct components vs 765 functions — functions carry most of this codebase.But the result is 27 leaf nodes, not the "hundreds of function entry points" the fix aims for, so the same chain still plays out:
Preparing 27 leaf nodes for clustering (3397 tokens, threshold 36369) Skipping LLM clustering; selected leaf nodes fit within the module token threshold Leaf-node entry points cover only 11 of 57 parsed files (19%). Created 0 modules; continuing in whole-repository documentation modeRoot cause
A second mechanism undoes the fix. In
topo_sort.py:concise_leaf_nodes = concise_node(leaf_nodes) if len(concise_leaf_nodes) >= LEAF_REDUCTION_THRESHOLD: # 400 for node, deps in acyclic_graph.items(): for dep in deps: leaf_nodes.discard(dep) concise_leaf_nodes = concise_node(leaf_nodes)
With functions included, umoria produces 770 candidates, which crosses the 400 threshold. Everything that is a dependency of anything else is then discarded, so what survives are only components nothing calls — 27 of 770. The two mechanisms work against each other: the better function inclusion gets, the more reliably the threshold is crossed and the harder the pruning cuts.
The pruning is logged at DEBUG, so at normal verbosity there is no indication that it happened.
Effect on the output
I grepped the generated Markdown for source file names: the documentation mentioned 12 of umoria's 77 source files.
Proposed fix
Compare file coverage before and after the pruning; if it drops below half, keep a capped, file-spread selection instead. On umoria this gives 400 leaf nodes across 46 files, clustering runs normally and produces 46 modules. File coverage in the generated docs goes from 12 to 31 of 77 files.
Existing tests pass except two in
test_gitignore_filtering.py, which fail onmainas well.See PR: #101
On a mid-sized C++ codebase, Phase 1 (dependency analysis) parses all files successfully but extracts only 5 leaf nodes (2065 tokens) from 786 functions. CodeWiki then silently falls back to whole-repository documentation mode, so the generated documentation would not be based on the actual source code. The run continues as if everything were fine.
Reproduction
Observed behavior
Parsing itself succeeds:
But only 5 leaf nodes survive, and clustering is skipped:
The "2065 tokens" conclusion is drawn from the leaf-node set, not from the actual repository size (~30k+ LOC), so the "repo can fit in the context window" decision is based on incorrect data.
Occurs independently of LLM/provider
The collapse happens entirely within Phase 1 static analysis, before any LLM request is made; the LLM clustering step is skipped as a consequence of the collapse. Model and provider configuration are therefore irrelevant for reproduction (verified: no request reached the configured backend during these runs).
Expected behavior
Leaf-node extraction should produce a node set roughly reflecting the function/file count. Or — if the graph analysis cannot produce a usable decomposition — the run should fail.
Environment
CodeWiki 1.0.1 · Python 3.12 · pydantic-ai >=1.0.6,<2