Large language models can now inspect substantial codebases, reconstruct business flows and reason about the likely consequences of a change. But if AI can already read the repository, what does a code graph still add? A benchmark comparing direct AI source analysis with Roam on a production-scale codebase offers a useful answer.
The benchmark used Neruba, AspectSoft‘s SaaS billing platform. The same engineering questions were approached in two ways: direct AI source analysis and Roam, an open-source code-intelligence tool focused on structural relationships. The runs were conducted independently, then the most important disagreements were checked against the source. Roam is available through its website and GitHub repository.
The result was not a clean victory for either side. Roam was stronger at systematic structural discovery; the AI was stronger at interpreting business meaning and filtering relevance. The source-level adjudication showed why both capabilities matter: broad graph coverage can surface important candidates, but structural reach is not the same as change impact.
A production-scale codebase
AspectSoft’s Neruba is large enough to make dependency discovery non-trivial, with about 1,110 project files relevant to the substantive analysis.
That scale matters. On a small repository, a capable model can often read enough source to build a useful mental map. At around a thousand project files, repeated discovery becomes more expensive, and the value of fast structural indexing becomes easier to test.
| Neruba benchmark fact | Value |
| Codebase | AspectSoft’s Neruba SaaS billing platform |
| Project files analyzed | ~1,110 |
| Primary stack | TypeScript / Node.js / NestJS / TypeORM |
| Core domains | Subscriptions, pricing, billing, payments, Stripe/webhooks |
| Claims checked against source | 28 |
How the comparison worked
The direct-AI run was limited to source reading and ordinary file search; it could not use Roam, an LSP, a static-analysis engine or another coding agent. Roam, by contrast, built a structural index and answered graph-oriented questions over symbols, callers, references, file dependencies, affected tests, cycles and impact sets. The direct-AI condition used GPT-5.6 Sol, while Roam 14.1.0 was tested independently in standalone CLI mode on Linux x86_64 with Python 3.13.5. Reported Roam timings are specific to that environment.
The benchmark focused on practical change-safety questions: which parts of the system participate in a business flow, what a pricing change might affect, which tests are likely to matter, and where important dependencies are mediated by framework behavior rather than obvious imports.
Three Roam retrieval experiments were then run on the same repository snapshot: TF-IDF/hybrid semantic search, dense ONNX with 12,030 of 12,030 symbols embedded, and an end-to-end tiered workflow that used TF-IDF plus ordinary retrieval first, invoked ONNX only when the cheaper pass looked weak, and then ran graph analysis. These experiments did not change the source-level findings used to adjudicate the original comparison.
Where direct AI analysis was stronger
The clearest advantage was semantic reconstruction. In Neruba, subscription renewal is not represented by one tidy function. The behavior is distributed across Stripe events, webhook persistence, an event bus, runtime handler registration, subscription status synchronization and invoice-payment handling.
The AI reconstructed that sequence from source even where the static graph did not connect every runtime-style relationship. Source inspection supported the reconstruction: the webhook path dispatches into the Stripe event bus, a Nest lifecycle hook registers the subscription handler, and that handler delegates into the relevant services.
This matters because modern frameworks contain relationships that are not simple function-call edges. Dependency injection, lifecycle hooks, decorators, ORM metadata and event registration all create behavior that a code graph must model explicitly.
The AI also surfaced a higher-level pricing issue that is difficult to express as a graph problem. Neruba has both a local subscription-plan price and a Stripe price identifier. The source showed a path where the local value could change while the existing Stripe Price remained attached. That turns a requirement such as “use a different price at renewal” into a source-of-truth and provider-transition problem, not merely a calculation change.
Where Roam was stronger
Roam’s clearest advantage was breadth. Once a useful pivot was known, it could produce structural candidate sets almost immediately. For the pricing function pickSubscriptionInvoicePrice, the benchmarked run returned a reverse-reachability set of 563 symbols across 298 files. Its affected-test query for the same pivot returned 122 tests across 83 files.
A model reading source directly is unlikely to enumerate hundreds of structurally reachable nodes with the same consistency without spending substantially more search and context. Roam also offered a reproducibility advantage: after indexing, known-symbol queries in the retained benchmark data were generally sub-second.
The test results show why high recall can matter. Across Roam’s broader union of related impact and test pivots, 15 of the 16 exact test filenames selected by the AI appeared somewhere in Roam’s candidate universe. Roam also surfaced additional defensive coverage that the direct-source analysis had not emphasized.
That makes the graph particularly interesting as a systematic “don’t forget this” layer for an engineer or coding agent.
What semantic retrieval changed
Roam’s semantic layer helped, but not in the obvious way. On a controlled set of 41 graph-useful pivots from the benchmark, ordinary retrieval placed 28 in the top 20 results, while TF-IDF semantic search found 37. That is a meaningful improvement in finding the right place to start. The target set is Roam-derived rather than independent accuracy ground truth, so the result measures retrieval effectiveness, not code-understanding accuracy.
The full dense ONNX backend was then tested at 100% embedding coverage. It was useful on some vocabulary-shift questions—especially paraphrases around checkout, event/listener wiring and schedulers—but it was not a general upgrade. On the same fixed queries, official ONNX retrieval found 17 of the 41 pivots in the top 20, and dense-hybrid found 18.
Semantic search made it easier to find the right starting points, but it did not change the underlying dependency graph. Relationships Roam had missed, including runtime/event links and a direct test relationship, remained missing. That matters because better retrieval can help an engineer or coding agent reach the right code faster, but it cannot compensate for evidence that is absent from the graph itself.
Dense retrieval also carried a material operational cost on this repository. The dense index took about 10.9 minutes to build versus 18.7 seconds for the TF-IDF condition, grew from roughly 18.2 MB to 118.2 MB, and averaged about 4.18 seconds per query versus roughly 0.95 seconds for TF-IDF. The tiered retrieval experiment then tested the combination end to end: cheap retrieval handled 12 of 13 fixed queries, conditional ONNX fallback closed the remaining query-level gap, and the mixed workflow averaged 2.60 seconds per query, about 38% lower than the dense-query average. On the vocabulary-shift stress set, dense clearly rescued two of seven cases. The result supports tiered retrieval as a practical robustness strategy without making dense search the default.
The key distinction: reachability is not relevance
The main qualification is that a large structural result does not automatically imply a large behavioral impact. The useful engineering question is not only “how many nodes can be reached from this symbol?” but “which of those nodes are likely to matter for this change?”
Source inspection made that distinction concrete. pickSubscriptionInvoicePrice had one production call site, plus a direct colocated test, while Roam’s broader candidate set extended into more distant parts of the application. Some of those paths may be structurally legitimate through shared dependencies without being behaviorally affected by a pricing change.
The same pattern appeared in test selection. Roam’s broad set included clearly relevant subscription and invoice tests, but also tests from more remote areas. The most useful interpretation is therefore a high-recall candidate universe, not an automatic change checklist.
A stronger product experience would rank those relationships explicitly: direct dependency, strong behavioral dependency, defensive regression candidate and merely structurally reachable.
What source-level validation changed
The adjudication phase also found several cases where a Roam result did not align with the relationship visible in source.
One impact query reported no reverse dependents for an invoice-payment handler even though source showed a direct call. Separately, test-map reported no tests for pickSubscriptionInvoicePrice even though its colocated spec imported and called it directly. Roam’s broader affected-tests query still included that spec as a colocated candidate, so the finding is the missed direct relationship, not the difference in result-set size.
Roam also reported an actionable cycle in one project/API-key area where the inspected source showed a one-way call chain rather than an obvious closed cycle.
These findings should not be generalized into a claim that Roam’s graph is broadly incorrect. They are specific gaps observed in one version, repository and benchmark configuration, but they matter because direct relationships visible in the source were not surfaced by queries used for downstream impact and test analysis; a reminder that consequential changes require senior engineering review.
The AI had important weaknesses too
The direct-source process was less deterministic. The model had to choose what to search, which files to open and where to stop. Another run could choose different pivots or simply fail to inspect a structurally connected area.
It also did not produce exhaustive graph-wide closures, fan metrics or hundreds of candidate tests. That was partly intentional because the benchmark prioritized semantic relevance, but it still leaves a false-negative risk: an important dependency may never be inspected.
There is also an operational cost question. Direct AI analysis requires repeated source search and reading. Roam pays an indexing cost and can then answer structural queries repeatedly. In this benchmark, the baseline and TF-IDF-oriented indexes were quick to build, while the dense ONNX condition was substantially more expensive. That makes retrieval strategy—not simply whether vectors are available—part of the operational economics.
The benchmark did not instrument AI token usage or total source-reading cost, so no numerical claim about agent savings is justified. Measuring whether Roam actually reduces model context, source reads and time-to-answer is the most commercially relevant follow-up experiment.
The stronger architecture may be AI + graph
The benchmark suggests that Roam is most compelling not as a replacement for a coding model, but as a retrieval and structural-evidence layer underneath one. A practical workflow suggested by the benchmark would look like this:
- The model identifies the business change and likely pivot symbols.
- Roam uses the cheapest effective retrieval path to find candidate pivots, then returns callers, references, importers, transitive candidates and candidate tests.
- Each result would carry the shortest evidence path and a relevance or confidence class.
- The model reads only the source needed to interpret the highest-value candidates.
- Ambiguous framework/runtime relationships would be explicitly flagged for verification.
This division of labor fits the strengths observed in the benchmark: structural tooling is good at systematic enumeration; the model is good at deciding why a relationship matters.
In an AI-plus-graph workflow, reusable local analysis and mechanical checks give coding agents evidence they can inspect. A dependency graph provides candidates for investigation, but reachability alone does not establish behavioral impact. The real test is whether the combined workflow reaches better-supported decisions with less repeated investigation while making missing evidence visible. Roam supports that model by supplying the reusable structural evidence and mechanical checks that the coding agent can then interpret.
What would make Roam more valuable
Three improvements stand out from this case study: precision, explainability and framework awareness.
First, broad reachability should be ranked rather than presented as one homogeneous form of impact. Second, every recommended file or test should ideally carry an evidence path explaining why it appears. Third, framework-specific relationships such as Nest dependency injection, lifecycle registration and event-bus wiring deserve first-class treatment.
The tool would also benefit from making missing direct evidence easier to spot. In this benchmark, a test that directly imported and called the target was surfaced only through broader affected-test discovery, while the direct relationship itself was not identified. Making that stronger source-level evidence visible would make the result easier to interpret.
So, does AI still need a code graph?
“Need” is too strong. The direct AI analysis understood Neruba well enough to reconstruct important business flows and identify a subtle pricing risk that the graph did not explain. For one-off analysis on a manageable repository, a strong coding model can get surprisingly far on source alone.
But Roam demonstrated fast, repeatable and systematic structural recall, a capability that is difficult to replace with model reasoning alone. The tiered retrieval experiment showed that dense ONNX can help when cheaper retrieval is weak, but semantic similarity remains a way to discover candidates, not evidence of dependency.
A code graph is most valuable as retrieval infrastructure for AI coding agents when it improves recall and reduces repeated source discovery. Cheap semantic retrieval can help find the right starting point; dense retrieval is a useful fallback when vocabulary shifts; the graph supplies structural evidence; and the model should still make the semantic judgment.
That is a different proposition from “AI that understands your code better.” Based on this benchmark, the stronger proposition is a layered system in which retrieval, graph evidence and model reasoning each do the job they are best suited to do.
Subscribe to our newsletter!
+ There are no comments
Add yours