Skip to content
VDAI with VD

September 25, 2026 · 7 min read

Svarupa: Architecture Diagrams That Carry Their Own Evidence

I published a CLI that renders architecture diagrams where every node and edge carries file:line evidence — and stays silent where it has none. This is the thesis behind it.

  • Architecture
  • Developer Tools
  • Static Analysis

Open the architecture doc for the last system you worked on. Now open the code. How long before you find the first box in the diagram that no longer exists — or the first dependency in the code that appears nowhere in the diagram?

For me the answer was measured in minutes, every time. So I built a tool that refuses to participate in that gap, and this week I published it: svarupa, a Python CLI that reads a repository and produces architecture diagrams where every node and every edge carries file:line evidence. The name is Sanskrit for "its own true form" — the tool renders what exists, not what someone meant to build.

The problem with pictures of intent

Architecture documentation is written at a moment of intent: the design review, the kickoff, the quarter when someone had time. The code then lives a different life. A layer gets bypassed under deadline. A service grows a second database. An AI coding agent lands a dozen commits a day, each one locally reasonable, none of them updating the diagram.

The result is worse than missing documentation. A plausible-but-wrong diagram gets believed. Reviewers reason from it. New engineers onboard against it. Agents ingest it as ground truth and confidently repeat its errors. The diagram stops being a map and becomes a rumor with a legend.

The obvious modern fix — point an LLM at the repo and ask for a diagram — produces exactly that rumor, faster. The model describes what the code probably does. "Probably" is the problem: the picture is fluent, confident, and unverifiable without doing the reading yourself.

The evidence contract

Svarupa takes the boring, strict route. It parses the repository with pinned tree-sitter grammars (Python 0.25.0, TypeScript/JavaScript 0.23.2), extracts facts, and builds a knowledge graph in which every fact is anchored to a source location. The renderers draw only from that graph. The invariant:

Every node and every edge in every diagram carries file:line evidence, or it does not render.

The svarupa evidence contract from source line to rendered node

Every rendered node carries file:line evidence; when nothing extractable exists, no box renders and SVA-R-004 states the omission.

Read diagram description

Source code passes through tree-sitter extraction, which yields facts with file:line spans. When evidence is found, a graph node carries that file:line evidence and the renderer draws a box for it. When nothing extractable exists, no box is rendered; the SVA-R-004 diagnostic states the omission instead. The rule is one invariant: every node and edge carries evidence or it does not render. This is an illustrative sketch, not a deployment topology.

View full-size diagram(opens in a new tab)

Explore it live: adaptive-rag — diagram report · graph · document-extraction-pipeline — diagram report · graph

No evidence, no box. And the failure mode is loud, not silent: when svarupa omits something, it emits a structured diagnostic. SVA-R-004 states that a module was omitted because there was no extractable source behind it — you see the hole and the reason for it, instead of a diagram that pretends the hole isn't there. Diagnostics follow the SVA-<area>-<nnn> scheme, so they are greppable and CI-actionable.

In the viewer, the evidence is one click away: every box traces back to the exact lines that justify it. You can disagree with a diagram, but you can no longer be fooled by one — the proof is attached.

Honesty about what isn't there

The harder discipline is not drawing things. Two examples from the live demo reports, both generated with the released 0.2.1 against public repos I maintain:

On adaptive-rag (268 nodes, 422 edges), svarupa drew five diagrams and generated no ERD, lifecycle, or workflow views — the code contains no SQL-schema, state-transition, or process evidence to back them. An LLM would happily invent a data model from the variable names. Svarupa reports the absence. It also flagged, via SVA-R-007, that a few modules (the repo root, app/eval, scripts) don't appear in the request-flow view because nothing reachable from a route handler imports them within two hops. That is a true and useful statement about the code, delivered as a diagnostic rather than silently smoothed over.

On document-extraction-pipeline (277 nodes, 406 edges), the module-dependencies root view came out at 836×1076 — larger than the reference viewports — so the viewer scrolls it at natural size and says so (SVA-G-016) instead of shrinking it into illegibility.

Scale gets the same treatment. When a view would be unreadable, svarupa draws the 12 most significant boxes and states the remainder: "N more not drawn; all of them are in graph.json." Nothing is silently dropped, and buttons that would open a withheld view are hidden rather than dead. This is what I mean by capability honesty: the tool's output is calibrated to what it can prove, at whatever scale you point it at.

What you actually get

One command — svarupa . — writes a .svarupa/ directory containing a self-contained index.html viewer, the full knowledge graph as graph.json, a human-readable REPORT.md, and per-diagram layout data.

The svarupa pipeline from repository walk to knowledge graph, rendered views, and lockfile

A repository walk feeds pinned tree-sitter grammars, builds graph.json, and renders six views, a byte-deterministic lockfile, and a report; queries and MCP read the same graph.

Read diagram description

A repository walk feeds tree-sitter extraction with pinned Python and TypeScript grammars. Extraction builds a knowledge graph stored as graph.json with nodes, edges, and file:line evidence. From that one graph svarupa renders an interactive viewer with 6 diagram types, a byte-deterministic architecture.lock, and a REPORT.md summary. A query layer of 7 graph query functions and an MCP server both read graph.json. This is an illustrative sketch, not a deployment topology.

View full-size diagram(opens in a new tab)

Explore it live: adaptive-rag — diagram report · graph · document-extraction-pipeline — diagram report · graph

Six diagram types render from the same graph — architecture, module dependencies, data flow, request flow, deploy topology, ERD — each gated on evidence. The viewer supports recursive drill-down (every box opens into its own evidence-backed interior), #tab/<box-id> deep links for pasting to a teammate, and byte-deterministic output, so the same repository state always produces byte-identical artifacts. That last property is what makes the lockfile and CI diffing in part 3 possible.

Where I would not use it

If your code isn't Python or TypeScript/JavaScript. Those are the two extraction languages, with pinned grammars. Other files are counted and honestly reported as not extractable — but a polyglot monorepo will get partial maps, and svarupa will tell you exactly how partial.

If you need design rationale. Svarupa renders what exists, with evidence. It does not know why the payments service talks to the ledger, and it will not speculate. That layer of documentation still has to be written by a human — the difference is that the human now writes it against a true map.

If you expect maturity. It is 0.2.1, alpha, and honest about it. On framework-heavy Python code, call-site pinning is modest — 29.3% on the adaptive-rag demo, where imports and references pin at 77–97% — because frameworks route calls through decorators and dependency injection that static analysis cannot always resolve. The tool prints these numbers instead of hiding them, but you should know them before you buy in.

If you try it

  • Open the adaptive-rag demo report and click a box through to its evidence before installing anything.
  • Run svarupa . on a repo you know well; the fastest trust test is checking it against your own mental model.
  • Find one SVA-R-004 in your report and confirm the omission is real — that is the contract working.
  • Check which views were not generated, and why.
  • Diff the architecture against your existing docs. Count the rumors.

The install and workflow walkthrough is in part 2: Using Svarupa, and the lockfile, CI diff loop, and the story of why I built it are in part 3. The measured numbers are in the case study, the package is on PyPI, and the project page has the short version. If you try it on something interesting — or it lies to you — I want to hear about it.

Written by

Vishvdeep Dashadiya

Lead AI Engineer. Agentic AI, real-time ML systems, and cloud-native infrastructure.

About me

Keep reading

More posts

September 25, 2026 · 6 min read

Using Svarupa: From Clone to Verified Map in Minutes

Install svarupa, scan a repo, and read the evidence-backed report: artifact contents, the viewer, the query CLI, and the MCP server — anchored on two live demo reports.

Developer ToolsTutorial

September 25, 2026 · 6 min read

Svarupa for Agents: Give Your Coding Agent a Verified Map

Your agent re-derives the architecture from grep on every task, and guesses confidently when the fragments lie. Hand it the verified graph instead — over MCP, with file:line citations.

MCPAI Agents

September 25, 2026 · 7 min read

Why I Built Svarupa: Architecture Diffing in Every PR

A clean-looking PR rerouted around a load-bearing layer and the diff didn't show it. The lockfile design, --diff, --drift-base, and CI setup that came out of that review.

ArchitectureCI/CD

Working on something like this?

Tell me about it. I reply within one working day with a first take and no sales pitch.