Svarupa: Verified Architecture Diagrams from Real Code
A Python CLI that renders architecture diagrams where every node and edge carries file:line evidence, plus a churn-free lockfile that puts architecture diffing into every PR.
Numbers are from the project's own test suite and validation runs on a 592-module production backend; no comparative benchmark against other tools is claimed.
856 tests
Gate every commit: unit, determinism, mutation harnesses
0 lines
Median architecture.lock churn per commit, measured on real history
592 modules
Validated on a production backend at that scale
6
Diagram types, each evidence-backed
At a Glance
Field
Detail
Work type
Self-built open tool, published on PyPI (MIT, alpha)
Sector
Developer tooling: architecture documentation and governance
Engagement
Design, build, and maintain a CLI that turns a repository into an evidence-backed architecture report
Platform
Python CLI (tree-sitter extraction, NetworkX graph), self-contained HTML viewer, MCP server
Focus
Every rendered claim carries file:line evidence; what cannot be proven is stated, not drawn
Outcome
856 tests gate every commit; median lockfile churn of 0 lines/commit measured on real history; validated on a 592-module production backend
The Problem
Architecture documentation describes what someone meant to build. The code describes what exists. Between the two sits drift: the layer that was bypassed "just this once", the dependency added under deadline, the service that quietly grew a second database.
AI-assisted codegen accelerates this. When an agent lands a dozen commits a day, the architecture diagram from last quarter is not stale — it is fiction. And a plausible-but-wrong diagram is worse than none at all, because it gets believed. Reviewers reason from it, onboarding engineers trust it, and agents read it back as ground truth.
The alternative I wanted was boring and strict: a diagram that only shows what it can prove, points at the proof, and says out loud what it left out.
The Evidence Contract
Svarupa's central invariant: every node and every edge in every diagram carries file:line evidence, or it does not render. There is no "probably connected to" box. When a module cannot be drawn because there is no extractable source behind it, the report says so — diagnostic SVA-R-004 names the omission and the reason.
Every rendered node carries file:line evidence; when nothing extractable exists, no box renders and SVA-R-004 states the omission.
Read diagram description
Source code passes through tree-sitter extraction, which yields facts with file:line spans. When evidence is found, a graph node carries that file:line evidence and the renderer draws a box for it. When nothing extractable exists, no box is rendered; the SVA-R-004 diagnostic states the omission instead. The rule is one invariant: every node and edge carries evidence or it does not render. This is an illustrative sketch, not a deployment topology.
This is what separates it from AI-generated architecture diagrams. A language model looks at a repository and produces a plausible picture of what the code probably does. Svarupa parses the code with pinned tree-sitter grammars (Python 0.25.0, TypeScript/JavaScript 0.23.2), builds a knowledge graph in which every fact is anchored to a source line, and renders only from that graph. The viewer lets you click any box through to the line that justifies it.
What It Produces
Running svarupa . in a repository writes a .svarupa/ directory: a self-contained index.html viewer (no server needed), the full knowledge graph as graph.json, a human-readable REPORT.md, and per-diagram layout data.
A repository walk feeds pinned tree-sitter grammars, builds graph.json, and renders six views, a byte-deterministic lockfile, and a report; queries and MCP read the same graph.
Read diagram description
A repository walk feeds tree-sitter extraction with pinned Python and TypeScript grammars. Extraction builds a knowledge graph stored as graph.json with nodes, edges, and file:line evidence. From that one graph svarupa renders an interactive viewer with 6 diagram types, a byte-deterministic architecture.lock, and a REPORT.md summary. A query layer of 7 graph query functions and an MCP server both read graph.json. This is an illustrative sketch, not a deployment topology.
Six diagram types come out of the same graph: architecture, module dependencies, data flow, request flow, deploy topology, and ERD. Each is drawn only where evidence exists. On one public demo repository, the ERD, lifecycle, and workflow views were simply not generated — the code contained no SQL-schema, state-transition, or process evidence to back them. That absence is the product working as designed.
The viewer handles scale honestly. When a view would be unreadable, it draws the 12 most significant boxes and states the remainder ("N more not drawn; all of them are in graph.json"). Every box supports recursive drill-down and a #tab/<box-id> deep link you can paste to a teammate. The same seven query functions behind the viewer are also exposed as a CLI (svarupa query) and as an MCP server over stdio or streamable-http, so an AI agent can answer architecture questions against the verified graph instead of guessing.
The Lockfile and the CI Diff Loop
Diagrams you cannot diff are documentation. Diagrams you can diff are governance. svarupa . --lock writes architecture.lock: a compact, canonically ordered, byte-deterministic text file (schema version 1.7) designed to be committed. Environment records are name-only and ordering is canonical, so the lockfile does not churn — measured over 15 commits of real project history, the median change is 0 lines per commit.
The base branch commits a byte-deterministic architecture.lock; CI regenerates it for each pull request and svarupa --diff posts the architecture delta as a review artifact.
Read diagram description
The base branch carries a committed, byte-deterministic architecture.lock. When a pull request opens, CI regenerates the lock from the PR code, and svarupa --diff compares it against a base lock regenerated with --drift-base. The architecture delta — added and removed modules and changed edges — is posted as a review artifact, and merging the pull request updates the committed lock. This is an illustrative sketch, not a deployment topology.
On every PR, CI regenerates the report and runs --diff against the committed lock: added and removed modules, changed edges, grammar-version changes. --drift-base solves the stale-base problem by regenerating the base lockfile from the base branch's actual code first, so a PR's delta is not polluted by other people's merged changes. svarupa setup ci_github (or ci_gitlab) installs the workflow in one command, pinned to the exact svarupa version for reproducibility.
Measured Numbers
All figures below are from the project's own test suite and validation runs. No comparative benchmark against other tools is claimed.
Metric
Result
Test gate per commit
856 tests — unit, determinism, and mutation harnesses — plus ruff and strict pyright
Lockfile churn
Median 0 lines per commit, measured over 15 commits of real project history
Validation scale
Full report on a 592-module production backend; viewer stays responsive
Viewer artifact on that backend
383 MB reduced to 113 MB via expansion caps (12 boxes per flow stage, 24 pre-rendered expansions)
The demo numbers are live, not illustrative: both interactive reports are hosted on this site, generated with the released svarupa 0.2.1, and correspond to the repositories at generation time.
Known Limits
Alpha software. The CLI is at 0.2.1; interfaces can still change.
Two extraction languages. Python and TypeScript/JavaScript, with pinned grammars. Other files are counted and honestly reported as not extractable.
Artifact weight on huge repos.index.html can still be tens of MB because all views are pre-rendered for offline use, and some views are withheld on very large repositories with an explicit note — the data is always complete in graph.json.
Modest call-pinning in framework-heavy code. On the framework-heavy adaptive-rag demo, 29.3% of Python call sites pinned to a definitive target (import- and reference-based edges pin at 77–97%). Svarupa prints these numbers rather than inflating them.
Links
Svarupa is on PyPI (uv tool install svarupa). The two live demo reports — adaptive-rag and document-extraction-pipeline — are the fastest way to judge the output before installing anything. The project page has the feature summary, and the blog series walks through the thesis, the workflow, and the CI story. If you want this kind of evidence-backed architecture governance in your own codebase, get in touch.
Architecture diagrams that carry their own evidence
Self-Built ProjectMeasured
A Python CLI that reads a codebase with tree-sitter (pinned Python and TypeScript/JavaScript grammars) and produces verified architecture diagrams plus a queryable knowledge graph. Every node and edge carries file:line evidence or it does not render; omissions are stated as diagnostics, never fabricated. Published on PyPI under the MIT license.
I published a CLI that renders architecture diagrams where every node and edge carries file:line evidence — and stays silent where it has none. This is the thesis behind it.
Install svarupa, scan a repo, and read the evidence-backed report: artifact contents, the viewer, the query CLI, and the MCP server — anchored on two live demo reports.
Your agent re-derives the architecture from grep on every task, and guesses confidently when the fragments lie. Hand it the verified graph instead — over MCP, with file:line citations.
A clean-looking PR rerouted around a load-bearing layer and the diff didn't show it. The lockfile design, --diff, --drift-base, and CI setup that came out of that review.
A protocol-neutral reference design for giving AI agents bounded authority, exact approvals, isolated credentials, safe outcome handling and auditable transaction evidence.
Design outputs
The counts below describe elements in this target reference design and its public documentation; they are not observed production results or performance measurements.
7 controls
Risks, enforcement points and validation methods mapped
4 identities
Human, agent, workload and downstream credential kept distinct
Dynamic prompt injection, a gold evaluation set and Amazon Nova fine-tuning cut a live agent's token spend by 90% and lifted tool-calling accuracy from 75.8% to 95%.
Measured outcomes
Measured with before-and-after per-query token counts and tool-call accuracy scored against the gold evaluation set.
A brittle multi-hop prompt chain became one observable LangGraph ReAct agent with a custom MCP layer over the existing FastAPI backend and Claude on Bedrock.
Production outcomes
Production status reflects the deployed client workflow and its live traceability; no comparative performance figure is claimed.