Skip to content
VDAI with VD

Self-Built Project

Svarupa: Verified Architecture Diagrams from Real Code

A Python CLI that renders architecture diagrams where every node and edge carries file:line evidence, plus a churn-free lockfile that puts architecture diffing into every PR.

Developer tooling, architecture governanceMeasuredAI-Assisted Custom Software DevelopmentAgentic Workflow DevelopmentMLOps

Measured outcomes

Numbers are from the project's own test suite and validation runs on a 592-module production backend; no comparative benchmark against other tools is claimed.

856 tests

Gate every commit: unit, determinism, mutation harnesses

0 lines

Median architecture.lock churn per commit, measured on real history

592 modules

Validated on a production backend at that scale

6

Diagram types, each evidence-backed

At a Glance

FieldDetail
Work typeSelf-built open tool, published on PyPI (MIT, alpha)
SectorDeveloper tooling: architecture documentation and governance
EngagementDesign, build, and maintain a CLI that turns a repository into an evidence-backed architecture report
PlatformPython CLI (tree-sitter extraction, NetworkX graph), self-contained HTML viewer, MCP server
FocusEvery rendered claim carries file:line evidence; what cannot be proven is stated, not drawn
Outcome856 tests gate every commit; median lockfile churn of 0 lines/commit measured on real history; validated on a 592-module production backend

The Problem

Architecture documentation describes what someone meant to build. The code describes what exists. Between the two sits drift: the layer that was bypassed "just this once", the dependency added under deadline, the service that quietly grew a second database.

AI-assisted codegen accelerates this. When an agent lands a dozen commits a day, the architecture diagram from last quarter is not stale — it is fiction. And a plausible-but-wrong diagram is worse than none at all, because it gets believed. Reviewers reason from it, onboarding engineers trust it, and agents read it back as ground truth.

The alternative I wanted was boring and strict: a diagram that only shows what it can prove, points at the proof, and says out loud what it left out.

The Evidence Contract

Svarupa's central invariant: every node and every edge in every diagram carries file:line evidence, or it does not render. There is no "probably connected to" box. When a module cannot be drawn because there is no extractable source behind it, the report says so — diagnostic SVA-R-004 names the omission and the reason.

The svarupa evidence contract from source line to rendered node

Every rendered node carries file:line evidence; when nothing extractable exists, no box renders and SVA-R-004 states the omission.

Read diagram description

Source code passes through tree-sitter extraction, which yields facts with file:line spans. When evidence is found, a graph node carries that file:line evidence and the renderer draws a box for it. When nothing extractable exists, no box is rendered; the SVA-R-004 diagnostic states the omission instead. The rule is one invariant: every node and edge carries evidence or it does not render. This is an illustrative sketch, not a deployment topology.

View full-size diagram(opens in a new tab)

Explore it live: adaptive-rag — diagram report · graph · document-extraction-pipeline — diagram report · graph

This is what separates it from AI-generated architecture diagrams. A language model looks at a repository and produces a plausible picture of what the code probably does. Svarupa parses the code with pinned tree-sitter grammars (Python 0.25.0, TypeScript/JavaScript 0.23.2), builds a knowledge graph in which every fact is anchored to a source line, and renders only from that graph. The viewer lets you click any box through to the line that justifies it.

What It Produces

Running svarupa . in a repository writes a .svarupa/ directory: a self-contained index.html viewer (no server needed), the full knowledge graph as graph.json, a human-readable REPORT.md, and per-diagram layout data.

The svarupa pipeline from repository walk to knowledge graph, rendered views, and lockfile

A repository walk feeds pinned tree-sitter grammars, builds graph.json, and renders six views, a byte-deterministic lockfile, and a report; queries and MCP read the same graph.

Read diagram description

A repository walk feeds tree-sitter extraction with pinned Python and TypeScript grammars. Extraction builds a knowledge graph stored as graph.json with nodes, edges, and file:line evidence. From that one graph svarupa renders an interactive viewer with 6 diagram types, a byte-deterministic architecture.lock, and a REPORT.md summary. A query layer of 7 graph query functions and an MCP server both read graph.json. This is an illustrative sketch, not a deployment topology.

View full-size diagram(opens in a new tab)

Explore it live: adaptive-rag — diagram report · graph · document-extraction-pipeline — diagram report · graph

Six diagram types come out of the same graph: architecture, module dependencies, data flow, request flow, deploy topology, and ERD. Each is drawn only where evidence exists. On one public demo repository, the ERD, lifecycle, and workflow views were simply not generated — the code contained no SQL-schema, state-transition, or process evidence to back them. That absence is the product working as designed.

The viewer handles scale honestly. When a view would be unreadable, it draws the 12 most significant boxes and states the remainder ("N more not drawn; all of them are in graph.json"). Every box supports recursive drill-down and a #tab/<box-id> deep link you can paste to a teammate. The same seven query functions behind the viewer are also exposed as a CLI (svarupa query) and as an MCP server over stdio or streamable-http, so an AI agent can answer architecture questions against the verified graph instead of guessing.

The Lockfile and the CI Diff Loop

Diagrams you cannot diff are documentation. Diagrams you can diff are governance. svarupa . --lock writes architecture.lock: a compact, canonically ordered, byte-deterministic text file (schema version 1.7) designed to be committed. Environment records are name-only and ordering is canonical, so the lockfile does not churn — measured over 15 commits of real project history, the median change is 0 lines per commit.

Per-PR architecture diff loop driven by a committed svarupa lockfile

The base branch commits a byte-deterministic architecture.lock; CI regenerates it for each pull request and svarupa --diff posts the architecture delta as a review artifact.

Read diagram description

The base branch carries a committed, byte-deterministic architecture.lock. When a pull request opens, CI regenerates the lock from the PR code, and svarupa --diff compares it against a base lock regenerated with --drift-base. The architecture delta — added and removed modules and changed edges — is posted as a review artifact, and merging the pull request updates the committed lock. This is an illustrative sketch, not a deployment topology.

View full-size diagram(opens in a new tab)

Explore it live: adaptive-rag — diagram report · graph · document-extraction-pipeline — diagram report · graph

On every PR, CI regenerates the report and runs --diff against the committed lock: added and removed modules, changed edges, grammar-version changes. --drift-base solves the stale-base problem by regenerating the base lockfile from the base branch's actual code first, so a PR's delta is not polluted by other people's merged changes. svarupa setup ci_github (or ci_gitlab) installs the workflow in one command, pinned to the exact svarupa version for reproducibility.

Measured Numbers

All figures below are from the project's own test suite and validation runs. No comparative benchmark against other tools is claimed.

MetricResult
Test gate per commit856 tests — unit, determinism, and mutation harnesses — plus ruff and strict pyright
Lockfile churnMedian 0 lines per commit, measured over 15 commits of real project history
Validation scaleFull report on a 592-module production backend; viewer stays responsive
Viewer artifact on that backend383 MB reduced to 113 MB via expansion caps (12 boxes per flow stage, 24 pre-rendered expansions)
Geometry validation on that backend2266 s reduced to 0.25 s
Demo: adaptive-rag (public repo)268 nodes / 422 edges; 9 modules, 19 dependencies; 6 routes; 5 diagrams drawn, 0 withheld
Demo: document-extraction-pipeline (public repo)277 nodes / 406 edges; 13 modules, 27 dependencies; 6 routes, 1 task; 5 diagrams drawn

The demo numbers are live, not illustrative: both interactive reports are hosted on this site, generated with the released svarupa 0.2.1, and correspond to the repositories at generation time.

Known Limits

  • Alpha software. The CLI is at 0.2.1; interfaces can still change.
  • Two extraction languages. Python and TypeScript/JavaScript, with pinned grammars. Other files are counted and honestly reported as not extractable.
  • Artifact weight on huge repos. index.html can still be tens of MB because all views are pre-rendered for offline use, and some views are withheld on very large repositories with an explicit note — the data is always complete in graph.json.
  • Modest call-pinning in framework-heavy code. On the framework-heavy adaptive-rag demo, 29.3% of Python call sites pinned to a definitive target (import- and reference-based edges pin at 77–97%). Svarupa prints these numbers rather than inflating them.

Links

Svarupa is on PyPI (uv tool install svarupa). The two live demo reports — adaptive-rag and document-extraction-pipeline — are the fastest way to judge the output before installing anything. The project page has the feature summary, and the blog series walks through the thesis, the workflow, and the CI story. If you want this kind of evidence-backed architecture governance in your own codebase, get in touch.

Work

Related project

Svarupa

Architecture diagrams that carry their own evidence

Self-Built ProjectMeasured

A Python CLI that reads a codebase with tree-sitter (pinned Python and TypeScript/JavaScript grammars) and produces verified architecture diagrams plus a queryable knowledge graph. Every node and edge carries file:line evidence or it does not render; omissions are stated as diagnostics, never fabricated. Published on PyPI under the MIT license.

Pythontree-sitterNetworkX

Keep reading

Related articles

September 25, 2026 · 6 min read

Using Svarupa: From Clone to Verified Map in Minutes

Install svarupa, scan a repo, and read the evidence-backed report: artifact contents, the viewer, the query CLI, and the MCP server — anchored on two live demo reports.

Developer ToolsTutorial

September 25, 2026 · 6 min read

Svarupa for Agents: Give Your Coding Agent a Verified Map

Your agent re-derives the architecture from grep on every task, and guesses confidently when the fragments lie. Hand it the verified graph instead — over MCP, with file:line citations.

MCPAI Agents

September 25, 2026 · 7 min read

Why I Built Svarupa: Architecture Diffing in Every PR

A clean-looking PR rerouted around a load-bearing layer and the diff didn't show it. The lockfile design, --diff, --drift-base, and CI setup that came out of that review.

ArchitectureCI/CD

More

More case studies

Independent Product R&D

TargetTarget Reference Architecture

Designing Secure Agent Transaction Infrastructure

A protocol-neutral reference design for giving AI agents bounded authority, exact approvals, isolated credentials, safe outcome handling and auditable transaction evidence.

Design outputs

The counts below describe elements in this target reference design and its public documentation; they are not observed production results or performance measurements.

7 controls

Risks, enforcement points and validation methods mapped

4 identities

Human, agent, workload and downstream credential kept distinct

Read case study

hypREspace

Measured

hypREspace: Cutting Token Costs 90% While Raising Tool-Calling Accuracy to 95%

Dynamic prompt injection, a gold evaluation set and Amazon Nova fine-tuning cut a live agent's token spend by 90% and lifted tool-calling accuracy from 75.8% to 95%.

Measured outcomes

Measured with before-and-after per-query token counts and tool-call accuracy scored against the gold evaluation set.

90%

Token cost reduction, measured per query

75.8% to 95%

Tool-calling accuracy on the gold eval set

Read case study

hypREspace

Production

hypREspace: Rearchitecting NL-to-SQL into an Agentic Analytics Engine

A brittle multi-hop prompt chain became one observable LangGraph ReAct agent with a custom MCP layer over the existing FastAPI backend and Claude on Bedrock.

Production outcomes

Production status reflects the deployed client workflow and its live traceability; no comparative performance figure is claimed.

One agent

Replaced a multi-hop prompt chain end to end

Every step traced

Failures diagnosed from the exact step trace

Read case study

Have a problem that looks like this?

Tell me about it. I reply within one working day with a first take and no sales pitch.