Why I Built Svarupa: Architecture Diffing in Every PR
A clean-looking PR rerouted around a load-bearing layer and the diff didn't show it. The lockfile design, --diff, --drift-base, and CI setup that came out of that review.
Architecture
CI/CD
AI Engineering
The PR was clean. Small, well-described, tests green, a tidy few hundred lines. What it also did — and what I only caught because I knew the system cold — was quietly route a new code path around a layer that everything else went through. No file deleted, no function removed. Just a new edge in the dependency graph where the architecture said none should exist.
A line diff cannot show you that. Diffs show lines; architecture is structure. And the structural change was exactly the kind that, six months later, someone calls "how did this get here?" That review is why svarupa has a lockfile.
The review artifact that didn't exist
AI-assisted codegen made this urgent rather than theoretical. When agents land commits daily — each one locally reasonable, typed, tested — the rate of structural drift goes up by an order of magnitude while review attention stays flat. The architecture diagram from last quarter becomes fiction, and nobody notices because nothing in the PR flow looks at structure.
What I wanted in every PR, next to the line diff, was the architecture delta: modules added or removed, edges changed, nothing else. For that to work, the artifact being diffed has to be deterministic and quiet — a lockfile, not a report.
Designing for zero churn
A lockfile that changes on every commit gets ignored within a week; it becomes noise that reviewers reflexively approve. So architecture.lock was designed backwards from the goal: the median diff on a commit that doesn't change architecture should be zero lines.
The decisions that get there:
Byte-deterministic generation. The same repository state always produces byte-identical artifacts. No timestamps, no hash-ordering accidents, no environment leakage into output.
Canonical ordering. Everything sorted, so equivalent graphs serialize identically regardless of traversal order.
Name-only environment records. The lockfile records that environments exist and what they're called — not incidental config values that would churn on unrelated changes.
Schema versioning. The format is versioned (currently 1.7), and grammar-version changes appear in the diff, so a svarupa upgrade that changes extraction is visible rather than silent.
Measured over 15 commits of real project history, the median change is 0 lines per commit. That number is the whole design: when the lockfile diff does appear in a PR, it means something.
What a diff looks like
svarupa . --diff architecture.lock regenerates from the working tree and prints the delta against the committed lock. The output shape:
Added and removed modules, changed edges, grammar-version changes. That is the entire vocabulary — deliberately. A reviewer can read it in seconds, and "new edge bypassing the billing layer" is a one-line observation instead of a caught-it-by-luck moment.
The stale-base problem
The naive PR diff has a flaw every lockfile workflow eventually hits: the committed base goes stale. If main has moved since the PR branched, diffing the PR against the committed lock mixes this PR's changes with everyone else's merged changes, and the delta is polluted.
--drift-base fixes it by regenerating the base instead of trusting it:
svarupa . --diff pr.lock --drift-base base.lock
Svarupa regenerates the base lockfile from the base branch's actual code first, reports whether the committed base is stale, and diffs against the regenerated truth. The PR author is then answering for their own structural changes, not inheriting the team's.
CI in one command
The whole loop installs with one command:
svarupa setup ci_github # GitHub Actions: architecture delta on every PR
svarupa setup ci_gitlab # GitLab CI: same, per merge request
The base branch commits a byte-deterministic architecture.lock; CI regenerates it for each pull request and svarupa --diff posts the architecture delta as a review artifact.
Read diagram description
The base branch carries a committed, byte-deterministic architecture.lock. When a pull request opens, CI regenerates the lock from the PR code, and svarupa --diff compares it against a base lock regenerated with --drift-base. The architecture delta — added and removed modules and changed edges — is posted as a review artifact, and merging the pull request updates the committed lock. This is an illustrative sketch, not a deployment topology.
The generated workflow pins svarupa==<your version>, so builds are reproducible and upgrades are a deliberate, visible commit — which, thanks to grammar-version reporting, shows up in the diff when extraction behavior changes. One honest caveat: the GitLab job semantics are verified against GitLab's documentation, not against a live GitLab pipeline. The GitHub path is the one I run myself.
The environment dimension
One piece of structure that shows up in reviews constantly: which environment does this touch? Svarupa detects deployment environments from evidence — env-suffixed config filenames (settings.production.py, .env.staging), eas.json build profiles, CI environment: keys, and Dockerfile ENV NODE_ENV — folding aliases automatically (dev → development, prod → production).
Environments surface as env:* nodes in graph.json, a section in REPORT.md, a line in the CLI summary, and a colored environment strip across the top of the deploy-topology diagram, with evidence-gated frames around the components deployed to each. When a PR adds a config file that only exists for production, that is now a visible architectural fact, not something a reviewer has to pattern-match from a filename.
Honest limits, and what would make it 1.0
Svarupa is 0.2.1 and alpha, and the road to 1.0 is defined by the limits, not by a feature wishlist:
Two extraction languages. Python and TypeScript/JavaScript with pinned grammars; everything else is counted and reported as not extractable. More grammars, with the same evidence bar, is the clearest 1.0 gap.
Artifact weight. All views are pre-rendered for offline use, so index.html on huge repos can be tens of MB, and some views are withheld at scale with an explicit note (the data stays complete in graph.json). The 592-module validation repo runs at 113 MB — workable, not elegant.
Call-pinning modesty. On framework-heavy Python code, 29.3% of call sites pinned to a definitive target in one demo (imports and references pin at 77–97%). Better interprocedural resolution would raise that, but never by guessing.
What is already there: 856 tests gating every commit — unit, determinism, and mutation harnesses — plus ruff and strict pyright, and a lockfile that stays quiet until the architecture actually moves.
Implementation checklist
Run svarupa . --lock on your main branch and commit architecture.lock — review the first lock by hand once.
Run svarupa setup ci_github (or ci_gitlab) and open a throwaway PR that adds a module; confirm the delta appears.
Open a second PR that changes no structure; confirm the lockfile diff is empty.
Rebase a stale branch and check that --drift-base reports the committed base honestly.
Decide your merge policy: is an architecture delta a blocking review item, or a required acknowledgment?
Check the environment section of REPORT.md against what you think your environments are.
The thesis and the honesty model are in part 1, and the install-to-MCP workflow is in part 2. The package is on PyPI, the live demo reports are at adaptive-rag and document-extraction-pipeline, the measured numbers are in the case study, and the feature summary is on the project page. If your team wants architecture diffing in review and your stack isn't Python or TypeScript, talk to me — that's the conversation that sets the 1.0 roadmap.
Written by
Vishvdeep Dashadiya
Lead AI Engineer. Agentic AI, real-time ML systems, and cloud-native infrastructure.
I published a CLI that renders architecture diagrams where every node and edge carries file:line evidence — and stays silent where it has none. This is the thesis behind it.
Your agent re-derives the architecture from grep on every task, and guesses confidently when the fragments lie. Hand it the verified graph instead — over MCP, with file:line citations.
Learn how to design decoupled LLM systems that handle 10,000 requests per minute. Covers async queues, caching strategies, RAG integration, and the three biggest bottlenecks you'll face.
LLM System DesignFastAPI
Working on something like this?
Tell me about it. I reply within one working day with a first take and no sales pitch.