Skip to content
VDAI with VD

Blog

Engineering notes and reference designs

Production lessons and independent reference-design research across agents, retrieval, fine-tuning and MLOps, with diagrams and implementation detail.

By topic

LangGraph

9 posts

View all posts

September 20, 2026 · 6 min read

Jev in LangGraph: Bounded Answers for Bounded Questions

Jev is a model that never generates text — it answers typed questions with calibrated probabilities. Where that fits in a LangGraph agent, with code, and where I would not use it.

LangGraphModel Routing

September 5, 2026 · 8 min read

Tracing and Scoring a RAG Pipeline Without Slowing It Down

Every run of the pipeline becomes a nested Langfuse trace, and a judge model scores it with DeepEval. Neither adds a millisecond to the response, because scoring starts after the client already has the answer. The wiring: a per-request callback, a fire-and-forget task that cannot fail the request, a judge decoupled from the graph LLM, and an offline golden set for the metrics live traffic cannot compute.

LangfuseDeepEval

September 4, 2026 · 7 min read

Skills as Sub-Agents: One Graph, Many Personas

A skill is a folder with two files. SKILL.md becomes the system prompt, skill.yaml declares which sources the run may read, and the shared LangGraph pipeline runs unchanged on an isolated thread. Slash-invocable, like a command. Here is how a domain persona parametrizes four points of one graph without forking it, and the allowlist bug a code review found along the way.

LangGraphSkills

AI Agents

6 posts

View all posts

September 25, 2026 · 6 min read

Svarupa for Agents: Give Your Coding Agent a Verified Map

Your agent re-derives the architecture from grep on every task, and guesses confidently when the fragments lie. Hand it the verified graph instead — over MCP, with file:line citations.

MCPAI Agents

FastAPI

6 posts

View all posts

September 3, 2026 · 7 min read

It's Not Adaptive RAG: Why I Let the Human Choose the Route

I built a RAG service for a regulated finance and legal practice and called it adaptive-rag. It isn't. There is no LLM router picking a path. Vector search always runs, web search is added only when a person flips a toggle, and four self-reflection gates with hard-capped loops decide whether the answer is good enough to ship. Here is why that design beat the classic one.

RAGLangGraph

Fine-Tuning

6 posts

View all posts

June 15, 2026 · 7 min read

Fine-Tuning a Tool-Calling Agent: SFT + QLoRA on Gemma 3 4B

How I turned a prompt-driven ReAct agent over a 40+ tool MCP backend into a fine-tuned one — collecting multi-turn tool trajectories, building the SFT dataset, training Gemma 3 4B with QLoRA, and serving it back behind the agent loop. The case study that ties the whole fine-tuning series together.

Fine-TuningSFT

June 14, 2026 · 7 min read

LoRA & QLoRA: Fine-Tuning a Model That Doesn't Fit on Your GPU

Full fine-tuning a 7B model needs ~112GB before you load a batch. LoRA trains ~1% of the weights; QLoRA squeezes the frozen base into 4 bits so a large model fits on a single card. Here's how both work — and the four pieces of QLoRA interviewers always probe.

LoRAQLoRA

All posts

September 25, 2026 · 7 min read

Why I Built Svarupa: Architecture Diffing in Every PR

A clean-looking PR rerouted around a load-bearing layer and the diff didn't show it. The lockfile design, --diff, --drift-base, and CI setup that came out of that review.

ArchitectureCI/CD

September 25, 2026 · 6 min read

Using Svarupa: From Clone to Verified Map in Minutes

Install svarupa, scan a repo, and read the evidence-backed report: artifact contents, the viewer, the query CLI, and the MCP server — anchored on two live demo reports.

Developer ToolsTutorial

September 25, 2026 · 6 min read

Svarupa for Agents: Give Your Coding Agent a Verified Map

Your agent re-derives the architecture from grep on every task, and guesses confidently when the fragments lie. Hand it the verified graph instead — over MCP, with file:line citations.

MCPAI Agents

September 20, 2026 · 6 min read

Jev in LangGraph: Bounded Answers for Bounded Questions

Jev is a model that never generates text — it answers typed questions with calibrated probabilities. Where that fits in a LangGraph agent, with code, and where I would not use it.

LangGraphModel Routing

September 17, 2026 · 9 min read

Where Secure Agent Transactions Fit: A Use-Case Map

A decision framework for applying governed transaction controls across procurement, travel, paid APIs, recurring operations, and service purchasing.

Agent TransactionsUse-Case Map

September 17, 2026 · 8 min read

What Exactly Did the Human Approve?

A practical model for turning an approval click into typed, transaction-bound authority that deterministic controls can verify.

Agent AuthorizationHuman Approval

September 17, 2026 · 8 min read

“Uncertain” Is a State, Not an Error

Why ambiguous transaction outcomes must pause blind retries and move through evidence-led reconciliation.

Outcome ReconciliationIdempotency

September 17, 2026 · 8 min read

One Authorization, One Commit Boundary

How to consume approval, reserve constraints, issue scoped authority, and hand off evidence without pretending an external side effect is atomic.

Agent AuthorizationCommit Boundary

September 17, 2026 · 8 min read

Keeping Provider Credentials Outside the Agent

A target design for workload identity, credential brokering, short-lived grants, audience restriction, and isolated execution.

Credential IsolationWorkload Identity

September 17, 2026 · 9 min read

Designing Spend Policies for Autonomous Buyers

How deterministic spend policy combines action limits, budgets, supplier rules, validity, velocity, risk, and separation of duties.

Spend PolicyAgent Procurement

September 17, 2026 · 8 min read

Building a Governed Procurement Agent

An illustrative procurement architecture that separates agent planning, spend policy, exact approval, isolated purchasing, and receipt evidence.

Procurement AgentSpend Controls

September 5, 2026 · 8 min read

Tracing and Scoring a RAG Pipeline Without Slowing It Down

Every run of the pipeline becomes a nested Langfuse trace, and a judge model scores it with DeepEval. Neither adds a millisecond to the response, because scoring starts after the client already has the answer. The wiring: a per-request callback, a fire-and-forget task that cannot fail the request, a judge decoupled from the graph LLM, and an offline golden set for the metrics live traffic cannot compute.

LangfuseDeepEval

September 4, 2026 · 7 min read

Skills as Sub-Agents: One Graph, Many Personas

A skill is a folder with two files. SKILL.md becomes the system prompt, skill.yaml declares which sources the run may read, and the shared LangGraph pipeline runs unchanged on an isolated thread. Slash-invocable, like a command. Here is how a domain persona parametrizes four points of one graph without forking it, and the allowlist bug a code review found along the way.

LangGraphSkills

September 3, 2026 · 7 min read

It's Not Adaptive RAG: Why I Let the Human Choose the Route

I built a RAG service for a regulated finance and legal practice and called it adaptive-rag. It isn't. There is no LLM router picking a path. Vector search always runs, web search is added only when a person flips a toggle, and four self-reflection gates with hard-capped loops decide whether the answer is good enough to ship. Here is why that design beat the classic one.

RAGLangGraph

June 15, 2026 · 7 min read

Fine-Tuning a Tool-Calling Agent: SFT + QLoRA on Gemma 3 4B

How I turned a prompt-driven ReAct agent over a 40+ tool MCP backend into a fine-tuned one — collecting multi-turn tool trajectories, building the SFT dataset, training Gemma 3 4B with QLoRA, and serving it back behind the agent loop. The case study that ties the whole fine-tuning series together.

Fine-TuningSFT

June 14, 2026 · 7 min read

LoRA & QLoRA: Fine-Tuning a Model That Doesn't Fit on Your GPU

Full fine-tuning a 7B model needs ~112GB before you load a batch. LoRA trains ~1% of the weights; QLoRA squeezes the frozen base into 4 bits so a large model fits on a single card. Here's how both work — and the four pieces of QLoRA interviewers always probe.

LoRAQLoRA

June 11, 2026 · 6 min read

The Two Axes of Fine-Tuning: A Mental Model That Stops the Confusion

LoRA, SFT, QLoRA, DPO, PPO, GRPO — they all blur together until you see they live on two independent axes. One decides how you touch the weights, the other decides what signal you train on. The map I use before every fine-tuning project.

Fine-TuningLoRA

April 22, 2026 · 8 min read

From Document Processing to LLM Resilience: Patterns That Scale

Building on Document Extraction Pipeline's Celery/Redis foundation, learn to extend async patterns to LLM-specific resilience: circuit breakers, multi-provider fallbacks, and token bucket rate limiting.

LLM ResilienceCircuit Breaker

April 13, 2026 · 14 min read

OpenClaw: A Self-Hosted AI Assistant with Ollama, Telegram & Discord

Install OpenClaw, wire Anthropic/OpenAI/Google or a local Ollama model, control it from Telegram and Discord, extend it with skills.sh, and turn it into a business gateway for lead qualification, customer support, and agent-ecosystem management.

OpenClawOllama

Reading is free. So is the first call.

Tell me about it. I reply within one working day with a first take and no sales pitch.