Skip to content
Back to blog
Guest blog / 2026-07-31

The Next Step in the Evolution of SDD

Guest Blog

By Zohar Stolar

A Brief History of Time

The Big Bang of LLMs (large language models) gave us the raw capabilities needed to develop software using AI. Then came coding agents, and their rapid, shapeless expansion, led to the creation of Rules and Skills: basically - useful packages of knowledge that teach an agent how to write tests, perform code reviews, work with a particular framework, or carry out a professional task.

Ages ago (as fast as early 2026!), Skills were considered cutting edge: smart loading of instructions, higher respect to the context window, and a relatively convenient way for developers to share knowledge. Then, once again a new, though entirely predictable problem emerged: Skills became scattered across tools, repositories, and private workflows.

For example: one developer might have an excellent rule in Cursor, another maintains a Skill for Claude Code, a senior architect knows which constraints the system must never violate but has never documented them, code-review standards are buried in Pull Request comments, and the terminology of the team, project, or domain exists mainly inside the overloaded heads of team members, gradually becoming a kind of private slang that no outsider could understand. It is like talking to teenagers: their words do not mean anything to you.

In short, the problem is not the LLM capabilities. It is human coordination.

The result is Agentic Cacophony®: impressive capabilities and performance at the individual and small-team level, but nothing consistent or engineered at the organizational level.

The natural next step is not another Prompt Marketplace.

The next step is Organizational Doctrine (!!!!!)

But wait. I am getting ahead of myself. First, a few words about SDD.

Spec-Driven Development: the Next Current Evolutionary Step

AI-assisted software development began with autocomplete. From there it progressed to conversational programming, autonomous agents, and “Vibe Coding.” Every generation increased the speed at which software could be produced. It did not necessarily improve our ability to build the right thing, or ensure that the outcome met architectural, security, quality, and operational requirements.

Specification-Driven Development, or SDD, changes the way we work (although those of us who have been around here for a while, can tell you that this is merely returning back to our roots): instead of treating the conversation with the agent as the Source of Truth (that indisputable source which Agile culture rewrote every morning during stand-ups), SDD creates a durable and explicit workflow:

Specification → Plan → Work Packages → Implementation → Review → Acceptance → Merge

If the first generation of AI-assisted development was an eruption of raw capabilities, SDD is the point at which those capabilities learn to operate inside an orderly and predictable development process.

And now we can finally move on to the next stage of evolution: the Doctrine!

Spec Kitty and Doctrine: the Next Hot Thing (and also the one after)

Spec Kitty is a framework for proper SDD with some very impressive and convenient features. It was built for Enterprises of the modern era (that is July 2026, which will become distant past in two weeks) that wish and need to race ahead with AI adoption across the organization, but simply cannot keep up with the pace of things. Because, if we are honest and admit the truth, it is frightening.

Put yourself in the shoes of a CISO, the head of a legal department, or even some poor software testers who suddenly have to deal with an uncontrolled flood of code that no human being has ever laid eyes on.

At its core, the system provides multi-agent orchestration, isolated work environments, lifecycle management, separation between implementation and review, and traceable human decisions (a.k.a. audit trail).

This makes Spec Kitty less dependent on any particular model or harness. Claude Code, Codex, Cursor, Gemini, Copilot, or tomorrow’s agent can perform the work. The organization’s standards and development process remain stable.

But Spec Kitty is not just about getting shit done. It goes one sinful step further and combines the obvious development skills with what every minimally self-respecting person wants to achieve in life: Governance! Organizational governance!!!!

A pop-art woman reacting to the word Governance.
And really, what’s not to love in… governance!!?

(Governance: a term that has spread like wildfire through a field of dry thorns since the linguistic plague of AI)

The framework was built by veteran software architects who understand the harsh reality of enterprise development: the need for clear boundaries, control planes, security, parallel development, testing, managed code releases, and the uncomfortable fact that plethora of policies and rules are often followed only when somebody remembers them.

When working with Spec Kitty, the source of truth remains the code itself. It is supplemented by a collection of artifacts: specifications, plans, acceptance criteria, work packages, review status, ADRs, all stored alongside the code. Agents work in isolated Git Worktree environments. Hooks and Gates enforce constraints at the appropriate points in the lifecycle. A cumulative, documented event log preserves the history of the work’s progress.

The goal is not deterministic model output since LLMs remain probabilistic.

The goal is a deterministic operating framework around that output: explicit inputs, controlled transitions, enforceable gates, and auditable decisions.

That distinction is critical in Enterprise environments.

Skills Provide Capabilities. Doctrines Create Culture.

Since our business is with organizations, let’s talk a bit about “organizational culture.” An agent Skill answers a question such as “How should I perform this task?” Whereas a doctrine answers broader questions such as:

  • How does this organization build software?
  • Which rules apply here?
  • Who is authorized to make the decision?
  • Which conditions must be satisfied before we can proceed?

These questions are less technical and more procedural. Together, their answers constitute the organizational culture, at least in the context of software development. Although, as we will see later, not exclusively.

Beyond its Healthy Coding® approach (a lighter alternative to “Specification-Driven Development”), Spec Kitty provides an additional layer intended to let the organization answer those questions and enforce its rules.

The surprising name of this layer in Spec Kitty is: Doctrine.

Doctrine encodes the team’s way of working as hierarchical, version-controlled, enforceable content. It may include:

  • Design and architectural principles and boundaries
  • Security and compliance policies
  • Coding and testing standards
  • Review and acceptance gates
  • Business terminology
  • Branding and communication guidelines
  • Escalation policies
  • Agent responsibilities and limitations
  • Required HITL points (Human in the Loop, or Human in Control)

Skills make a single agent more capable.

Doctrine makes different agents behave like members of the same organization.

Doctrine Is a Graph, Not One Giant Spaghetti Prompt

Organizational policy cannot be managed through a single enormous instruction file, especially when teams use multiple AI coding harnesses. Different tasks require different expertise and constraints. A security review does not need the same context as a documentation update. A Python implementation does not require every branding guideline. An architectural decision should be subjected to stricter governance than a mechanical file edit.

Spec Kitty models doctrine as a graph of mandatory directives, procedures, capabilities, shared knowledge, and agent profiles: Rules can cascade through organizational, project, mission, and role-specific layers. Profiles inherit shared capabilities while adding their own specializations and boundaries. Each operation loads the context relevant to that operation instead of injecting the organization’s entire body of knowledge into every agent during every session. This provides consistency without flooding agents with irrelevant information.

This approach also addresses a common source of policy drift. Security requirements no longer depend on somebody remembering to add the appropriate Skill. Brand rules do not disappear because a different tool was used. Architectural constraints no longer remain trapped in the mind of a single senior engineer.

The code repository carries the organization’s institutional knowledge.

The Right Professionals for Each Operation

Real engineering organizations do not ask one person to act simultaneously as product analyst, architect, implementer, security specialist, tester, and final reviewer.

Spec Kitty applies the same principle to AI agents.

Agent profiles define more than personalities. They capture expertise, responsibilities, relevant Doctrine, collaboration patterns, and explicit boundaries.

Spec Kitty can select profiles appropriate to the current action:

  • Annie the Analyst helps clarify requirements.
  • Alphonso the Architect evaluates boundaries and trade-offs.
  • Ivan the implementer works within the approved design.
  • Larry, the language specialist, applies ecosystem-specific practices.
  • Renata the reviewer independently challenges the implementation.
  • Cecilia the security specialist looks for threats others may overlook.

This creates a mixture of expertise around the work. It is not a model-level mixture-of-experts architecture. It is an organizational pattern: different agents contribute different professional perspectives to a shared delivery process.

The Red Team: Adversarial Review Squad

The principle above is demonstrated particularly well by the adversarial review squad. A conventional AI review often asks one agent to inspect a change and return a verdict. But one reviewer brings one framing—and therefore one set of blind spots.

An adversarial squad assigns separate review lenses to different agents. One may defend architectural integrity. Another challenges security assumptions. Another examines test credibility. Another looks for unnecessary complexity or operational risk.

Their conclusions may conflict.

That is a strength.

Strong human engineering teams do not achieve quality because everyone immediately agrees. They achieve it because professionals with different responsibilities challenge each other’s assumptions before those assumptions reach production.

The architect may argue for conceptual consistency. The implementer may expose its operational cost. The security reviewer may reject a convenience both accepted. The tester may demonstrate that a green test suite never exercised the real behavior.

Spec Kitty provides a structure in which that productive disagreement can happen repeatedly and in a traceable manner.

What This Looks Like in Practice Inside an Enterprise SDLC

Well, it looks exactly how you wished it would look like:

  1. Product intent becomes a specification with explicit acceptance criteria.
  2. Planning turns the specification into an architectural approach.
  3. Work is decomposed into bounded packages.
  4. Relevant agent profiles receive the Doctrine needed for their tasks.
  5. Implementers work in isolated execution environments.
  6. Hooks and tests enforce local constraints.
  7. Independent or adversarial reviewers challenge the result.
  8. Acceptance gates verify evidence before progression.
  9. Human authority remains explicit at consequential decisions.
  10. The repository retains an audit trail of what happened and why.

This is not autonomy for autonomy’s sake.

It is a controlled delegation.

The Closest Thing to a Professional AI Engineering Team

Spec Kitty does not pretend that a single agent can replace an entire software organization. It does something more credible: it models the structures that make professional teams effective. The future of AI-assisted development will not be decided by which system produces the largest amount of code from the shortest prompt. It will be decided by the ability of systems to preserve intent, coordinate specialists, enforce organizational constraints, and remain trustworthy as models and tools change.

The next generation after Agent Skills is Doctrine.

And it is open source.

Check it out: spec-kitty.ai

GitHub: Priivacy-ai/spec-kitty