Fast Initialization of a Basic SDD Pipeline

A practical guide to building an SDD loop from specify and clarify through implementation and evidence-based verification, integrating OpenSpec, and handling work that returns with a new bug.

Fast Initialization of a Basic SDD Pipeline

An AI agent can produce code in minutes. The harder problem is making sure it produces the right code, for the right reason, and leaves enough evidence for another person or agent to continue the work safely.

A request such as “add a discount” leaves important questions unanswered:

  • Is the discount fixed or percentage-based?
  • Can the final price become negative?
  • How should half-cent values be rounded?
  • Which API and UI behaviors must remain compatible?
  • What proves that the implementation is complete?

Specification-driven development makes these decisions visible before implementation. This guide builds a lightweight loop that can be copied into a new repository without introducing a large process platform.

The complete working example is available in CyrilStrone/demo-sdd.

What We Will Build

The pipeline has six explicit phases:

specify → clarify → plan → tasks → implement → verify

It also handles a less convenient but very common path:

verified change
      ↓
new production bug
      ↓
classify the gap
      ↓
invalidate stale evidence
      ↓
return to spec, plan, or implementation
      ↓
verify again

At the end, the repository will contain:

AGENTS.md
specs/
  README.md
  phases/
    specify.md
    clarify.md
    plan.md
    tasks.md
    implement.md
    verify.md
  _template/
    spec.md
    clarify.md
    plan.md
    tasks.md
    verify.md
openspec/
  config.yaml
  schemas/
    sdd-loop/
.agents/
  skills/
    sdd-start/
    sdd-rework/
    sdd-verify/
scripts/
  check-sdd.mjs

One Source of Rules

The most important design decision is to avoid copying the full process into every agent command.

Location Responsibility
AGENTS.md Repository constitution, architecture boundaries, commands, prohibitions, and Definition of Done
specs/README.md Short human-readable overview of the SDD loop
specs/phases/*.md Canonical instructions for each phase
specs/_template/*.md Required shape of generated artifacts
Agent skills Thin operation-specific entry points
OpenSpec schema Machine-readable artifact graph and completion rules
Guard script Mechanical checks that prevent an invalid transition

If the process changes, update the canonical phase instruction or schema. A skill should point to that rule instead of maintaining a second version of it.

The Six Phases

1. Specify: define what and why

The specification describes observable behavior, not an implementation guess.

A minimal template:

# Change title

## Goal

What user or system outcome should change?

## Context

What exists today, and why is it insufficient?

## In scope / out of scope

### In scope

- Required behavior.

### Out of scope

- Explicitly excluded work.

## Acceptance criteria

- GIVEN a concrete state
  WHEN an action happens
  THEN an observable result follows

## Assumptions

- Assumptions that still need validation.

Good criteria can be checked without knowing which class or file will implement them.

2. Clarify: make uncertainty a gate

Clarification is not a ceremonial list of questions. It determines whether planning is allowed to start.

Separate unknowns into two groups:

  • Blockers can materially change behavior, compatibility, data, security, or the public contract.
  • Non-blockers have a safe default that can be documented and revisited later.
## BLOCKER questions

1. How are half-cent values rounded?
   - Status: resolved
   - Decision: round half away from zero
   - Source: product decision

## Non-blockers

- Empty UI state uses the existing application pattern.

## Gap register

| Unknown | Impact | Assumption / status | Owner |
| --- | --- | --- | --- |
| Rounding | Price can differ by one cent | Resolved | Product |

Planning starts only when every blocker has an answer or the change is explicitly paused.

3. Plan: describe how the system will change

The plan connects desired behavior to the real repository.

## Approach

## Affected files and modules

## Decisions and alternatives

## Compatibility, risks, and rollback

## Open assumptions

## Verification

Name existing modules and boundaries. Record rejected alternatives when the reason will matter during review or rework.

4. Tasks: create an executable checklist

Tasks should be small enough to verify independently and should name their evidence.

## 1. Implementation

- [ ] 1.1 Add fixed-discount calculation.
  - Check: focused unit tests
- [ ] 1.2 Expose the result through the existing public API.
  - Check: type-check and compatibility test

## 2. Verification

- [ ] 2.1 Run the complete quality gate.
- [ ] 2.2 Compare every acceptance criterion with evidence.

Do not mark a checkbox merely because code was written. Mark it when its check passes.

5. Implement: execute the approved plan

Implementation consumes the artifacts instead of silently replacing them.

During this phase, the agent should:

  1. Follow the task order unless a dependency requires a change.
  2. Keep changes inside the declared scope.
  3. Record any deviation that affects the specification or plan.
  4. Run the task-level check before completing a checkbox.
  5. Return to the appropriate earlier phase when a new blocker appears.

6. Verify: collect evidence

Verification is a separate artifact because “the agent believes it works” is not durable evidence.

## Acceptance criteria

- Criterion: fixed discount never produces a negative price
  - Evidence: `discount.test.ts`, case `clamps result to zero`
  - Result: pass

## OpenSpec scenarios

- Scenario: half-cent result
  - Result: pass

## Checks

- `npm test` — pass
- `npm run typecheck` — pass
- `npm run lint` — pass

## Specification deviations

- None.

## Residual gaps

- None.

Verification should answer three questions: what was checked, where the evidence lives, and what remains uncertain.

Quick Vendor-Neutral Setup

Create the directories:

mkdir -p specs/phases specs/_template scripts

Add the six phase instructions and five artifact templates. Then define the loop in specs/README.md:

## SDD loop

1. `specify` defines observable behavior and scope.
2. `clarify` resolves every blocker.
3. `plan` maps behavior to repository changes.
4. `tasks` creates independently verifiable work items.
5. `implement` executes the approved tasks.
6. `verify` records evidence for every acceptance criterion.

Do not skip a phase unless the repository rules explicitly classify the change as exempt.

In AGENTS.md, make the entry condition explicit:

## Specification-driven development

Use the SDD loop for features, behavior changes, compatibility changes,
and bug fixes that expose a missing or incorrect requirement.

Before implementation:

1. Read `specs/README.md`.
2. Complete specify, clarify, plan, and tasks.
3. Stop if a blocker remains unresolved.

Before completion:

1. Complete verify with command output and criterion-level evidence.
2. Run `npm run check:all`.

This is enough to establish a tool-independent baseline. Codex, Claude Code, another agent, or a human contributor can follow the same files.

Why Add OpenSpec

OpenSpec adds useful mechanics on top of the basic loop:

  • structured changes and permanent capability specifications;
  • dependency-aware artifact generation;
  • strict validation of requirements and scenarios;
  • generated agent skills;
  • apply, sync, and archive operations;
  • custom schemas for project-specific workflows.

The goal is not to replace project rules. OpenSpec makes those rules easier to execute and validate.

Install and Initialize OpenSpec

Install the CLI according to the current OpenSpec documentation, then initialize it for the agent used by the project:

openspec init

For a Codex repository, initialization creates or updates the OpenSpec integration and its skills. Verify the result:

openspec --version
openspec list
openspec validate --all --strict

Keep generated integration files in version control when the team expects the same commands and behavior on every machine.

How Codex Knows When to Use the Loop

Codex does not infer a complete SDD policy from the presence of a specs directory. Three layers work together.

AGENTS.md defines when SDD is mandatory

Repository instructions describe which changes enter the loop, which commands must pass, and which actions are forbidden. Codex reads these rules as project-level instructions.

Skills define how to perform an operation

OpenSpec provides general skills such as propose, apply, sync, and archive. The demo adds three thin project skills:

Skill Purpose
$sdd-start Create a change with the sdd-loop schema and complete planning up to review-ready state
$sdd-rework Classify a newly observed bug, invalidate stale verification, and return to the correct phase
$sdd-verify Run checks, update fingerprints, create verify.md, and execute the complete gate

The skills do not duplicate the entire workflow. They route the agent to AGENTS.md, the schema, phase instructions, and guard.

Example usage:

$sdd-start Add fixed discount calculation.
$sdd-rework Reopen add-fixed-discount after incorrect half-cent rounding.
$sdd-verify Verify add-fixed-discount.

Schema and guard make the agreement testable

Instructions influence agent behavior. A schema and guard prove that required artifacts exist, dependencies are satisfied, tasks are complete, and verification is still current.

This distinction matters:

AGENTS.md tells the agent what must happen.
Skills tell the agent how to start a specific operation.
Schema and guard reject an invalid result.

Create the Custom sdd-loop Schema

The schema represents the artifact graph:

proposal ───────────────┐
specs ──────────────────┼→ clarify → plan → tasks → verify
                        │              │       │
                        └──────────────┴───────┘

A simplified schema configuration looks like this:

name: sdd-loop
version: 1
description: Six-phase specification-driven development loop

artifacts:
  proposal:
    generates: proposal.md
  specs:
    generates: specs/**/*.md
  clarify:
    requires:
      - proposal
      - specs
    generates: clarify.md
  plan:
    requires:
      - clarify
    generates: plan.md
  tasks:
    requires:
      - plan
    generates: tasks.md
  verify:
    requires:
      - specs
      - plan
      - tasks
    generates: verify.md

Use the exact schema syntax supported by the installed OpenSpec version. Pin that version in repositories where reproducibility matters.

Create a change explicitly with the schema:

openspec new change add-fixed-discount --schema sdd-loop

Then inspect its state:

openspec status --change add-fixed-discount
openspec instructions proposal --change add-fixed-discount

Verification Must Detect Stale Evidence

A verify.md file can exist and still be obsolete. The implementation or planning artifacts may have changed after it was produced.

The demo prevents this with fingerprints:

planning fingerprint = hash(specs + clarify + plan + tasks)
implementation fingerprint = hash(managed source and test files)

verify.md records both values. The guard recalculates them and rejects the change when either value differs.

Conceptually:

if (recordedPlanning !== currentPlanning) {
  fail('verify.md is stale: planning changed');
}

if (recordedImplementation !== currentImplementation) {
  fail('verify.md is stale: implementation changed');
}

This turns “please verify again” into a mechanical rule.

When the Same Task Returns With a New Bug

A new bug does not automatically mean “edit the code and rerun tests.” First classify what the bug teaches us.

Classification Meaning Return to
Specification gap Required behavior was missing or wrong specs and clarify
Planning gap Behavior was correct, but the design missed a case plan
Implementation defect Spec and plan already covered the behavior implement
Verification gap The implementation may be correct, but evidence was insufficient verify

For example, a half-cent rounding bug reveals an incomplete pricing rule. That is a specification gap, not merely a code defect.

The rework protocol is:

observe the failure
      ↓
record reproduction evidence
      ↓
classify the gap
      ↓
append rework history
      ↓
update affected artifacts
      ↓
invalidate verify.md
      ↓
add a regression scenario and test
      ↓
implement and verify again

Keep a short history inside the active change:

## Rework history

| Cycle | Trigger | Classification | Return to | Invalidated artifacts | Regression scenario |
| --- | --- | --- | --- | --- | --- |
| 2 | Incorrect half-cent rounding | Specification gap | specs | plan, tasks, verify | Half-cent price |

Do not erase the first cycle. The history explains why requirements and implementation changed.

If the original change has already been archived, create a new change that modifies the permanent capability specification. An archive is history, not a mutable scratch directory.

Test the Pipeline on a Real Change

The demo-sdd repository validates the loop with a fixed-discount capability.

Cycle one covers:

  • a fixed discount amount;
  • a result clamped to zero;
  • invalid numeric input;
  • reusable application logic;
  • criterion-level verification evidence.

Cycle two introduces a rounding regression. The change is reopened through the rework protocol, the permanent requirement gains a half-cent scenario, tests reproduce the failure, and the fingerprints force a new verification.

The repository gate combines normal engineering checks with SDD checks:

{
  "scripts": {
    "test": "node --test",
    "check:sdd": "node scripts/check-sdd.mjs",
    "check:openspec": "openspec validate --all --strict",
    "check:all": "npm test && npm run check:sdd && npm run check:openspec"
  }
}

The important outcome is not the number of artifacts. It is that an incorrect transition becomes visible and fails the gate.

A Practical Daily Workflow

When the request is still unclear:

$openspec-explore Investigate the requested behavior and unknowns.

When the task is ready to enter the loop:

$sdd-start Add a new capability.

Review proposal, scenarios, clarification decisions, plan, and tasks before implementation. Then apply the approved change:

$openspec-apply-change <change-name>

When implementation is complete:

$sdd-verify Verify <change-name>.

If verification is green, archive separately:

$openspec-archive-change <change-name>

If a new bug appears:

$sdd-rework Reopen <change-name> after <observed failure>.

The explicit commands make phase transitions visible in the conversation and repository history.

What Belongs in CI

Start with warnings while the team learns the workflow, then make the critical rules blocking.

Useful checks include:

  1. Every active SDD change uses the expected schema.
  2. Clarification contains no unresolved blockers before planning.
  3. Every task has a verification method.
  4. Completed changes have no unchecked tasks.
  5. verify.md contains evidence for every acceptance criterion.
  6. Planning and implementation fingerprints are current.
  7. OpenSpec strict validation passes.
  8. Normal tests, lint, types, and build checks pass.

CI should validate state, not attempt to make product decisions.

When SDD Helps — and When It Does Not

The full loop is useful for:

  • new capabilities;
  • externally observable behavior changes;
  • compatibility or data-model changes;
  • cross-module work;
  • bugs that expose an incomplete requirement;
  • work performed by multiple people or agents.

For a typo or a mechanical rename, six artifacts are excessive. OpenSpec supports changes without spec-level behavior; the repository can define a lightweight exemption instead of inventing a fake capability.

The objective is not to maximize Markdown. The objective is to make decisions visible before code and verifiable afterward.

Practical Rules

  1. Store specifications beside the code and review them like code.
  2. Keep phase instructions canonical instead of copying them between agents.
  3. Treat clarification as a real gate.
  4. Write requirements in terms of behavior, not classes and filenames.
  5. Give every task its own check.
  6. Never complete a checkbox before its check passes.
  7. Record evidence in verify.md, not confidence.
  8. Preserve deviations and residual gaps explicitly.
  9. Begin with warn-only CI, then strengthen the gate.
  10. Pin OpenSpec when schema reproducibility matters.
  11. Do not rewrite archived history; create a modifying change for a later bug.
  12. Invalidate verification when planning or implementation fingerprints change.

Conclusion

A useful SDD pipeline does not require a heavyweight platform. AGENTS.md, six phase instructions, five templates, and a small guard already create a shared path for humans and AI agents.

OpenSpec adds a dependency graph, strict requirement validation, persistent capability specifications, agent skills, and archiving. A custom schema preserves the exact specify → clarify → plan → tasks → implement → verify loop.

The demo-sdd project proves two consecutive cycles: the original capability and a rounding bug discovered afterward. The schema validates, the guard rejects stale verification, regression evidence is preserved, and the permanent specification evolves instead of hiding the failure.

That is the useful version of SDD: not more ceremony, but a reproducible way to turn intent into reviewed decisions, implementation, and evidence.

Links

  1. Working example: CyrilStrone/demo-sdd
  2. OpenSpec repository
  3. OpenSpec installation
  4. OpenSpec Getting Started
  5. How OpenSpec commands work
  6. OpenSpec CLI reference
  7. OpenSpec troubleshooting
  8. Codex custom instructions with AGENTS.md
  9. Build and use Codex skills