Fast Initialization of a Basic SDD Pipeline
A practical guide to building an SDD loop from specify and clarify through implementation and evidence-based verification, integrating OpenSpec, and handling work that returns with a new bug.
An AI agent can produce code in minutes. The harder problem is making sure it produces the right code, for the right reason, and leaves enough evidence for another person or agent to continue the work safely.
A request such as “add a discount” leaves important questions unanswered:
- Is the discount fixed or percentage-based?
- Can the final price become negative?
- How should half-cent values be rounded?
- Which API and UI behaviors must remain compatible?
- What proves that the implementation is complete?
Specification-driven development makes these decisions visible before implementation. This guide builds a lightweight loop that can be copied into a new repository without introducing a large process platform.
The complete working example is available in CyrilStrone/demo-sdd.
What We Will Build
The pipeline has six explicit phases:
specify → clarify → plan → tasks → implement → verifyIt also handles a less convenient but very common path:
verified change
↓
new production bug
↓
classify the gap
↓
invalidate stale evidence
↓
return to spec, plan, or implementation
↓
verify againAt the end, the repository will contain:
AGENTS.md
specs/
README.md
phases/
specify.md
clarify.md
plan.md
tasks.md
implement.md
verify.md
_template/
spec.md
clarify.md
plan.md
tasks.md
verify.md
openspec/
config.yaml
schemas/
sdd-loop/
.agents/
skills/
sdd-start/
sdd-rework/
sdd-verify/
scripts/
check-sdd.mjsOne Source of Rules
The most important design decision is to avoid copying the full process into every agent command.
If the process changes, update the canonical phase instruction or schema. A skill should point to that rule instead of maintaining a second version of it.
The Six Phases
1. Specify: define what and why
The specification describes observable behavior, not an implementation guess.
A minimal template:
# Change title
## Goal
What user or system outcome should change?
## Context
What exists today, and why is it insufficient?
## In scope / out of scope
### In scope
- Required behavior.
### Out of scope
- Explicitly excluded work.
## Acceptance criteria
- GIVEN a concrete state
WHEN an action happens
THEN an observable result follows
## Assumptions
- Assumptions that still need validation.Good criteria can be checked without knowing which class or file will implement them.
2. Clarify: make uncertainty a gate
Clarification is not a ceremonial list of questions. It determines whether planning is allowed to start.
Separate unknowns into two groups:
- Blockers can materially change behavior, compatibility, data, security, or the public contract.
- Non-blockers have a safe default that can be documented and revisited later.
## BLOCKER questions
1. How are half-cent values rounded?
- Status: resolved
- Decision: round half away from zero
- Source: product decision
## Non-blockers
- Empty UI state uses the existing application pattern.
## Gap register
| Unknown | Impact | Assumption / status | Owner |
| --- | --- | --- | --- |
| Rounding | Price can differ by one cent | Resolved | Product |Planning starts only when every blocker has an answer or the change is explicitly paused.
3. Plan: describe how the system will change
The plan connects desired behavior to the real repository.
## Approach
## Affected files and modules
## Decisions and alternatives
## Compatibility, risks, and rollback
## Open assumptions
## VerificationName existing modules and boundaries. Record rejected alternatives when the reason will matter during review or rework.
4. Tasks: create an executable checklist
Tasks should be small enough to verify independently and should name their evidence.
## 1. Implementation
- [ ] 1.1 Add fixed-discount calculation.
- Check: focused unit tests
- [ ] 1.2 Expose the result through the existing public API.
- Check: type-check and compatibility test
## 2. Verification
- [ ] 2.1 Run the complete quality gate.
- [ ] 2.2 Compare every acceptance criterion with evidence.Do not mark a checkbox merely because code was written. Mark it when its check passes.
5. Implement: execute the approved plan
Implementation consumes the artifacts instead of silently replacing them.
During this phase, the agent should:
- Follow the task order unless a dependency requires a change.
- Keep changes inside the declared scope.
- Record any deviation that affects the specification or plan.
- Run the task-level check before completing a checkbox.
- Return to the appropriate earlier phase when a new blocker appears.
6. Verify: collect evidence
Verification is a separate artifact because “the agent believes it works” is not durable evidence.
## Acceptance criteria
- Criterion: fixed discount never produces a negative price
- Evidence: `discount.test.ts`, case `clamps result to zero`
- Result: pass
## OpenSpec scenarios
- Scenario: half-cent result
- Result: pass
## Checks
- `npm test` — pass
- `npm run typecheck` — pass
- `npm run lint` — pass
## Specification deviations
- None.
## Residual gaps
- None.Verification should answer three questions: what was checked, where the evidence lives, and what remains uncertain.
Quick Vendor-Neutral Setup
Create the directories:
mkdir -p specs/phases specs/_template scriptsAdd the six phase instructions and five artifact templates. Then define the loop in specs/README.md:
## SDD loop
1. `specify` defines observable behavior and scope.
2. `clarify` resolves every blocker.
3. `plan` maps behavior to repository changes.
4. `tasks` creates independently verifiable work items.
5. `implement` executes the approved tasks.
6. `verify` records evidence for every acceptance criterion.
Do not skip a phase unless the repository rules explicitly classify the change as exempt.In AGENTS.md, make the entry condition explicit:
## Specification-driven development
Use the SDD loop for features, behavior changes, compatibility changes,
and bug fixes that expose a missing or incorrect requirement.
Before implementation:
1. Read `specs/README.md`.
2. Complete specify, clarify, plan, and tasks.
3. Stop if a blocker remains unresolved.
Before completion:
1. Complete verify with command output and criterion-level evidence.
2. Run `npm run check:all`.This is enough to establish a tool-independent baseline. Codex, Claude Code, another agent, or a human contributor can follow the same files.
Why Add OpenSpec
OpenSpec adds useful mechanics on top of the basic loop:
- structured changes and permanent capability specifications;
- dependency-aware artifact generation;
- strict validation of requirements and scenarios;
- generated agent skills;
apply,sync, andarchiveoperations;- custom schemas for project-specific workflows.
The goal is not to replace project rules. OpenSpec makes those rules easier to execute and validate.
Install and Initialize OpenSpec
Install the CLI according to the current OpenSpec documentation, then initialize it for the agent used by the project:
openspec initFor a Codex repository, initialization creates or updates the OpenSpec integration and its skills. Verify the result:
openspec --version
openspec list
openspec validate --all --strictKeep generated integration files in version control when the team expects the same commands and behavior on every machine.
How Codex Knows When to Use the Loop
Codex does not infer a complete SDD policy from the presence of a specs directory. Three layers work together.
AGENTS.md defines when SDD is mandatory
Repository instructions describe which changes enter the loop, which commands must pass, and which actions are forbidden. Codex reads these rules as project-level instructions.
Skills define how to perform an operation
OpenSpec provides general skills such as propose, apply, sync, and archive. The demo adds three thin project skills:
The skills do not duplicate the entire workflow. They route the agent to AGENTS.md, the schema, phase instructions, and guard.
Example usage:
$sdd-start Add fixed discount calculation.
$sdd-rework Reopen add-fixed-discount after incorrect half-cent rounding.
$sdd-verify Verify add-fixed-discount.Schema and guard make the agreement testable
Instructions influence agent behavior. A schema and guard prove that required artifacts exist, dependencies are satisfied, tasks are complete, and verification is still current.
This distinction matters:
AGENTS.md tells the agent what must happen.
Skills tell the agent how to start a specific operation.
Schema and guard reject an invalid result.Create the Custom sdd-loop Schema
The schema represents the artifact graph:
proposal ───────────────┐
specs ──────────────────┼→ clarify → plan → tasks → verify
│ │ │
└──────────────┴───────┘A simplified schema configuration looks like this:
name: sdd-loop
version: 1
description: Six-phase specification-driven development loop
artifacts:
proposal:
generates: proposal.md
specs:
generates: specs/**/*.md
clarify:
requires:
- proposal
- specs
generates: clarify.md
plan:
requires:
- clarify
generates: plan.md
tasks:
requires:
- plan
generates: tasks.md
verify:
requires:
- specs
- plan
- tasks
generates: verify.mdUse the exact schema syntax supported by the installed OpenSpec version. Pin that version in repositories where reproducibility matters.
Create a change explicitly with the schema:
openspec new change add-fixed-discount --schema sdd-loopThen inspect its state:
openspec status --change add-fixed-discount
openspec instructions proposal --change add-fixed-discountVerification Must Detect Stale Evidence
A verify.md file can exist and still be obsolete. The implementation or planning artifacts may have changed after it was produced.
The demo prevents this with fingerprints:
planning fingerprint = hash(specs + clarify + plan + tasks)
implementation fingerprint = hash(managed source and test files)verify.md records both values. The guard recalculates them and rejects the change when either value differs.
Conceptually:
if (recordedPlanning !== currentPlanning) {
fail('verify.md is stale: planning changed');
}
if (recordedImplementation !== currentImplementation) {
fail('verify.md is stale: implementation changed');
}This turns “please verify again” into a mechanical rule.
When the Same Task Returns With a New Bug
A new bug does not automatically mean “edit the code and rerun tests.” First classify what the bug teaches us.
For example, a half-cent rounding bug reveals an incomplete pricing rule. That is a specification gap, not merely a code defect.
The rework protocol is:
observe the failure
↓
record reproduction evidence
↓
classify the gap
↓
append rework history
↓
update affected artifacts
↓
invalidate verify.md
↓
add a regression scenario and test
↓
implement and verify againKeep a short history inside the active change:
## Rework history
| Cycle | Trigger | Classification | Return to | Invalidated artifacts | Regression scenario |
| --- | --- | --- | --- | --- | --- |
| 2 | Incorrect half-cent rounding | Specification gap | specs | plan, tasks, verify | Half-cent price |Do not erase the first cycle. The history explains why requirements and implementation changed.
If the original change has already been archived, create a new change that modifies the permanent capability specification. An archive is history, not a mutable scratch directory.
Test the Pipeline on a Real Change
The demo-sdd repository validates the loop with a fixed-discount capability.
Cycle one covers:
- a fixed discount amount;
- a result clamped to zero;
- invalid numeric input;
- reusable application logic;
- criterion-level verification evidence.
Cycle two introduces a rounding regression. The change is reopened through the rework protocol, the permanent requirement gains a half-cent scenario, tests reproduce the failure, and the fingerprints force a new verification.
The repository gate combines normal engineering checks with SDD checks:
{
"scripts": {
"test": "node --test",
"check:sdd": "node scripts/check-sdd.mjs",
"check:openspec": "openspec validate --all --strict",
"check:all": "npm test && npm run check:sdd && npm run check:openspec"
}
}The important outcome is not the number of artifacts. It is that an incorrect transition becomes visible and fails the gate.
A Practical Daily Workflow
When the request is still unclear:
$openspec-explore Investigate the requested behavior and unknowns.When the task is ready to enter the loop:
$sdd-start Add a new capability.Review proposal, scenarios, clarification decisions, plan, and tasks before implementation. Then apply the approved change:
$openspec-apply-change <change-name>When implementation is complete:
$sdd-verify Verify <change-name>.If verification is green, archive separately:
$openspec-archive-change <change-name>If a new bug appears:
$sdd-rework Reopen <change-name> after <observed failure>.The explicit commands make phase transitions visible in the conversation and repository history.
What Belongs in CI
Start with warnings while the team learns the workflow, then make the critical rules blocking.
Useful checks include:
- Every active SDD change uses the expected schema.
- Clarification contains no unresolved blockers before planning.
- Every task has a verification method.
- Completed changes have no unchecked tasks.
verify.mdcontains evidence for every acceptance criterion.- Planning and implementation fingerprints are current.
- OpenSpec strict validation passes.
- Normal tests, lint, types, and build checks pass.
CI should validate state, not attempt to make product decisions.
When SDD Helps — and When It Does Not
The full loop is useful for:
- new capabilities;
- externally observable behavior changes;
- compatibility or data-model changes;
- cross-module work;
- bugs that expose an incomplete requirement;
- work performed by multiple people or agents.
For a typo or a mechanical rename, six artifacts are excessive. OpenSpec supports changes without spec-level behavior; the repository can define a lightweight exemption instead of inventing a fake capability.
The objective is not to maximize Markdown. The objective is to make decisions visible before code and verifiable afterward.
Practical Rules
- Store specifications beside the code and review them like code.
- Keep phase instructions canonical instead of copying them between agents.
- Treat clarification as a real gate.
- Write requirements in terms of behavior, not classes and filenames.
- Give every task its own check.
- Never complete a checkbox before its check passes.
- Record evidence in
verify.md, not confidence. - Preserve deviations and residual gaps explicitly.
- Begin with warn-only CI, then strengthen the gate.
- Pin OpenSpec when schema reproducibility matters.
- Do not rewrite archived history; create a modifying change for a later bug.
- Invalidate verification when planning or implementation fingerprints change.
Conclusion
A useful SDD pipeline does not require a heavyweight platform. AGENTS.md, six phase instructions, five templates, and a small guard already create a shared path for humans and AI agents.
OpenSpec adds a dependency graph, strict requirement validation, persistent capability specifications, agent skills, and archiving. A custom schema preserves the exact specify → clarify → plan → tasks → implement → verify loop.
The demo-sdd project proves two consecutive cycles: the original capability and a rounding bug discovered afterward. The schema validates, the guard rejects stale verification, regression evidence is preserved, and the permanent specification evolves instead of hiding the failure.
That is the useful version of SDD: not more ceremony, but a reproducible way to turn intent into reviewed decisions, implementation, and evidence.