Whitepaper · 52 pages · October 2026

Driving the AI SDLC with Shipwright

Ship right, not just fast.

Cover of the whitepaper Driving the AI SDLC with Shipwright

AI writes the code now. What did not get solved is what comes after: knowing that the software is still true to what you decided, on every change, not just the first one. This paper is the thinking behind Shipwright, shown on a real example.

The argument

Control is a process that can be built and measured.

Free PDF, no email. Ten chapters, fifteen exhibits, one real change.

Cover of the whitepaper Driving the AI SDLC with Shipwright

Masterclass coming soon

Rather skip the theory and practice on real projects?

Rather skip the theory and practice on real projects?

We are currently building the Shipwright Masterclass. It takes you straight from the theory in this paper to applying it on real projects.

$49

Founding seat · first 50

Then $97 in early access, $497 at full price.

Fully refundable until launch. Founders get in first and help shape what the Masterclass covers.

Where this paper starts · one small change

“I needed a handful of mobile fixes in the Command Center. I typed what I wanted, answered two questions, and went on with my day.”

When I looked again, exactly what I wanted had happened, the way I wanted it. None of that was luck, and I had not asked for any of it. This is what was waiting:

A short specification of the change

Tests written before the code, green

Documentation updated

The decision recorded for the next session

Every piece traced to a requirement

Pull request opened, checks green, merged

The harness around the AI enforced each step, and the loop steered the change through them. That is the work I would otherwise have to remember to do by hand, on every change, forever. It raises the question every technical leader eventually asks about AI-built software:

Who checked it, and against what?

Who checked it, and against what?

Who checked it, and against what?

01 · The problem

AI moved speed. Nothing moved control with it.

Studies from 2025 and 2026 disagree on how much faster AI makes teams. They agree on this much: AI-written code arrives with more issues per change, and teams report losing track of what they own.

1.7x

1.7x

issues per AI-co-authored pull request, against human-only ones. Security issues up to 2.74x.

CodeRabbit · 470 PRs · Dec 2025

44%

44%

of AI code-generation tasks introduced a risky vulnerability. Average security pass rate: 56%.

Veracode · 2026 report

+19%

+19%

time taken with AI by experienced developers, who believed they had been 20% faster.

METR trial · 16 devs · 246 issues · 2025

Exhibit 4: The iceberg. Above the waterline: fast. Below: no record of what was decided, no check that a change did not undo an earlier one, no way to prove what the code does.

The part under the water

Three questions show where control goes missing.

01

Could you prove what this software does today?

02

Is every behavior covered by a test that runs?

03

Did the last change quietly undo something you decided months ago?

If you cannot answer all three without asking the AI, you do not control the software. The AI knows what is in front of it right now. The decisions you made last quarter are only in front of it if somebody put them there.

02 · The answer, in four layers

02 · The answer, in four layers

Control is a process that can be built and measured.

Control is a process that can be built and measured.

The industry found its words in 2026: harness engineering, the system around the agent, and loop engineering, the system that prompts it. Hashimoto and Böckeler described the harness, Osmani and Cherny the loop. The paper orders those ideas into four layers. Each one is a chapter.

The industry found its words in 2026: harness engineering, the system around the agent, and loop engineering, the system that prompts it. Hashimoto and Böckeler described the harness, Osmani and Cherny the loop. The paper orders those ideas into four layers. Each one is a chapter.

Layer 1 · Baseline · Chapter 3

Start with what is true.

Start with what is true.

Start with what is true.

A record only helps while it stays true. A requirement settled in spring can be undone in autumn by a change that nobody connected to it. The AI that made the change did not know, and the one that reviewed it did not know either.

So the first layer is a baseline: requirements with IDs, an architecture document, decision records and an append-only event log, linked to the code and the tests. Two gates keep it true. A change that edits an acceptance criterion has to name it, and the tests bound to it run again before the merge. A change that affects behavior but names no requirement is refused.

Exhibit 06 · The chain from requirement to release

The chain from requirement to release

In one line

A baseline is a record that gets checked. Without the checking it is documentation.

Layer 2 · Harness · Chapter 4

Instructions are read. Gates are run.

Instructions are read. Gates are run.

Instructions are read. Gates are run.

An instruction is advice. A hook that blocks an action is a gate. Picture a cookie factory: nobody asks the filling machine politely for the right number of cookies. A scale weighs every pack, a counter counts, and a pack that fails is blown off the belt.

Shipwright places each rule where it holds. Guides advise. Local hooks stop an action, and some let a person continue, with compliance overrides logged. The six required checks on every pull request stop, full stop. A check that cannot run does not count as a pass: fail-closed.

Exhibit 07 · Where rules live

Where rules live

In one line

Enforce the rules that matter, and keep the harness light enough for a long run.

Layer 3 · Loop · Chapter 5

Put every change through the same loop.

Put every change through the same loop.

Put every change through the same loop.

Instead of prompting an agent task by task, a system prompts it and keeps going until the checks say the change is done. In Shipwright that loop is called iterate, and every change after the first build runs through it.

The loop starts before any code exists. The plan is reviewed first: by internal reviewers, and from medium-sized changes upward also by models from other vendors, because a finding at this point costs minutes, while the same finding after the build costs the build. Only then is the change built, with the tests written first. Once it stands, reviewers with a fresh context check the result, the tests run again, the full suite for anything medium-sized and up, and six required checks in CI decide whether it can merge. The change counts as done when it is merged and green, not when it is sent.

The review in the middle is a cascade rather than a single pass. A specification reviewer acts as a hard gate, a code reviewer follows, and for risky changes a third reviewer deliberately tries to disprove the work. From medium-sized changes upward, models from other vendors look at the code as well. Every review is recorded, and the loop does not finish while one is still open: a review that did not run has to say why.

The change loop, one change from spec to merge

In one line

A few extra minutes per change, and what comes out is what you wanted: cleanly engineered, tested, and backed by evidence.

Layer 4 · Evidence · Chapter 6

Make the result readable.

Make the result readable.

Make the result readable.

How do you know, without asking the AI, that the software is under control? The Control Grade scores a repository from A to F on how much control it can prove, not on how good the code is. Seven weighted dimensions, modeled on the OpenSSF Scorecard.

A dimension it cannot measure is left out, so you never get an unfair F. Two caps keep the number honest, so an average cannot hide a dark spot. The same records produce compliance documents as a by-product: traceability matrix, test evidence, SBOM, change history.

Control gradeWeight, out of 100
A
0/100
Shipwright's own
repository
28.09.2026
requirement traceability25
test health20
change traceability15
change reconciliation15
security10
maintainability and size discipline10
dependency hygiene5
99 ▼
F below 50
D 50+
C 70+
B 80+
A 90+
Capped at 49 if a load-bearing pillar collapses.Capped at 89 if a measured control is dark.

The Control Grade card

In one line

Evidence you have to ask for is not evidence, so build it to be read.

Case · The method, applied to itself

The first honest grade.

The first honest grade.

On 27 June 2026 I ran Shipwright’s own Control Grade on Shipwright. I had built the thing, so I expected a good result. It gave me a B, 88 out of 100, and put the biggest gap right at the top: my changes had quietly stopped tracing back to a requirement.

Two more things it could not measure at all yet: whether changed behavior had been re-verified, and security. Over the next days I closed those three gaps, and the grade climbed to an A. The number matters less than what it did: it handed me a list that showed where control had slipped.

A caution belongs next to that number: it is the tool grading its own repository. Read it as a demonstration of the method, not as independent proof. The grader is open source, so run it on your own repository.

B

B

88 / 100

27.06.2026

A

A

99 / 100

28.09.2026

03 · Chapter 7 · The Command Center

03 · Chapter 7 · The Command Center

Loop engineering you can see.

Loop engineering you can see.

A loop that runs without you needs a place where you can see where each change stands, step in on an exception, and decide. That turns a loop from something you launch into something you steer. The terminal and the files stay the source of truth.

A loop that runs without you needs a place where you can see where each change stands, step in on an exception, and decide. That turns a loop from something you launch into something you steer. The terminal and the files stay the source of truth.

Case · Why I built it

Working in VS Code, I eventually lost the overview: several windows, and several tabs in each. I also needed one simple place where the CI and the loops could bring something to my attention, which became the triage. And I wanted to see at a glance when an iterate needs something from me. That became the inbox.

$

npx @svenroth-ai/shipwright@latest

Optional. Every Shipwright plugin also works in VS Code or the terminal where Claude Code runs.

Optional. Every Shipwright plugin also works in VS Code or the terminal where Claude Code runs.

1

2

3

1

2

3

The task board

Every task has its card on the board.

Pipeline and campaign lanes, one card per task. Each card shows where the task stands and launches or resumes it from there.

1

All projects. Every piece of work, on one board.

2

Backlog to done. An iterate is a card with its state.

3

Inbox and Triage. Every permission prompt in one place, and findings from hooks and CI, ready to become tasks.

1

2

3

1

2

3

A task, with files and terminal

The live session streams right on the task.

Launch a pipeline or iterate from any task, and output streams in place in the embedded terminal on the task page. The Command Center does not wrap Claude Code. It starts the standard Claude command in the terminal and follows the transcript.

1

File tree. The project files, right beside the terminal.

2

A real terminal. It runs Claude Code itself, on the standard Claude Code subscription.

3

Smart viewer. Markdown renders, code is highlighted, diffs are marked.

1

2

3

1

2

3

The Ship’s Log of a project

The Ship’s Log is the project’s home.

Every run is an entry. The Control Grade sits on top, with the specs, agent docs and compliance reports one click away, and a box to start the next change.

1

The Control Grade, with the dimensions behind it.

2

Every run is an entry, with the requirements it touched.

3

The project documents. Read-only: it shows the evidence and does not change it.

Kanban Task Board

Backlog to done, one button to start a run.

Global Inbox

Every permission prompt, in one place.

Multi-project

Many projects, one board.

In one sentence

“Generation is solved. Verification, judgment, and direction are the new craft.”

“Generation is solved. Verification, judgment, and direction are the new craft.”

“Generation is solved. Verification, judgment, and direction are the new craft.”

Osmani, Saboo, Kartakis · Google · The New SDLC With Vibe Coding · May 2026

Ship right, not just fast.

Masterclass coming soon

Founding access to the Shipwright Masterclass is open right now.

Founding access to the Shipwright Masterclass is open right now.

Founding access to the Shipwright Masterclass is open right now.

The harness is free. The Masterclass is where you master it: the discipline and judgment to run Shipwright across real projects. It is in active development, and founders get in first and help shape it.

The harness is free. The Masterclass is where you master it: the discipline and judgment to run Shipwright across real projects. It is in active development, and founders get in first and help shape it.

The harness is free. The Masterclass is where you master it: the discipline and judgment to run Shipwright across real projects. It is in active development, and founders get in first and help shape it.

Founding seats still available

$49

Founding · first 50

$97

Early access

$497

Full price

Get the founding offer · $49

Lowest price it will ever be. Fully refundable until launch.

Ship right, not just fast.

Questions

Questions

Before you decide.

Before you decide.

Is Shipwright free?

Yes. Shipwright is open source on GitHub, and every plugin works without paying anything. The $49 is a founding seat in the Masterclass.

What does it run on?

Claude Code, for now. Shipwright is a full plugin suite wired across the whole lifecycle, and that has to run reliably end to end, so one platform comes first. A Codex path, Codex Light, has been started. Separately, Codextender, an open-source side project, lets Claude Code sessions run on the models of a Codex plan through a local LiteLLM proxy, which increases the available usage limits.

What is the Masterclass?

The place where you learn to drive the harness on real projects. It is in active development. Founders get in first and help shape it.

What if it is not for me?

The founding seat is fully refundable until launch. After that, the price goes to $97 for early access and $497 in full.

Do I need the Command Center?

No. It is optional. Every plugin also works in VS Code or in the terminal where Claude Code runs. It needs no API key of its own.

Does the Control Grade judge my code?

It scores how much control a repository can prove, from A to F, across seven weighted dimensions. A dimension it cannot measure is left out, so you never get an unfair F.

Who’s behind it

Ship right, not just fast.

I’ve spent 20+ years on the process of shipping software cleanly, including FINMA-regulated banking and digital-asset environments, where every change comes with proof that it is safe. Shipwright encodes that discipline so you don’t have to carry it in your head.