Skip to content

Instantly share code, notes, and snippets.

@rhuss
Last active September 14, 2026 11:22
Show Gist options
  • Select an option

  • Save rhuss/008dae570b35a3413e7ac0439f306dfc to your computer and use it in GitHub Desktop.

Select an option

Save rhuss/008dae570b35a3413e7ac0439f306dfc to your computer and use it in GitHub Desktop.

Audience Q&A: Code Europe, 2026-09-14

Questions from the session chat for "Spec-Driven Development: Making AI Coding Assistants Build What You Actually Want". Some were answered live, the rest afterwards. Names are omitted.

Questions are grouped by theme rather than by the order they arrived, because several people were circling the same thing from different directions.

Specs over time

After a spec is merged and the requirement changes, do we update the existing spec or create a new one?

Create a new one, and move the old spec out of the way so it does not confuse agents that inspect the source tree later.

Can we agree that the specification files are the single source of truth for requirements? And when a fundamental requirement changes in a mature project, affecting several others, do we update the original spec plus everything it touched, or write a new User Story that supersedes the earlier arrangements?

No, not the single source of truth. The running source code is. Specs are a tool, not the record of what the system does.

That makes the second half easier: write a new spec rather than rewriting a web of old ones, and move the superseded one aside.

OpenSpec takes the opposite position, and takes it explicitly. Its specs live in openspec/specs/ and describe current behaviour, while in-flight work sits in openspec/changes/<name>/ carrying delta specs that record only what is added, modified or removed. Its glossary calls that specs directory "the source of truth", holding "the current, agreed-upon behavior of your system".

That is a coherent design and it works. I still think the running code is the more reliable answer to "what does this actually do", because it is the only artifact that cannot be out of date.

How do you keep the specification from becoming outdated when the implementation changes during development?

speckit-spex-evolve updates the spec from the changed code. Very old specs do not get updated at all: they move to an attic/ directory or get deleted.

Can the AI update the spec when implementation decisions change?

Yes, and that is the intended direction of travel. You describe what felt wrong, the agent finds the cause and updates the spec, so the correction survives the next run.

How do you keep code and specifications in sync? Should code only ever change via the specification? What about hotfixes?

Fix first, evolve the spec after. Nobody routes a production incident through a spec cycle, and pretending otherwise just means the spec quietly rots. Ship the fix, then run evolve so the requirement lands in the spec and the next implementation run cannot silently drop it again.

Suggestions for adding new features to an existing SDD project?

Same shape as a changed requirement: a new spec per feature rather than growing the old one. Superseded specs move to attic/. Specs stay per-feature and therefore stay human-sized, which matters for the reading-volume worry below.

Working with Spec Kit

Spec Kit sometimes creates more code than I can review. Can it be forced to produce small increments?

Yes. The spex collab mode proposes multiple implementation phases and opens or updates a PR after each one. You can also tell Spec Kit to implement only phase 1 or phase 2, which the tasks file already splits out, and review each in isolation.

Do you recommend Spec Kit for greenfield projects, where there is no code and therefore no context? It hallucinates a lot and produces unnecessary code, roughly 30% of which could be removed.

Yes for greenfield. The demo in this talk was greenfield, built from an empty directory.

On the over-generation: that is usually a spec problem rather than a hallucination problem. If the agent produced 30% more than you wanted, the spec permitted 30% more than you wanted. Tighten the scope in the spec, and push the standing constraints into the constitution so they apply to every run rather than being re-argued each time. Phased implementation helps too, because an increment small enough to review is an increment small enough to reject.

Can the content of tasks.md be improved? Looking at it, I cannot tell whether Spec Kit will build what I actually wanted.

Talk to the agent about the concern and ask it to update the tasks. Neither tasks.md nor plan.md gets edited by hand, it is all conversation with the agent about what should change.

If you do not read the plan or the tasks, how do you know the agent prepared a proper plan?

Honestly: plan.md and tasks.md are assembler as far as I am concerned. What counts is the outcome, and how well the implementation fits the spec. That is what I tighten and test as hard as I can, automatically where possible and manually where not.

So the attention goes to the spec and to the code. The spec is the higher-level language and gets the larger share of it, but the code is still the thing that actually runs.

Spec Kit handles custom skills badly. Instead of applying a skill such as effective-go while writing code, it writes code its own way and then retrofits the skill, which comes out worse than a plain custom agent. Can a skill be forced in during generation?

No clean answer yet. It should be enforceable through an extension or an adapted Spec Kit template. Spec Kit is very customizable in practice, so this looks like a template problem rather than a limitation.

Should the spec and plan be mirrored into Jira or Trello using a company template, and kept updated through the SDD phases?

Link, don't copy. Two copies drift, and the spec is already versioned alongside the code.

The spex collab workflow offers to create a GitHub issue from a brainstorming document, which is useful when you do not want to specify immediately, when the project is brownfield and does not want specs, or when you are handing the feature to someone else. The next step is then just /speckit-specify <issue-url>.

The spec itself does belong in a tracking system when humans need to review it, so Jira is a reasonable venue if that is the company standard. Most frameworks today are centred on GitHub and GitLab, but there is nothing hard about wiring in other systems.

Models and tooling

Which agents or models work best with SDD?

I have used Opus 4.6 and now 5 for most of my SDD work, mostly because that is what I have access to. From what I hear, GPT 5.6-sol and above are equally good.

The general shape I would suggest: invest more in the first phase with the more elaborate models, up to and including Fable or Astra, and downgrade for implementation, down to Sonnet but not much further. Intent is formed early, and that is where a better model pays for itself.

Beyond that it depends on your own experience and how well you get on with a given model. I find Opus 5 considerably chattier than 4.6, so 4.6 or 4.8 is still my sweet spot, mostly because my skills are tuned to it.

Would the GSD framework fit into this? (github.com/open-gsd/gsd-core)

First, a correction worth making: there are two projects using the GSD name. The one on my landscape slide is gsd-build/get-shit-done, a meta-prompting and context-engineering system for Claude Code by TÂCHES. The one asked about is open-gsd/gsd-core, "Git. Ship. Done". Different projects, same acronym.

On the substance: yes. The point of the landscape slide is that these tools all implement the same underlying pattern, discover what, design how, break into tasks, implement. The methodology is what matters, and any framework following that shape fits.

Is the /grill-me skill enough? In the AI era it feels like we are forced to read a lot of specs and docs.

Whatever the tool, having an agent interrogate you before it writes anything is the right instinct, and it is exactly what the brainstorm phase is for. The knowledge that causes the intent gap is knowledge you already have and did not say. Something has to drag it out of you, and a series of pointed questions does that better than a blank template.

I have written about a related move from the other direction: Know Your Limits: Quiz Yourself Before You Trust AI, where you ask the AI to quiz you on the domain, to find out whether you are even equipped to review what it produced. Scoring badly is useful information rather than failure.

On the reading volume specifically: AI Wrote It. Nobody Read It. One person saves an hour drafting, and many readers collectively lose days. The answer is not to read more, it is to keep specs per-feature and human-sized, and to be honest about which artifacts are written for people and which are written for machines. plan.md and tasks.md are the latter.

Scope and value

Is spec-driven development the answer to the flaws of AI-generated code?

No, it is not a silver bullet, but it speeds things up considerably.

A concrete data point: the Go SDK for OpenShell. Eighteen specs landed in the first week, carrying 47k lines of Go. Six weeks in it stands at 24 specs and 62.7k lines, which splits as 21.5k generated protobuf, 23.1k test code and 18.1k hand-written production code.

The harder problem is what comes next. Getting that upstream means review, and 41k lines of non-generated code is not humanly reviewable, nor is it reviewable split across 6 PRs. That is where the industry is heading regardless, and the likely shape of the answer is that review itself becomes mostly agent work, triaged by other agents. That is already partly true today.

How many tokens did the demo project consume?

Nothing was instrumented during the build, so this is reconstructed from the session transcripts afterwards.

The original devdash build, across 907 assistant turns:

tokens
Output 731,032
Cache writes 2,992,276
Input (uncached) 1,436
Billable total 3,724,744
Cache reads 178,477,247
Total including cache reads 182,201,991

That produced 5,505 lines of Rust, 14 specification files and 11 commits across the demo arc.

The rehearsal runs in the demo worktrees added roughly another 2.66M billable tokens on top.

Two things worth reading off that table. Output tokens are a small fraction of the bill: most of the cost is context being written and re-read. And the gap between the billable figure and the 182M total is what prompt caching is doing for you. Without it this would be a very different conversation.

Logistics

Will there be a recording? Yes. All lectures are recorded and posted to the Code Europe Recordings page.

Where can we see answers to these questions? Here. This document is published alongside the resources list at https://bit.ly/sdd-resources.

Follow-up

Questions that arrived after the session, or that deserve a longer answer than a chat reply, are welcome via LinkedIn or ro14nd.de.

Spec-Driven Development: Resources & Frameworks

Companion links for the Code Europe 2026 talk by Roland Huß.

Slides & Demo

  • Slides (PDF, 28 pages)
  • Audience Q&A - every question from the session chat, answered. Names omitted.
  • rhuss/devdash - the project built live during the talk. Every demo stage is tagged (demo/00-start through demo/07-after-evolve), so you can diff any two steps to see exactly what changed.

SDD Frameworks

Tool Repository Focus
Spec Kit github.com/github/spec-kit Structured workflow, 30+ agent integrations
BMAD Method github.com/bmad-code-org Simulated agile team
OpenSpec github.com/Fission-AI/OpenSpec Delta specs, change-focused
GSD github.com/gsd-build/get-shit-done Minimal overhead, solo devs
Kiro kiro.dev AWS IDE with SDD built in
spex github.com/rhuss/cc-spex Composable Spec Kit extensions
Superpowers github.com/obra/superpowers Process discipline, plans + quality gates
cc-review github.com/rhuss/cc-review Standalone multi-agent code review

Reading

Books

Papers

Articles

Courses

Speaker

Roland Huß, Distinguished Engineer, Red Hat

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment