DevSpark Dogfooding: Escaping the Scope-Creep Spiral
I used DevSpark on real development and watched a system built to control complexity start to generate its own. This essay follows the scope-creep spiral, the rules that broke it, and the ideas I wrote down instead of building.
DevSpark Series — 33 articles
- Taking DevSpark to the Next Level
- From Oracle CASE to Spec-Driven AI Development
- Fork Management: Automating Upstream Integration
- Why I Built DevSpark
- Getting Started with DevSpark: Requirements Quality Matters
- DevSpark: Constitution-Based Pull Request Reviews
- DevSpark: The Evolution of AI-Assisted Software Development
- DevSpark: Months Later, Lessons Learned
- DevSpark in Practice: A NuGet Package Case Study
- DevSpark: From Fork to Framework — What the Commits Reveal
- DevSpark v0.1.0: Agent-Agnostic, Multi-User, and Built for Teams
- DevSpark Monorepo Support: Governing Multiple Apps in One Repository
- Bring Your Own AI: DevSpark Unlocks Multi-Agent Collaboration
- Dogfooding DevSpark: Building the Plane While Flying It
- DevSpark: Constitution-Driven AI for Software Development
- The DevSpark Tiered Prompt Model: Resolving Context at Scale
- Workflows as First-Class Artifacts: Defining Operations for AI
- Closing the Loop: Automating Feedback with Suggest-Improvement
- Observability in AI Workflows: Exposing the Black Box
- Autonomy Guardrails: Bounding Agent Action Safely
- Designing the DevSpark CLI UX: Commands vs Prompts
- A Governed Contribution Model for DevSpark Prompts
- The Alias Layer: Masking Complexity in Agent Invocations
- Prompt Metadata: Enforcing the DevSpark Constitution
- DevSpark Blogging Workflow: How I Built Better Articles
- DevSpark and Agent Skills: Beyond Portable AI Capabilities
- The Methodology Tax: Why Grassroots Innovation Gets Rejected
- DevSpark's Next Evolution: Rethinking Where Knowledge Lives
- DevSpark v4: The Spec Was Never the Record
- DevSpark: Change Must Start With a Real Need
- I'm Not a Game Developer. I Built a Game Anyway.
- QA Was Always My Last Battle Front
- DevSpark Dogfooding: Escaping the Scope-Creep Spiral
Filed under Governed AI-assisted delivery
Topic cluster
DevSpark and Spec-Driven DeliverySpec-driven development, AI-assisted delivery workflows, governance, and the DevSpark toolkit.
Using It for Real
There is a particular irony in building a system meant to control software-development complexity and then watching that system fall into scope creep of its own.
I had finished the latest round of work on DevSpark, my spec-driven development framework, around QA, verification, and how a specification carries its meaning through planning and implementation. The feature was good, the tests were good, and the ideas held up. Then I used it, and I mean really used it: not against a toy example built to show the framework working, but against real development, where requirements are imperfect, repositories carry history, architecture pushes back, and somebody eventually has to decide, "Are we done yet?"
How Scope Creep Starts: The Spec That Ate the Feature
One idea behind DevSpark has always been that a plan deserves scrutiny before implementation starts. A specification becomes a plan, the plan becomes tasks, and before an agent disappears into the repository to change things, DevSpark asks some hard questions. Does this solve the problem? Does the plan match the repository? Did I quietly introduce assumptions, or turn something explicitly out of scope into work? Those questions have saved me more than once.
There is a failure mode hiding inside good review, though. A reviewer finds a potential problem, and the natural response is to mitigate it. The mitigation raises a new concern, so I mitigate that, and the fix adds a capability. That capability needs lifecycle rules, the rules need validation, and the validation needs another state. Before long the elegant little feature I started with looks like it is applying for a zoning variance.
I have watched this happen in software projects for decades, and I knew what it looked like. What I had not done was teach DevSpark to recognize it. What I finally named was that I was letting risk prescribe architecture.
"If this fails, what happens?" is a good question. "We could add a recovery operation" is a plausible answer. But then what happens if recovery fails? And off it goes.
Sometimes the more interesting answer is "if it fails, try it again." That is not always right. But when the product owner has already accepted that answer, the planning agent's job is not to rescue me from my own decision by inventing another subsystem. That became one of my favorite rules from this release:
An accepted limitation is a decision, not an unfinished feature request.
It sounds painfully obvious once it is said out loud, which in my experience is a decent sign of a good rule.
KISS Wasn't the Problem
I am a strong believer in KISS and YAGNI, and DevSpark is opinionated about both. I do not want speculative architecture or infrastructure for theoretical futures; I want the smallest change that solves the problem. There is a trap in that, and I walked into it.
The case that exposed it was a versioning feature for a collection of configuration documents. The existing system treated those documents as separate things, each with its own identity and each addressable on its own.
During planning, the planner noticed that the whole collection was small enough to fit in a single stored object, and that made several hard problems disappear at once. Atomicity, reading a version and promotion all became easy, and the result was a beautifully simple implementation that was wrong. The feature was supposed to create versioned copies of the existing documents, not redefine the collection as one new aggregate. The planner had optimized away part of the domain.
KISS and YAGNI had not failed here. They had optimized perfectly, inside the wrong solution space, and nothing stopped them because I had never told the planner that one source document staying one versioned document was an invariant rather than an implementation preference. What was missing was a boundary around what they were allowed to simplify, and that gave me language for something I had always practiced without encoding well:
Implementation freedom begins after domain invariants and architectural constraints are established.
The specification does not need to say which class to create or which partition key to use. But if one source entity must remain one independently addressable, versioned entity, that belongs in the specification. If changing it would make me say, "That's not the feature I asked for," then it was never merely an implementation choice. The spec defines what may vary and what must not, and the plan gets to be creative inside that boundary. Now KISS can do its job without simplifying away the product.
Context Is Not Scope
Agents are very good at noticing things, sometimes too good at it. Hand one a repository, a specification, historical context, future ideas, accepted limitations, architecture notes and background, and it sees all of it. Then it tries to help, which is where the trouble starts. Knowing about something is not the same as being authorized to work on it, and that distinction became the rule I now care about most:
Context is not scope.
DevSpark now has a clearer vocabulary for this. Some knowledge describes the work I am actually doing. Some constrains how I may do it. Some gives context that helps me decide well, and some is not needed for the current command at all. The labels matter less than the fact that only one of those categories can create work. A constraint can stop a planner from choosing a solution, but it cannot invent a feature. Context can explain why a decision exists, but it cannot quietly become a requirement.
When the Safe Path Feels Like Punishment
This discovery was less about architecture and more about human behavior. I got through planning. The plan looked good, the tasks looked good, and then Planning Review ran and stopped me with three problems. They turned out to be a task that should happen later, a missing test fixture, and a verification command that was too narrow. There was nothing wrong with the architecture and no missing product decision, and I knew exactly what to fix.
The workflow, though, essentially said to go backward: run the previous stage again, then come back. It felt like punishment, and I recognized the temptation such a process creates: fix the thing and move on. That is poor UX for an assurance framework, because when following the safe process is much more annoying than bypassing it, people eventually bypass it. I would have too. A gate that people route around no longer protects scope at all, so its friction is a scope-control problem as much as an annoyance.
It exposed a distinction I should have had all along, between a substantive revision and a bounded consistency correction. If a requirement changes, go back to the specification. If the architecture is wrong, go back to the plan. But if Task 9 needs to happen after Task 12, fix Task 9, verify the correction, and keep moving. A gate exists to stop bad work from proceeding. It does not exist to demand a ceremonial reenactment of work that is already settled.
That improvement did not go into this release, and that matters. I had found a good, related improvement and did not build it on the spot, which was the new scope controls working as intended.
The Prompt That Proved the Point
Near the end I made a tiny clarification to the specification. Nothing downstream had changed, but the spec file had, so the provenance check, which records which version of the spec each plan and task list was built from, flagged an upstream difference. The mechanical answer would have been to regenerate everything.
Instead I wrote what I called the classic dogfood prompt. Compare the new spec to the plan: did the meaning change? If not, leave the plan alone. Compare the plan to the tasks: did the meaning change? If not, leave the tasks alone. Do not refresh a hash merely to make a warning disappear; establish semantic equivalence first, and update provenance only if it is still needed.
The prompt came back with exactly the answer I hoped for:
Spec → Plan: NO CONTENT CHANGE.
Plan → Tasks: NO CONTENT CHANGE.
It inspected both artifacts, established that the clarification had not changed their meaning, and changed nothing. That may be the least exciting output an AI agent has ever produced, and I was delighted by it. The agent did what a careful engineer would: it checked, confirmed the meaning was intact, and left the files alone. Regenerating everything, or updating the hashes so the warning went away, would have been the mechanical answers.
The exercise captured the whole day for me, because I had been treating lifecycle stages and files as the substance of the process. The substance is meaning, and the artifacts exist to preserve and communicate it.
What Changed, and What I Didn't Build
The biggest improvement in this release is not a particular command or test. DevSpark got better at telling information from authority. A specification is more than a bag of requirements. Alongside the work, the constraints and the context, it also records the things I have deliberately decided not to solve. Planning should consume the smallest amount of authoritative context that does the job, but it cannot reach simplicity by dropping the constraints that define the job.
Analyze should catch semantic drift, saying that the plan changed what the specification said, rather than designing three more components in response. And tests are now part of the durable description of the system: code says what it does, knowledge says what I believe is true about it, and tests challenge those claims.
The better evidence of improvement is what I did with the new ideas I found along the way. A review that finds three tiny task corrections should probably be able to fix them and continue forward. A harmless specification clarification should probably trigger semantic reconciliation rather than ceremonial regeneration. DevSpark could get smarter about freshness, dependency propagation and bounded remediation, and Planning Review could be less rigid. Every one of those ideas was plausible, and none of them was necessary to finish the feature in front of me.
I also had only a handful of dogfood examples. That was enough evidence to write the ideas down, but not enough to design a generalized mechanism with confidence. Building one immediately would have repeated the mistake I was trying to escape: discover a legitimate adjacent problem, promote it into the current scope, and inherit all the complexity needed to solve it. So I deferred them, because a good idea and a feature that belongs in this release are two different decisions. An earlier DevSpark, and an earlier me, might have turned them into five more requirements. This time I wrote them down and shipped.
That may sound trivial, but the scope-creep spiral rarely starts with someone deciding to add features irresponsibly. It starts with a smart person finding a legitimate problem, solving it, and then solving the problems the solution created. The escape is to stop assuming every problem I notice belongs to the work in front of me.
I did not learn KISS or YAGNI today, and I did not discover that domain meaning matters, that constraints differ from requirements, or that a review process should help people rather than irritate them into bypassing it. I have known all of that for years, though only implicitly, and the agent did not know any of it. What I learned is that knowing it is not enough when agents do more of the planning, because the judgment that used to live in my head has to become part of the system. Not all of it, and not a huge rules engine trying to encode thirty years of instinct. Just enough: the minimum sufficient change.
Explore More
- Dogfooding DevSpark: Building the Plane While Flying It — Using a prompt tool to refine a prompt tool means changing the wrench while you're swinging it. What dogfooding DevSpark teaches that the tool can't.
- DevSpark: Change Must Start With a Real Need — DevSpark treats change as a response to proven need, making agile product development explicit enough for an AI agent to execute with judgment.
- DevSpark v4: The Spec Was Never the Record — Assimilation was DevSpark's open question. In v4, the fix wasn't a new mechanism — it was asking the right question of prompts I already had.
- QA Was Always My Last Battle Front — Handing QA a build and a ticket is not a handoff. How DevSpark connects acceptance criteria to independent validation with one conditional artifact.
- Autonomy Guardrails: Bounding Agent Action Safely — How DevSpark's act/plan execution modes and per-step tool scoping let me expand agent autonomy incrementally — starting with review, earning toward execution.
Related project evidence

DevSpark: Constitutional AI Governance Framework
DevSpark is a standalone AI-assisted development framework that extends Specification-Driven Development with constitution-based PR reviews, codebase-wide compliance auditing, adversarial risk analysis, brownfield constitution discovery, and adaptive lifecycle management. DevSpark makes project constitutions valuable throughout the entire development lifecycle — from greenfield planning through continuous constitutional governance.
Used across most of my active repositories, including this site
DocSpecSpark
Documentation-driven specification system for turning architectural intent into implementation context.

API Test Spark
A NuGet package that embeds an interactive API test harness into any .NET 10 Minimal API with one line of code. It reads the API's own OpenAPI contract, so there is nothing extra to keep in sync.
Thousands of NuGet downloads
Working through something similar?
If this maps to something you're working on, I'm glad to compare notes. Start with the system and the constraint that matters most.
