Back to insights

QA Was Always My Last Battle Front

October 7, 202610 min read

I used to throw releases over the wall and let QA reconstruct what I meant. This essay follows how DevSpark connected acceptance criteria at the start of a change to independent validation at the end, without adding a single command, agent, or gate.

DevSpark Series — 32 articles
  1. Taking DevSpark to the Next Level
  2. From Oracle CASE to Spec-Driven AI Development
  3. Fork Management: Automating Upstream Integration
  4. Why I Built DevSpark
  5. Getting Started with DevSpark: Requirements Quality Matters
  6. DevSpark: Constitution-Based Pull Request Reviews
  7. DevSpark: The Evolution of AI-Assisted Software Development
  8. DevSpark: Months Later, Lessons Learned
  9. DevSpark in Practice: A NuGet Package Case Study
  10. DevSpark: From Fork to Framework — What the Commits Reveal
  11. DevSpark v0.1.0: Agent-Agnostic, Multi-User, and Built for Teams
  12. DevSpark Monorepo Support: Governing Multiple Apps in One Repository
  13. Bring Your Own AI: DevSpark Unlocks Multi-Agent Collaboration
  14. Dogfooding DevSpark: Building the Plane While Flying It
  15. DevSpark: Constitution-Driven AI for Software Development
  16. The DevSpark Tiered Prompt Model: Resolving Context at Scale
  17. Workflows as First-Class Artifacts: Defining Operations for AI
  18. Closing the Loop: Automating Feedback with Suggest-Improvement
  19. Observability in AI Workflows: Exposing the Black Box
  20. Autonomy Guardrails: Bounding Agent Action Safely
  21. Designing the DevSpark CLI UX: Commands vs Prompts
  22. A Governed Contribution Model for DevSpark Prompts
  23. The Alias Layer: Masking Complexity in Agent Invocations
  24. Prompt Metadata: Enforcing the DevSpark Constitution
  25. DevSpark Blogging Workflow: How I Built Better Articles
  26. DevSpark and Agent Skills: Beyond Portable AI Capabilities
  27. The Methodology Tax: Why Grassroots Innovation Gets Rejected
  28. DevSpark's Next Evolution: Rethinking Where Knowledge Lives
  29. DevSpark v4: The Spec Was Never the Record
  30. DevSpark: Change Must Start With a Real Need
  31. I'm Not a Game Developer. I Built a Game Anyway.
  32. QA Was Always My Last Battle Front

Filed under Governed AI-assisted delivery

Topic cluster

DevSpark and Spec-Driven Delivery

Spec-driven development, AI-assisted delivery workflows, governance, and the DevSpark toolkit.

Ready for QA

For most of my career, QA has been my last battle front.

That is not because I undervalue them. QA has saved me from myself more times than I can count. They find the assumption I forgot I made, the edge case I never considered, the workflow that worked perfectly on my machine and nowhere else, and the requirement that drifted slightly while passing through my developer brain. They keep me from shooting myself in the foot.

The problem was never appreciation. The problem was that I kept sending them into battle without enough ammunition or shields.

My routine was familiar. I would finish the feature, write the unit tests, tidy the code, get the pull request ready, and announce the magic words:

"Ready for QA."

Then I would throw the release over the wall. Here is the ticket. Here are the acceptance criteria. Here is the build. Good luck.

And QA would start reconstructing what I actually meant. What should happen here? Is this a defect or expected behavior? Does this replace an old regression scenario? Does it behave the same on Web and Mobile? What happens in CRM? What exactly did Product mean by "done"?

Those are not QA problems. They are development-process problems, and more specifically they were mine. Fixing them turned out to be less about testing and more about the lifecycle I had already built into DevSpark, my spec-driven development framework.

QA Was Supposed to Be There From the Beginning

The strange part is that I already knew this.

Look at how a good user story gets written. Nobody writes "Build Feature X" and stops. The story describes the behavior, defines acceptance criteria, draws a line around what is in scope, and names the success conditions. QA thinking is already present at the very start of software development, because acceptance criteria are really the first draft of a test. They say: if this is built correctly, these things will be true.

Then something happens between that early conversation and the finished implementation. Developers take over. The architecture gets interesting. The API takes shape, the models appear, the unit tests get written, the integration tests get added, CI goes green. By the time QA receives the release, those original acceptance criteria can feel more like historical notes than an active contract.

That is backwards. If QA is a first-class citizen, they belong at both ends of the lifecycle: at the beginning, helping define what success means, and at the end, independently proving that what got built is what was promised. The middle, which is implementation, should not be allowed to quietly redefine the contract.

The Trap of Tests Generated From Finished Code

AI made this problem sharper, and I did not see it coming.

It is remarkably easy now to finish a feature and say, "Generate test cases for this code." The AI will happily comply, and the tests may even be good. But there is a philosophical problem hiding in that convenience.

If I generate acceptance tests from the completed implementation, I am really asking:

What tests prove that this code behaves like this code?

What I need QA to ask is a different question:

Does this implementation satisfy the behavior that was agreed to before the code existed?

Those two questions can produce identical-looking test suites and mean completely different things. The first treats the implementation as the source of truth. The second treats it as the subject of the test. A test derived from the code can only confirm the code, including its mistakes. That distinction ended up as one of the foundations of the QA work in DevSpark.

My First Instinct Was the Expensive One

While evolving DevSpark, I started taking this weakness seriously. My first instinct was predictable: add QA to DevSpark.

That sentence gets expensive quickly. Add a QA command. Add a QA prompt. Add a QA agent. Add another lifecycle stage, another approval gate, a test-management system, some state, an integration. Suddenly QA becomes a parallel architecture bolted onto development, with its own seams where things can drift out of sync.

That did not feel right, because DevSpark has been moving in nearly the opposite direction, toward changes that start from a real need: fewer special cases, fewer ceremonies, fewer systems that can disagree with each other. So I tried a different question. Instead of asking how to add QA, I asked where QA already belongs in the lifecycle I have.

That changed the problem completely.

The Places Were Already There

DevSpark rests on three durable pillars: Code, Tests, and Knowledge. Code is what the system does. Knowledge is what I say is true about it, and, as I wrote in The Spec Was Never the Record, it outlives any single spec. Tests are how I prove it.

I had been thinking of Tests too narrowly, mostly as developer tests: unit, component, integration, contract. That definition was too small. QA is evidence too. A manual regression scenario is a Test. So is a browser automation script, an API acceptance script, a CRM workflow validation, a Mobile-to-API-to-CRM scenario. Once I accepted that, QA no longer needed its own architectural kingdom. It already lived inside the model.

The lifecycle points were there too. Specify is where user stories, requirements, and acceptance criteria are already defined. Plan is where there is finally enough context to ask how those requirements will be independently proven. Implement already updates Code, Tests, and Knowledge together, and PR Review already has a place to check that the promised durable test changes actually happened. After deployment, QA still performs the independent validation that keeps everyone honest.

I did not need to bolt QA onto DevSpark. I needed to connect the QA thinking that was already present at the start with the validation that happens at the end.

One Small Artifact

The connection turned out to be a single, deliberately small file: qa-acceptance.md.

It is conditional. If a change does not benefit from independent acceptance planning, the file does not exist. There is no empty form and no "N/A" document. When it does matter, the artifact is a table with six columns:

ID | Trace | Surface | Scenario | Runner | Regression

Trace ties the scenario back to the requirement or success criterion that produced it. Surface says where the behavior lives: API, Web, Mobile, CRM, or Cross-channel. Scenario expresses the observable behavior as Given / When / Then. Runner is automated, manual, or hybrid. And Regression says what should happen to the permanent baseline:

add
update <existing test>
retire <existing test>
none <reason>

The Trace column is the one I care about most, because it fixes the direction of the arrow. The test does not originate from the implementation. It originates from the requirement:

User Story
   ↓
Acceptance Criteria
   ↓
QA Acceptance Scenario
   ↓
Implementation
   ↓
Independent Validation

Not this:

Implementation
   ↓
"Hey AI, generate some tests for this."

Those are very different relationships to the truth.

Two Bookends

I now think of QA as two bookends around development. At the beginning the question is: what would prove that I built the thing I am promising to build? At the end it is: did the thing I actually built pass that proof?

Development lives between those questions, and the bookends keep implementation from quietly moving the goalposts. They also make QA stronger. Instead of arriving after the code exists and asking me what to test, QA already has the ammunition. They know which behavior matters, where it appears, what success looks like, and whether it should join the regression baseline. That frees them to do the thing I actually want from QA, which is to challenge the implementation rather than confirm it.

Regression Is Not a Landfill

The same thinking helped with a frustration I have carried for years: regression suites that only grow.

A sprint produces twenty test cases and fifteen get marked "include in regression." The next sprint adds fifteen more, and the one after that another fifteen. Nobody wants to delete a test, because deleting a test feels irresponsible. Eventually the suite becomes a historical archive of every feature anyone ever shipped.

The distinction that cleared it up for me was simple. Acceptance tests validate the delta. Regression tests validate the resulting system. A new acceptance scenario might become a new regression test, but it might just as easily update an existing one, retire one, fold into existing coverage, or disappear after the release because it only existed to check a one-time condition. That is why the Regression column has four verbs instead of one checkbox. Maintaining the baseline becomes curation instead of accumulation.

Testing the Idea Against a Real Workbook

I did not want to design this in a vacuum, so I compared the model against the QA regression workbook used by a QA team I work with. The workbook carries a lot of legitimate operational detail: Web, Mobile, API, and CRM tabs, environments, execution status, test users, comments, release history, regression indicators.

The easy mistake would have been to reproduce all of that inside DevSpark. I resisted. The workbook taught me exactly one thing my artifact was missing, which was Surface. So I added Surface and stopped.

DevSpark does not need to become the QA execution system. It does not need to manage test users, track whether something passed in staging on iOS, or replace a spreadsheet or Azure Test Plans. Its job comes earlier. It has to make sure that what lands in front of QA is not only "here is what I built," but also "here is what I said I would build, and here is how it was agreed it could be proven." That is a much better handoff.

What the Change Cost

When I tallied the design afterward, the count looked like this:

New commands:          0
New prompts:           0
New agents:            0
New lifecycle stages:  0
New gates:             0
New databases:         0
New model passes:      0

One conditional artifact, dropped into seams the lifecycle already had. QA was not bolted onto the end. It became part of the definition at the beginning and part of the proof at the end.

Architecture Over Reminders

I could have solved this with a checklist item: Mark, remember to involve QA earlier. I know exactly how that would have gone. I would remember most of the time. Then a deadline would show up, or an interesting architecture problem would swallow my attention, the code would become the center of gravity again, and QA would be my last battle front all over.

That is why I keep preferring architecture to reminders. A process that relies on remembering is a process that fails under load. An architecture gives the right behavior a natural place to happen. DevSpark already had the user stories, the acceptance criteria, the Tests pillar, Planning Review, implementation, and PR review. QA did not require reinvention, only connection.

Now QA is no longer on the far side of a wall waiting for my release to land. They are present when I say what I intend to build, present when I decide how it will be proven, and present at the end to ask the question that matters: did I build what I said I would? They still get to do what they have always done best for me, which is stop me from shooting myself in the foot.

Explore More

Working through something similar?

If this maps to something you're working on, I'm glad to compare notes. Start with the system and the constraint that matters most.