Back to insights

I'm Not a Game Developer. I Built a Game Anyway.

October 3, 20267 min read

After one ad too many in a mobile puzzle game, I built my own in a week with Godot and DevSpark, despite never having written a game. The most useful moment was when the code passed every test and the puzzles still weren't fun. I'm asking readers to play it, read the code, and tell me whether shipping makes me a game developer.

DevSpark Series — 31 articles
  1. Taking DevSpark to the Next Level
  2. From Oracle CASE to Spec-Driven AI Development
  3. Fork Management: Automating Upstream Integration
  4. Why I Built DevSpark
  5. Getting Started with DevSpark: Requirements Quality Matters
  6. DevSpark: Constitution-Based Pull Request Reviews
  7. DevSpark: The Evolution of AI-Assisted Software Development
  8. DevSpark: Months Later, Lessons Learned
  9. DevSpark in Practice: A NuGet Package Case Study
  10. DevSpark: From Fork to Framework — What the Commits Reveal
  11. DevSpark v0.1.0: Agent-Agnostic, Multi-User, and Built for Teams
  12. DevSpark Monorepo Support: Governing Multiple Apps in One Repository
  13. Bring Your Own AI: DevSpark Unlocks Multi-Agent Collaboration
  14. Dogfooding DevSpark: Building the Plane While Flying It
  15. DevSpark: Constitution-Driven AI for Software Development
  16. The DevSpark Tiered Prompt Model: Resolving Context at Scale
  17. Workflows as First-Class Artifacts: Defining Operations for AI
  18. Closing the Loop: Automating Feedback with Suggest-Improvement
  19. Observability in AI Workflows: Exposing the Black Box
  20. Autonomy Guardrails: Bounding Agent Action Safely
  21. Designing the DevSpark CLI UX: Commands vs Prompts
  22. A Governed Contribution Model for DevSpark Prompts
  23. The Alias Layer: Masking Complexity in Agent Invocations
  24. Prompt Metadata: Enforcing the DevSpark Constitution
  25. DevSpark Blogging Workflow: How I Built Better Articles
  26. DevSpark and Agent Skills: Beyond Portable AI Capabilities
  27. The Methodology Tax: Why Grassroots Innovation Gets Rejected
  28. DevSpark's Next Evolution: Rethinking Where Knowledge Lives
  29. DevSpark v4: The Spec Was Never the Record
  30. DevSpark: Change Must Start With a Real Need
  31. I'm Not a Game Developer. I Built a Game Anyway.

Topic cluster

DevSpark and Spec-Driven Delivery

Spec-driven development, AI-assisted delivery workflows, governance, and the DevSpark toolkit.

Thirty Seconds of Fun

I like the arrow puzzle games on my phone. A screen full of tangled arrows, one of which has a clear path. You tap it, watch it slide away, and the knot slowly comes apart. It's a simple pleasure, right up until the game decides I've had thirty seconds of fun and now owe it an advertisement.

After one interruption too many, I had the thought that has gotten programmers into trouble for decades: how hard could it be to build this myself?

The honest answer was that I had no idea, because I'm not a game developer. I've spent thirty years designing and building software systems. I know architecture, requirements, testing, and the countless small decisions between an idea and working software. Godot, puzzle generation, game loops, and level design were all unfamiliar territory.

A week later I had ArrowSpark, an ad-free arrow puzzle game with no lives, no waiting, and no monetization machinery between you and the puzzle.

That leaves me with a question I still can't answer cleanly: does shipping a game make me a game developer?

Even Nothing Turned Out to Be Something

I wanted a tangled board of arrows and nothing else. That sounds like the easy part, and it wasn't.

An arrow in ArrowSpark isn't an icon pointing in one of four directions. It has a tail, and the tail can turn. Every segment has to connect cleanly to the next, without gaps or branches, and an arrow's own tail must never trap it. The moment arrows have bodies, the board stops being a grid of symbols and becomes a geometry problem.

Scoring followed, then multiple levels, hints, larger boards, zooming, and panning. Each one was a small feature on its own, and each one changed what the others needed to do.

Along the way the game picked up the design rule I'm proudest of: hard is fine, unsolvable is not.

The puzzles I enjoy make me stare at the screen until I spot the one arrow that frees the next, and the next, until the whole board comes apart. What I can't stand is a puzzle that looks hard only because the generator accidentally built something impossible. The first kind respects the player. The second is a bug wearing a difficulty setting.

Where the Process Came In

The other half of this story is how someone outside the domain got there at all. That's where DevSpark, my approach to spec-driven, AI-assisted development, did its work.

A prompt like "build me an arrow puzzle game" leaves every important question unanswered. What makes a puzzle valid? How do we measure difficulty? What does "done" mean? Those are software-development questions, and I know how to investigate software-development questions even when the stack is new to me.

So I treated ArrowSpark the way I'd treat any system. I broke each capability into a bounded spec, challenged the requirements, tested the implementation, reviewed the results, and recorded why each decision was made. The AI helped me find my way around Godot. I supplied the judgment about what we were building and why.

I suspect this is where a lot of the current conversation about AI and expertise goes wrong. It treats the tool as either a replacement for knowing things or a toy. In this project it was neither. It was closer to a very fast colleague who knew the engine better than I did and had no opinion about whether the game was any good.

The Test That Passed and Was Still Wrong

The most useful moment of the week came when the AI and I built something that worked and was still wrong.

We wrote a difficulty analyzer. It measured dependency chains between arrows, confirmed that every puzzle was solvable, and passed every test. By every measure I had written down, it was done.

Then I played the puzzles. They weren't much fun.

By my own rough rating they landed around two out of five on difficulty, regardless of what the analyzer reported. The code had faithfully implemented an incomplete understanding of the problem. Tests could tell me whether the system did what I specified. Only playing could tell me what the specification had missed.

I've seen this pattern in business software for years, usually with a dashboard that is technically correct and practically useless. I didn't expect to meet it again in a puzzle game, but the lesson transferred whole. A green test suite answers the question you asked. It stays silent about the question you should have asked. I wrote more about that gap in The AI Confidence Trap.

Evidence for Both Sides

That moment cuts both ways, and I'd rather lay out both readings than pick the flattering one.

A seasoned game designer might have recognized the problem earlier. They might never have trusted a dependency-chain metric to stand in for difficulty in the first place. That's evidence for "not a game developer."

But I caught it. I played the output, believed what I was seeing over what the analyzer reported, and revised the generation and validation approach instead of arguing with the results. That feels like evidence for the other side.

The specs, milestones, and account of that pivot are part of the project. The process left receipts, which means you don't have to take my word for any of this.

You Decide

ArrowSpark is young, and nearly all of its playtesting so far has been done by the person who designed it. That is a poor basis for any claim about whether a puzzle is good.

I need players who don't know what I intended. People who will tap the wrong arrow, find a level delightful or tedious, and discover that something I thought was obvious isn't. That's the one kind of feedback I can't generate myself.

You can weigh in from either side. Play the game, judge whether it holds up, and tell me through the survey. Or open the repository and judge the developer: the specs, the commits, the tests, the decisions, and the things that didn't work.

Does shipping this game make me a game developer? I have a working game and an incomplete answer, and I'd like your verdict.

Explore More

Working through a similar architecture decision?

If this article maps to a problem in your system, send a short note with the constraint, the risk, and what decision is blocked.