Better Prompts Aren't Enough for Reliable AI Coding
Oct 05, 2026
When a coding agent gives you a confident answer, it's easy to treat that confidence as evidence.
I learned that on SkillPass, one of my open-source projects. An agent told me the GitHub check on a pull request was green. I trusted the report and merged it without opening the check myself.
The check had failed. Main was red, and we had to repair the lint problem in the next pull request.
Nothing catastrophic happened. What mattered was seeing exactly where I let the process fail. The agent gave me an answer, and I treated it like proof. That one is on me. I'm the person who clicked merge.
A better prompt might have produced a more accurate report. It wouldn't have replaced the need to inspect the check, define what evidence was required and prevent a failed check from being merged.
Good prompts still matter, but they're only one part of the development process.
None of this is new. Developers have always had to manage scope, document decisions and test their work. Coding agents haven't changed that. Their speed can make those steps easier to skip while the project still looks like it's moving forward.
A Good Prompt Handles the Next Interaction
A useful prompt should describe the outcome, point to the relevant context and define the check you expect the agent to run.
For a small change, that may be enough. A typo, a spacing adjustment or a focused bug fix does not need a formal planning process.
The problem starts when the same prompt-first approach is stretched across a longer project.
A project has to survive a context reset. It may need to move between agents. I might forget why I made a decision several weeks earlier. Chat history and larger context windows help, but the work shouldn't depend entirely on what one tool remembers.
If a decision matters later, it should live somewhere the next agent will actually read.
Important Context Has to Outlive the Chat
One of the early problems that shaped AI Blueprint, the open-source workflow I built for coding agents, was a disagreement about scaffolding.
Blueprint was meant to be added after an application had already been scaffolded. One agent treated the Blueprint directory as the place where the application should be created. Another followed the intended app-first process.
The correct decision existed in the README, but it was not in the project context both agents loaded when they started. I moved the rule into the files they read first.
A cleverer prompt wouldn't have solved that. I needed to put the decision somewhere both agents would see it.
This applies beyond any particular workflow. Repository instructions, feature notes or a short decision record can all work. The important part is that the information survives the conversation and shows up when the work starts.
Give Every Feature a Finish Line
AI makes it easy for a feature to grow while it is being implemented.
You ask for a password-reset form. The agent notices an auth module it could clean up. It changes an email template, moves a shared utility and updates another component to match the new pattern.
Those changes might be reasonable. That doesn't mean they belong in the same feature.
Before implementation starts, define a small finish line:
- Scope: the password-reset request form and reset flow.
- Outside the scope: a broader authentication refactor.
- Acceptance: a valid link works, while an expired or reused token fails safely.
- Verification: run the relevant tests, then try each path in the application.
This isn't a giant specification. It's enough information to keep the agent and the developer working toward the same result.
Without it, review becomes a series of corrections. The agent broadens the change, so you narrow it. It makes an assumption, so you undo it. After several rounds, the original feature is difficult to see inside the diff.
Match the Evidence to the Claim
A green build doesn't prove that every part of a feature works. A unit test provides different evidence than using the feature in a browser.
If the agent says a bug is fixed, reproduce the original bug. If it changes an interface, open it and use it. If it touches authentication or user data, test the failure cases as well as the happy path.
The level of verification should match what the change can break. A small visual adjustment doesn't need the same process as a payment or authentication change.
It also matters who produced the evidence. If the same agent writes both the implementation and its tests, both can reflect the same incorrect assumption. Passing tests are useful evidence, but they are not a substitute for understanding what was tested.
Keep Consequential Actions Separate
Permission to implement a feature should not automatically mean permission to merge it. A passing test should not automatically mean the work is accepted. A completed branch should not silently become permission to deploy or publish.
These boundaries may feel unnecessary when everything is going well. They become useful when the scope was misunderstood or the evidence is uncertain.
You don't need to type every line of code to remain responsible for the result. Someone still has to decide that the evidence is good enough and that the work should move forward.
Use Only as Much Process as the Work Needs
The strongest argument against a structured workflow is that it can become busywork.
You can spend more time writing plans and updating status files than it would take to make the change. Old instructions can accumulate. Agents can waste context reading documents that have nothing to do with the current task.
That's a valid criticism.
A throwaway prototype may need only one clear prompt. A focused bug with a reliable reproduction may need one regression test. A long-running feature that touches user data needs more structure.
Use enough process for the risk, scope and lifespan of the work.
A Small Workflow You Can Use
You don't need a framework to apply these ideas. For the next feature you build with an agent:
- Define the exact outcome and what is outside the scope.
- Write down the acceptance criteria before implementation begins.
- Put important project decisions somewhere the agent will load them.
- Run the checks that match the claim, then use the changed behavior yourself.
- Review the evidence before approving a merge, deployment or publication.
Better prompts still belong in that process because they help the agent understand the next task. But they can't preserve every decision, control every action or prove that the finished work is correct.
When I merged the failed SkillPass check, the prompt wasn't the final authority. I was. That's the part of AI-assisted development I don't want to lose.
Stay connected with news and updates!
Join our mailing list to receive the latest news and updates from our team.
Don't worry, your information will not be shared.
We hate SPAM. We will never sell your information, for any reason.