Personal Software Factory
How I built a workflow that moves software work from idea to evidence without handing judgment to the machine
Adam Shupe — author and operator
- program:
- Personal Software Factory
- work:
- Operational
- record:
- Published
- standard:
- Contemporaneous
May 22 – August 6, 2026 · selected source window; the broader origin period remains approximate
System
What it is
A session would solve the problem and forget why I had asked. I built the Personal Software Factory to stop losing the reason. The method is Operational today: request, work, evidence, and my decision stay on one chain, across several projects.
Commissioner Dev Agent is the local software; it holds the material on disk. The factory is the method around it.
The software's first build lands on May 22, 2026 — the earliest date I can point to, not the day the method began, and not a day it worked well.
The trouble was never volume. It was judgment.
How it works
I start with one bounded goal and a list of what may change. That list becomes the prompt. The work happens in the named project; validation records what was checked; a review holds the result against the request and the paths I allowed. Only then do I decide: commit, complete, or archive.
Every handoff loosens the one before it. A ticket loses its boundary in a prompt. An implementation loses its reason in a patch. Passing checks start standing in for the decision that the work is done. Held together, the chain is something I can look at when I am unsure.
It did not hold at first. Stale context. File targets no one had pinned down. Closeout evidence missing. Repairs on July 16 and July 27, 2026 patched closeout, evidence, and scope attribution — repairs, not a clean line between a bad era and a good one.
What changed was what I thought it was for. The factory keeps judgment; it does not automate it. Leave the uncertainty visible and let a project move without losing its history.
Sanitized representation assembled from a real completed run. This is not the full original packet.
Redaction statement: absolute local paths, user and machine names, project and ticket identifiers, unrelated prompt material, account or usage information, private repository state, internal operator notes, and confidential or personal data were omitted. The surviving sequence supports only the claim that the method preserved a chain from goal and bounded authority through prompt, validation, human review, and closeout.
Goal — Repair a closeout rule so a documentation-only design artifact could finish through its declared checks without inventing a focused test, while source and runtime work remained subject to focused validation.
Bounded authority — Changes were limited to validation-classification source and its focused tests. Product repositories, unrelated lifecycle commands, and commit, push, completion, and archive actions remained outside the implementation authority.
Implementation prompt — The handoff required fresh context, a named repository boundary, inspection of the existing probe, a closed set of allowed paths, independent test evidence, and a fail-safe result when eligibility could not be established. Raw prompt text is omitted.
Implementation — The classifier admitted an allowlisted non-code design case only when deterministic repository evidence found no applicable focused test. An inconclusive probe continued to require focused validation.
Validation evidence — Focused readiness checks, focused closeout checks, the full test suite, the build, and the diff check passed. Raw command output and machine-specific details are omitted.
Human review decision — The review found the change inside its bounded authority and found that the stricter source and runtime path remained enforced. The decision still required a human; the packet did not approve itself.
Completed and archived — After operator action, the preserved closeout retained the goal, authority, prompt, implementation summary, validation evidence, diff, and review together. No original identifier or private repository detail is reproduced here.
One completed run, sanitized for reading · August 6, 2026
A text representation I prepared from a completed factory run; it is not the full original packet.
The surviving sequence supports only the claim that the method preserved a chain from goal and bounded authority through prompt, validation, human review, and closeout.
Redaction statement: absolute local paths, user and machine names, project and ticket identifiers, unrelated prompt material, account or usage information, private repository state, my internal notes, and confidential or personal data were omitted.
What it costs
The factory became a burden before it became useful.
For roughly two weeks I did little but repair and simplify it. Readiness checks that were not ready. Closeouts that failed. Threads that crawled. Patch-based edits that fought me. For a stretch most tickets turned up another factory defect before they moved their own project — the factory heavier than the work it carried.
The defects were ordinary and relentless. Stale context bent a pass. A list of likely files could be read as permission. An empty diff could be filed wrong, and completion evidence could go missing.
Underneath was something worse. A factory that grades its own prompts, its own validation, its own closeout cannot treat its own packet as proof. Somewhere it needs a decision from outside itself.
Limits became part of the day. I ran Claude, Codex, and Cursor out of capacity working in parallel, and the better the factory got, the faster the more capable models burned through what I had. Routing and bounded tickets became constraints, not preferences.
The overhead has not gone away. The method recommends and prepares. Every decision that matters still waits on me.
What is unbuilt
There is no autonomous loop, and there was never going to be. No automatic commits or pushes, no deployment, nothing that marks work complete, no scheduled background runs, no reaching into a target repository. Exclusions, not a roadmap dressed as a deficit.
Commissioner Dev Agent may recommend work, prepare artifacts on my machine, inspect the projects I configure, and gather evidence. It does not choose policy, act on its own recommendation, or decide when work is finished. Ambiguous evidence stays unknown instead of becoming permission.
The line holds because the factory inspects its own machinery. A machine can tell me what the evidence contains. Not what risk to accept, or what the work means.
The gate is not missing functionality. Take it out and I have a different system, not a finished one.
RECORD NOTES
- source:
- Contemporaneous — working artifacts were created while the method operated; later synthesis is identified by date and speech class
- evidence:
- One sanitized text representation assembled from a completed run.
- claim boundary:
- No speed, quality, throughput, cost-saving, business-result, or professional-value claim is made.
- publication:
- Published August 11, 2026 · Revised —