“Done.” Prove it.

23 skills for Claude Code and Codex that make your coding agent prove its work before it says done.

SWE Stack by Bles Software. One-time payment, free updates for every future version, instant download.

gate
  1. agentFixed. All checks pass. Done.
  2. gatererunning the readback the agent recorded
  3. gateexit 1: it did not pass
  4. gatedone claim sent back with the output

Measured on our own agents, 24 to 27 Sep 2026

36 times our agent said “pass” and the gate’s own rerun did not confirm it.

Each square is one QA check. Marked squares are claims the gate rejected.

real requests
103
QA checks
1,156
“pass” claims rejected
36

30 readback commands that failed when the gate ran them

6 screenshots the overlap scan flagged (3 were false alarms, cleared by opening the crop)

23 skills. Five layers. One rule under all of them.

Each skill is a SKILL.md folder your agent loads, and many ship the scripts and thresholds that do the checking. Not prompts: gates that came from real failures.

Layer 1

Plan and prove

Every request becomes rows you can check, and nothing is done until each row has proof.

  • prd-to-proofTurns a request into numbered requirements, then proves every row on the real product.
  • request-gateA mechanical gate. It says ok only when every row passed a screenshot scan or a readback it ran itself.
  • long-horizon-buildPlans multi-day builds into 15 to 60 minute milestones, each with a validation command.
  • iterative-loopOne focused change per round, the same honest check, keep only clear wins.
  • proof-principlesThe working rules: verified is 1, unverified is 0. Load it always.
Layer 2

Catch breakage

Find what a change breaks outside its own diff before your users do.

  • what-could-breakFinds what a change could break outside its diff, then runs real code to prove the one fact that makes it safe.
  • hard-reviewA strict review where every claim comes with a command the author can rerun.
  • make-a-verifierGenerates a verify skill for your repo that drives the real app the way a user does and captures proof.
Layer 3

Visual and video QA

Check pages, screens and video the way a person sees them, not the way the markup reads.

  • visual-qaOpens the rendered result, judges it like the viewer, and scans screenshots for text drawn over text.
  • vision-qcChecks generated graphics at scale with any vision model and gets a strict JSON verdict back.
  • mobile-no-zoomStops a web app zooming on phones and proves it in a phone-emulated browser.
  • smooth-web-videoDrives the page frame by frame instead of screen recording, and gates every video for stutter.
Layer 4

Production bugs

The bugs that only show up with retries, several workers and real users.

  • idempotent-retriesFixes duplicate posts, double charges and messages sent twice after timeouts, retries and restarts.
  • single-writer-stateStops workers, processes and agents overwriting each other's files, rows, keys and branches.
  • prompt-override-debuggingFixes an LLM feature that ignores the user because its system prompt pushes the other way.
  • multilingual-text-debuggingTraces broken non-English text: stray punctuation, invisible direction marks, words cut in half.
  • error-to-fix-loopWatches a service log and hands each new code error to an agent that must reproduce it before fixing it.
Layer 5

Agent engineering

Build the agents and deciders themselves so they measure, not guess.

  • data-kernelAsks questions of big logs and JSON without flooding the agent's context.
  • llm-judge-evalsLLM-as-judge evaluations that give the same answer twice.
  • autoresearch-loopUnattended optimization of one asset against one number, one change per round, every round logged.
  • agent-cli-fallbackCalls claude -p or codex exec from a service, with a fallback provider and schema checks.
  • fast-decider-architecturesPuts a cheap fast model in front of a frontier model, with a gate that never serves a lower score.
  • skill-forgeWrites or refines your own skills to a standard that holds up.

Try one before you buy: what-could-break

Point it at your uncommitted change. It names what the change does, the one fact it is safe because of, and then it runs the code to check that fact.

Download what-could-break (free)
~/shop
> use what-could-break on my uncommitted change
What it does: total() now reads i["amount"] instead of i["price"].
The one fact it's safe because of: none, the only caller still passes "price" dicts.
Proof (ran it): app.py:2 calls total([{"price": 2}, {"price": 3}]) -> KeyError: 'amount'
Before you ship: revert the key, or update app.py:2.

New, free, for front‑end work: tastegate

Taste and impeccable in one skill, plus a gate that opens your page in a real browser at phone and desktop width. Same brief on Opus 5.5, five pages each: impeccable’s own detector found 29.2 flaws a page without it and 2.2 with it.

~/tidewell
> build the landing page, use tastegate
FAIL equal_icon_cards: 3 same-size icon + heading + text cards in a row
FAIL low_contrast: avatar initials 1.23:1 on teal, needs 4.5:1
FAIL small_tap_target: footer links 19px tall on a phone, needs 24px
Fixed in one batch, gate run again, both screenshots opened.
tastegate PASS: 0 fail, 0 warn at 390px and 1440px.

Works with

  • Claude Code
  • Codex

Not for

  • People who want a magic button.
  • Teams that won't run a check before merging.
  • ChatGPT copy-paste prompt collectors.

SWE Stack v1

$99

One-time. Free updates for every future version. Instant download.

  • All 23 skills, ready for Claude Code and Codex
  • The scripts inside: Python 3.10 standard library or Node 18, nothing to sign up for
  • Use it in your own projects and your clients' projects
  • If it does not install and load, we refund you. Ask within 14 days.
Buy SWE Stack, $99

Questions

Do I need Claude Code?

No. Tested with Claude Code, and Codex reads the same SKILL.md folders. Other agents can read a SKILL.md as plain instructions; we have not tested them.

What do the scripts need?

Python 3.10 or newer with the standard library only, or Node 18. The screenshot scans in visual-qa and request-gate need the tesseract binary. smooth-web-video needs ffmpeg and Playwright.

Can I use it for client work?

Yes, in any number of your own and your clients' projects. Please do not resell or republish the pack.

Is it just prompts?

No. It is gates, scripts and thresholds that came from real failures in our own agents. The gate runs the check itself instead of trusting the agent's word.

What if it doesn't work for me?

If it does not install and load, we refund you. Ask within 14 days.

Can you set this up for my team?

Yes. Bles Software sets up proof-driven agent workflows for engineering teams. Book a team setup call.

Want it running in your team?

Bles Software sets up proof-driven agent workflows for engineering teams and builds fast AI deciders: a small cheap model that makes the routine calls, with a frontier model behind it for the hard ones.

Book a team setup call