Four words carry the whole message, SHIPPED, PARTIAL, STALLED and NEEDS INPUT, with no softer option available and no fifth. Alongside them sit refusals: no estimating what is left, no describing a change it cannot point at, no starting or terminating a session. Pointing OpenCode at a discipline this narrow is quick, and how firmly it holds is down to the engine you put behind it.
Run a staged conversation — no account needed. The agent handles it for real while a simulated world answers its tool calls; nothing touches real accounts, and nothing is actually sent.
Migrate the notifications service off Moment.js to date-fns
Ran 47 minutes. Opened pull request 318 in fleetwing/notifications across 61 files, CI green. The session summary calls the migration complete, while a later message notes one helper, the business day offset used by the digest scheduler, still calls Moment because date-fns carries no direct equivalent. The run ended with the pull request left as a draft.
Using this template drops you into a guided setup. It asks exactly this, nothing else:
Connect Devin
One sign-in. The agent acts through your account, scoped to what this template uses.
Connect Slack
One sign-in. The agent acts through your account, scoped to what this template uses.
Review rules
What finished actually means to you: what a session has to produce before you would call it shipped, the shortcuts you refuse to accept as a fix, whether the agent may message a stalled session back, and who owns the queue when one goes wrong.
Runs on OpenCode
Preselected for this page — connect your OpenCode account during setup, or switch to NoClick's built-in models with one click.
Watch it handle a test run
A staged conversation against a simulated world — then it’s live.
A verdict vocabulary only works if it is never quietly widened. Most drift shows up as encouraging prose wrapped around the word, which is a thing you can read and correct.
Because the engine is swappable, two models can be pointed at the same finished sessions and their verdicts compared directly. The rules do not move, so the difference you see is the model.
That is exactly where a weaker one slips, usually by softening a stall into something hopeful or by inventing a fifth outcome in prose. Being model agnostic makes this cheap to check: point two models at the same three Test Runs, including the session solved by deletion, and read both before either reaches your channel.
Free to start. Guided setup, a test run against staged conversations, and it's live.