Skip to content

creative-testing-os: the creative verdict board

Run the week Cowork-ready

Decides which ads live and which die, and shows its work. It pulls ad-level performance, applies the winner math, and renders a board: scale / iterate / kill / hold.

The winner math is the part that matters. An ad earns a verdict only once it has cleared both legs: at least 10x the account’s median ad spend and at least $500 total. Below that, it goes to hold and is explicitly not judged. That is what keeps you from killing a winner on 80 dollars of noise, which is the most expensive habit in paid media.

It is sunk-cost-free in the other direction too: an ad you love with a real verdict against it gets killed, and a suspiciously high hit rate gets flagged as under-testing rather than celebrated.

  • Weekly, to decide next round’s launches.
  • Before a creative sprint, so you brief against what actually worked.
  • When someone asks “is this ad a winner?” and you want an answer with a floor under it.
  • When your hit rate looks great and you suspect you are not testing hard enough.
Cowork ✓ Desktop ✓

Both. It reads from the Meta connector, a Motion CSV, or an Ads Manager export.

“Which ads should we scale or kill for examplebrand?”

/um-toolkit:creative-testing-os examplebrand, weekly board”

  • Ad-level performance data, from any of the three sources above.
  • The account, confirmed by platform-reported name.
  • (Helpful) Your concept and angle labels, so it can roll verdicts up beyond individual ads.
  • The board: every ad with its verdict and the spend that earned it.
  • Concept and angle rollups, so you learn which idea is working, not just which file.
  • Next-round briefs for the iterate column.
  • A capacity plan: how many creatives your spend tier should be launching per week, and whether you are hitting it.

You: “Weekly creative readout for examplebrand.”

Board: 2 scale, 1 kill, 3 iterate, 9 hold (under floor, not judged). Angle rollup: the “upkeep, not price” territory is carrying both scale verdicts across two different visual concepts, so the winner is the angle, not the execution. Capacity flag: at $34k/mo you should be launching roughly 12 creatives a week; you launched 4.

  • Hold is not a soft kill. It means the ad has not spent enough to say anything. Killing holds is how accounts end up with nothing but small, safe, mediocre ads.
  • Both spend legs are required. 10x median alone is not enough on a low-spend account, and $500 alone is not enough on a high-spend one.
  • A high hit rate is a warning. If 80% of your ads are winners, your tests are too similar to each other.
  • It never touches the account. Launch and budget changes route through the approval lane.

Runs inside weekly-account-os. Feeds next round through concept-lab and hook-matrix-forge. Execution goes to meta-scaling-copilot.