creative-testing-os: the creative verdict board
What it does
Section titled “What it does”Decides which ads live and which die, and shows its work. It pulls ad-level performance, applies the winner math, and renders a board: scale / iterate / kill / hold.
The winner math is the part that matters. An ad earns a verdict only once it has cleared both legs: at least 10x the account’s median ad spend and at least $500 total. Below that, it goes to hold and is explicitly not judged. That is what keeps you from killing a winner on 80 dollars of noise, which is the most expensive habit in paid media.
It is sunk-cost-free in the other direction too: an ad you love with a real verdict against it gets killed, and a suspiciously high hit rate gets flagged as under-testing rather than celebrated.
When to reach for it
Section titled “When to reach for it”- Weekly, to decide next round’s launches.
- Before a creative sprint, so you brief against what actually worked.
- When someone asks “is this ad a winner?” and you want an answer with a floor under it.
- When your hit rate looks great and you suspect you are not testing hard enough.
Where it runs
Section titled “Where it runs”Both. It reads from the Meta connector, a Motion CSV, or an Ads Manager export.
How to run it
Section titled “How to run it”“Which ads should we scale or kill for examplebrand?”
“/um-toolkit:creative-testing-os examplebrand, weekly board”
What it needs from you
Section titled “What it needs from you”- Ad-level performance data, from any of the three sources above.
- The account, confirmed by platform-reported name.
- (Helpful) Your concept and angle labels, so it can roll verdicts up beyond individual ads.
What you get back
Section titled “What you get back”- The board: every ad with its verdict and the spend that earned it.
- Concept and angle rollups, so you learn which idea is working, not just which file.
- Next-round briefs for the iterate column.
- A capacity plan: how many creatives your spend tier should be launching per week, and whether you are hitting it.
A worked example
Section titled “A worked example”You: “Weekly creative readout for examplebrand.”
Board: 2 scale, 1 kill, 3 iterate, 9 hold (under floor, not judged). Angle rollup: the “upkeep, not price” territory is carrying both scale verdicts across two different visual concepts, so the winner is the angle, not the execution. Capacity flag: at $34k/mo you should be launching roughly 12 creatives a week; you launched 4.
Tips & gotchas
Section titled “Tips & gotchas”- Hold is not a soft kill. It means the ad has not spent enough to say anything. Killing holds is how accounts end up with nothing but small, safe, mediocre ads.
- Both spend legs are required. 10x median alone is not enough on a low-spend account, and $500 alone is not enough on a high-spend one.
- A high hit rate is a warning. If 80% of your ads are winners, your tests are too similar to each other.
- It never touches the account. Launch and budget changes route through the approval lane.
Related skills
Section titled “Related skills”Runs inside weekly-account-os. Feeds next round through concept-lab and hook-matrix-forge. Execution goes to meta-scaling-copilot.