# Pre-Launch Creative Scoring: Can You Predict a Winner Before Spending?

> Every creative tool gets asked if it can rank your ads before you spend. The honest answer is no -- and what's left once you drop the prediction claim is more useful than what was promised.
- **Author**: Marxx
- **Published**: 2026-09-26
- **Category**: Creative strategy
- **URL**: https://marxx.ai/posts/pre-launch-creative-scoring

---

Every creative tool gets asked the same question in the demo: *can it tell me which of these five ads will win before I spend on them?*

It's the right question. A Meta learning phase on five variants is real money and two weeks you don't get back. If software could rank them first, that's the whole product.

Here's the honest answer, and it's the one we give on our own product page: **no.** We tested it. Creative embeddings, scored against outcomes, the standard approach. Once brand and objective base rates were controlled for, the signal did not hold.

That's worth sitting with, because most of the category will tell you otherwise. What follows is what "pre-launch scoring" actually means once you strip the prediction claim out of it -- and why the honest version is more useful than the version that was promised.

## Why prediction doesn't survive contact with base rates

A model looks at ten thousand ads, learns which ones performed, and scores your new one. Sounds sound. The problem is what the model is really learning.

It learns that ads from big brands do better. It learns that retargeting ads outperform prospecting ads. It learns that certain categories convert at multiples of others. Feed it your new static and it will confidently tell you the ad is strong -- because you're a known brand running a mid-funnel objective in a high-intent category, which is a fact you already knew and paid nothing to discover.

Strip those out, and the creative signal that remains is thin. Not zero. Thin enough that acting on it costs more than ignoring it.

Look at what this means in practice. One brand, one product, three ads:

<div style="display:flex;align-items:flex-start;gap:8px;overflow-x:auto;padding-bottom:6px;margin:8px 0">
<img src="https://marx-ad-assets.s3.amazonaws.com/ads/110590742060085/2090922951835830/251e009c.jpg" alt="PetChef Meta ad: the founder eats his own dog food on camera to prove its quality" style="flex:0 0 32%;height:auto;border-radius:8px">
<img src="https://marx-ad-assets.s3.amazonaws.com/ads/110590742060085/1438452574361347/e04f84dc.jpg" alt="PetChef Meta ad: a comedic skit asking whether you would add harmful ingredients to your own food" style="flex:0 0 32%;height:auto;border-radius:8px">
<img src="https://marx-ad-assets.s3.amazonaws.com/ads/110590742060085/1505037301381590/6159a1ec.jpg" alt="PetChef Meta ad: India&#39;s biggest companies have entered the dog food market, and the buzz is all about cheaper dog food" style="flex:0 0 32%;height:auto;border-radius:8px">
</div>

Same brand, same product, same claims -- preservative free, human grade, freshly made. Three different mechanisms: a founder proof stunt, a comedy skit, a category-news hook. They ran **30 days, 14 days and 4 days** respectively.

Now ask yourself honestly: before launch, which one would you have scored highest? The founder eating dog food on camera is the kind of idea that either lands or is deeply strange, and no pre-launch model was going to tell you which. It ran seven times longer than the one built on a genuine news event.

That spread -- 4 to 30 days inside one account, one week -- is the thing a predictive score has to beat. It mostly can't.

## What you can actually know before you spend

Prediction is the wrong frame. **Elimination** is the right one. You can't rank five ads by future performance, but you can usually find the two that shouldn't launch -- and that's most of the money saved.

Five checks, all answerable in ten minutes, none of them requiring a model:

### 1. Has this hook already run in this account?

The most common pre-launch failure isn't a bad idea. It's an idea the account already tested in March and everyone forgot. If your ads are decoded by hook type, this is a filter, not an archaeology project.

### 2. Is the hook already the category's default?

Here's the same claim from three different brands in one category:

<div style="display:flex;align-items:flex-start;gap:10px;flex-wrap:wrap;margin:8px 0">
<img src="https://marx-ad-assets.s3.amazonaws.com/ads/100805262270782/929787656344150/c778add0.jpg" alt="Benny&#39;s Bowl Meta ad: India&#39;s best fresh dog food, fresh and nutritious meals for dogs" style="flex:0 0 46%;max-width:320px;height:auto;border-radius:10px">
<img src="https://marx-ad-assets.s3.ap-south-1.amazonaws.com/100506976019555/2525685487766867/thumbnail_0.jpeg" alt="Blep Meta ad: text overlays questioning the safety of common pet food ingredients -- natural, no sugar, no preservatives" style="flex:0 0 46%;max-width:320px;height:auto;border-radius:10px">
</div>

Benny's Bowl, Blep and PetChef are all running "no preservatives, real ingredients, fresh". Three brands, one argument. If your fourth ad in that category leads with clean ingredients, it isn't a hook -- it's the price of entry, and the viewer has already scrolled past it twice this week.

This is a pre-launch check, and it's a hard fail. No model needed: search the category, count how many competitors are already saying it.

### 3. Does the format match the base rate you actually have?

<div style="margin:8px 0">
<img src="https://marx-ad-assets.s3.ap-south-1.amazonaws.com/236672023872356/24477442135264452/image_0.jpg" alt="Snitch Meta ad: a clean high-contrast static of a man wearing sunglasses against a clear blue sky, with an app discount code" style="width:100%;max-width:380px;height:auto;border-radius:10px;display:block;margin:0 auto">
</div>

Snitch's sunglasses line runs statics and videos side by side. This clean product static has run **107 days**. Their high-energy video cuts for the same collection ran four. A creator-led piece on face shapes ran 36.

That isn't "statics beat video". It's *this brand's* base rate for *this product*, and it's the only base rate that should inform a pre-launch decision. A global model that has learnt "video outperforms static" would have scored the 107-day ad down.

Before launch, the question isn't "is this format good". It's: **in this account, at this funnel stage, what has this format historically been worth?**

### 4. Does the offer have a named job?

An offer is a hook, a reward, an accelerator, or a hesitation-reducer. It should have exactly one of those roles, and you should be able to say which before the ad is built.

<div style="margin:8px 0">
<img src="https://marx-ad-assets.s3.amazonaws.com/ads/1598012850481507/1367261102136277/8693e4b5.jpg" alt="Porcellia Meta ad: a force-versus-sharp-knife metaphor contrasting brute-force marketing with precision, offering the first 10 pages of a growth guide free" style="width:100%;max-width:380px;height:auto;border-radius:10px;display:block;margin:0 auto">
</div>

"Get the first 10 pages of the guide free." The offer's job here is unambiguous: it's a hesitation-reducer for a cold B2B audience that won't hand over an email for a full download. The ad has been **live 96 days**.

If you can't name the offer's job in one word before launch, the ad usually has two competing arguments in it, and that is visible on a sheet of paper.

### 5. Are you testing five ads, or one ad five times?

<div style="display:flex;align-items:flex-start;gap:8px;overflow-x:auto;padding-bottom:6px;margin:8px 0">
<img src="https://marx-ad-assets.s3.amazonaws.com/ads/111430194744033/975447548558579/ab88f207.jpg" alt="vamayukt Meta ad: Infinity Love, Modern Style -- Rubykara minimal mangalsutra with free delivery and 50% off" style="flex:0 0 32%;height:auto;border-radius:8px">
<img src="https://marx-ad-assets.s3.amazonaws.com/ads/111430194744033/2492515577848567/17052960.jpg" alt="vamayukt Meta ad: Elegance that shines in every look -- three-layer mangalsutra with free delivery and 50% off" style="flex:0 0 32%;height:auto;border-radius:8px">
<img src="https://marx-ad-assets.s3.amazonaws.com/ads/111430194744033/1617444716574774/a0bcbf64.jpg" alt="vamayukt Meta ad: the relatable struggle of wearing heavy uncomfortable jewellery to work" style="flex:0 0 32%;height:auto;border-radius:8px">
</div>

Three ads, 11, 7 and 4 days. Two of them are product beauty shots with an aesthetic line. The third -- the discomfort of wearing heavy jewellery to the office -- is a genuinely different argument about a real moment.

Only the third one can teach you anything the other two can't. A five-ad test where four share a hook type and an emotion is a one-ad test with a bigger invoice.

This is the check that scoring tools are genuinely good at, and it has nothing to do with prediction. Decode your five ads across hook type, offer type, ad type, emotion, social proof type, authenticity type and intended audience, and duplication is visible instantly. If four rows are identical, you don't have a test.

## The pre-launch scorecard

Ten minutes per batch, before anything goes live:

| Check | Fail condition | What it saves |
|---|---|---|
| **Hook novelty, account** | This hook type ran here in the last 180 days | Re-testing a settled question |
| **Hook novelty, category** | Three or more competitors are running the same argument | Paying to be ignored |
| **Format base rate** | Format underperforms in this account at this stage | Learning-phase spend on a known loser |
| **Offer job** | You can't name it in one word | An ad with two arguments |
| **Test spread** | Two or more variants share hook type and emotion | A test that can't produce a learning |
| **Attribute coverage** | The batch leaves a dimension untested | A quarter of guessing |

Note what's missing from that table: a number between 1 and 100 predicting performance. Every row is a fact you can verify, not a forecast.

<div style="font-family:-apple-system,Helvetica,Arial,sans-serif;background:#0f1218;border-radius:12px;padding:20px 22px;margin:14px 0;color:#e8eaf0">
<div style="font-size:11px;letter-spacing:.14em;text-transform:uppercase;color:#8b93a7;font-weight:700;margin-bottom:12px">What the score is, and what it isn't</div>
<div style="display:flex;gap:18px;flex-wrap:wrap">
<div style="flex:1 1 230px;background:#1d1518;border-left:3px solid #d1566b;border-radius:8px;padding:13px 15px">
<div style="color:#ff93a4;font-weight:700;font-size:14px;margin-bottom:7px">Sold as</div>
<div style="font-size:13.5px;line-height:1.6;color:#d8c3c8">"This creative scores 82. Launch it." A single number, no stated basis, no way to be wrong out loud.</div>
</div>
<div style="flex:1 1 230px;background:#13201b;border-left:3px solid #2f9e6e;border-radius:8px;padding:13px 15px">
<div style="color:#7fd1a8;font-weight:700;font-size:14px;margin-bottom:7px">Actually useful</div>
<div style="font-size:13.5px;line-height:1.6;color:#bfd4c9">"Three of these five share a hook type you last used in April, when it ran nine days. Two competitors are running the same claim. Nothing in this batch tests a new emotion."</div>
</div>
</div>
<div style="margin-top:13px;font-size:13px;color:#8b93a7;line-height:1.55">The second one is checkable. That's the entire difference -- you can find out it was wrong.</div>
</div>

## What replaces prediction: designing the test properly

If you can't know the winner in advance, the return comes from making the test cheap and fast rather than making the guess smarter.

1. **Test mechanisms, not wording.** Four headline variations of the same idea produce one learning. Four different human tensions produce four.
2. **Make one variant deliberately uncomfortable.** The founder eating the dog food is that variant. Batches where every ad is safe never surprise you, and a test that can't surprise you isn't worth running.
3. **Set the kill rule before launch, in writing.** How many days, at what spend, against which metric. Deciding this after you've seen the numbers is how a 4-day ad gets nursed for three weeks.
4. **Let the long-runners write the next brief.** Anything past 60 days in your account or your category is a verdict. That's a stronger input than any model output.
5. **Record what the ad was testing, not just what it said.** Six months on, the hypothesis is the asset. The JPG isn't.

Once the test has run, the useful question changes from "will this win" to "is this decaying" -- and *that* is measurable. Fatigue is a weighted composite of click-through rate, CPM, ROAS and frequency, with CTR carrying the most weight and frequency the least, sorted into five levels from healthy to critical. A falling creative has a shape. A future creative doesn't.

## Where Marxx fits

Marxx decodes ads along seven dimensions -- hook type, offer type, ad type, emotion, social proof type, authenticity type and intended audience -- across your own full creative history from the connected account, not just the last thirty days, and across competitors in your category.

That gives you the six scorecard rows above as lookups rather than memory. It flags fatigue across five levels, and names eight diagnostic states when performance moves: offer fatigue, burnout starting, high CTR with low sales, costs rising but still working, clicks too expensive, only existing buyers left, wrong audience, and cold start failure.

What it does not do is hand you a number predicting a winner. We tried; the signal didn't survive base-rate control, and shipping it anyway would have been a nicer demo and a worse product.

**[Book a demo](/book-a-demo)** and we'll decode your own last quarter of creative live -- hook types, duplication, and what your account has never tested.

## FAQs

### Can AI predict ad performance before launch?

Not reliably. Models trained on creative features mostly learn brand size, objective and category -- and once those base rates are controlled for, the remaining creative signal is too thin to act on. Tools that claim otherwise are usually reporting a base rate back to you as a creative insight.

### So what is pre-launch creative scoring good for?

Elimination rather than ranking. It catches hooks you've already tested, claims your whole category is already making, formats that underperform in your account, and batches where several variants are secretly the same ad.

### How many creatives should I launch in a test?

Fewer than you think, and more different from each other than you think. Five ads sharing a hook type teach you less than three that each test a distinct human tension.

### What's the most common pre-launch mistake?

Duplication that nobody notices. Four of five variants share a hook and an emotion, the test returns one learning, and the budget bought four copies of it.

### Is longevity a reliable signal?

It's the most honest public one. You can't see a competitor's spend, but you can see how long an ad has run and how many near-identical copies are live. A 107-day static is a verdict; a 4-day one is a question.

### When should I kill a new creative?

At the threshold you wrote down before launch. The number matters less than having set it while you had no numbers to rationalise around.

### Does creative scoring replace testing?

No. It makes the test worth running. The score tells you what's redundant; the test tells you what works.

---
- [More Creative strategy articles](https://marxx.ai/posts/category/creative-strategy)
- [All articles](https://marxx.ai/posts)