# Do Predicted-Performance Scores Work? Testing Ad Scoring Tools Against Real Results

> The same video ran 52 days and 4 days in the same account. Four exhibits against predicted-performance scores, and the five questions to ask any vendor selling one.
- **Author**: Marxx
- **Published**: 2026-10-01
- **Category**: Creative strategy
- **URL**: https://marxx.ai/posts/do-ad-scoring-tools-work

---

Pick any creative tool's homepage and you'll find a version of the same sentence: *know which ads will win before you spend.* Sometimes it's a number out of 100. Sometimes it's a traffic light. Sometimes it's "AI-predicted performance", which is the same thing wearing a coat.

The claim is testable. Almost nobody tests it, including the people buying it.

This is an opinion piece, so here's the opinion up front: **predicted-performance scores are mostly base rates in a trench coat, and the public record is full of evidence against them.** Not because the vendors are lying -- most believe it -- but because the thing they're predicting has far more variance than the thing they're measuring.

Below is the evidence, the test you can run on any vendor in one meeting, and the part of this category that genuinely works.

## What a "score" is actually made of

Three kinds of tool get sold under one word, and they fail differently.

| Type | How it works | Where it breaks |
|---|---|---|
| **Attribute scorers** | Tags your ad (hook, format, offer, emotion) and scores against what those tags historically did | Two ads with identical tags routinely produce opposite results |
| **Embedding models** | Converts the creative to a vector, scores against outcome-labelled training data | Learns brand size, objective and category; the creative signal is thin once those are controlled |
| **Benchmark comparators** | Compares your ad to a category average | Tells you about the category, not the ad |

All three share a structural problem. They're trained on ads that ran, which means they're trained on ads that *someone decided to run* -- already filtered by human judgment, already skewed toward brands with budget. The model inherits that filter and calls it insight.

## Exhibit A: the same video, 52 days and 4 days

This is the piece of evidence I'd put in front of any vendor.

<div style="display:flex;align-items:flex-start;gap:10px;flex-wrap:wrap;margin:8px 0">
<img src="https://marx-ad-assets.s3.amazonaws.com/ads/108047081396228/1640146824074499/9189dad3.jpg" alt="Foreplay Meta ad: &quot;If you run ads on Facebook, you have to see this&quot; -- a swipe file and ad tracking pitch" style="flex:0 0 46%;max-width:320px;height:auto;border-radius:10px">
<img src="https://marx-ad-assets.s3.amazonaws.com/ads/108047081396228/2712289325773486/9189dad3.jpg" alt="Foreplay Meta ad: the same video, hooked as &quot;Supplement stores. If you run ads on Facebook, you have to see this&quot;" style="flex:0 0 46%;max-width:320px;height:auto;border-radius:10px">
</div>

Same brand. Same video file. Same body copy, word for word. The only difference is that the second one names the audience in the opening line -- "Supplement stores" -- before saying the identical sentence.

One ran **52 days**. The other ran **4**.

Every attribute a scoring tool can see is the same across these two: hook type, ad type, offer type, emotion, social proof, authenticity, intended audience. Any attribute-based model gives them the same score. Any embedding model gives them near-identical vectors -- it's literally the same frames.

The outcome differed by 13x.

Whatever produced that gap -- audience, placement, timing, bid, competition in the auction that week, or a budget decision made by a person on a Tuesday -- it lives entirely outside the creative. That's not an edge case. That's the normal condition of a Meta account.

## Exhibit B: same brand, same format, 216 days against 6

<div style="display:flex;align-items:flex-start;gap:10px;flex-wrap:wrap;margin:8px 0">
<img src="https://marx-ad-assets.s3.amazonaws.com/ads/361216754245455/3341990512643035/e4b8b3d9.jpg" alt="Ace Blend Meta ad: a pilot describes how magnesium improved his sleep quality and morning focus" style="flex:0 0 46%;max-width:320px;height:auto;border-radius:10px">
<img src="https://marx-ad-assets.s3.amazonaws.com/ads/361216754245455/1351033376899590/d2ea0cc5.jpg" alt="Ace Blend Meta ad: a founder-led conversational piece addressing fishy burps from fish oil supplements" style="flex:0 0 46%;max-width:320px;height:auto;border-radius:10px">
</div>

Ace Blend, both real-person testimonial video, both supplements, both the same account.

The pilot talking about sleep has been **live 216 days**. The founder-led piece answering a very specific objection -- fishy burps, a genuine purchase blocker people actually type into search -- ran **6**.

Here's the uncomfortable part: if you handed both to a creative strategist cold, a good number would pick the objection-handler. It does more jobs. It names a real hesitation. It has the founder's face on it. Every heuristic a scoring model encodes says it should be the stronger ad.

It wasn't. The pilot was.

## Exhibit C: the better-written hook lost

<div style="display:flex;align-items:flex-start;gap:8px;overflow-x:auto;padding-bottom:6px;margin:8px 0">
<img src="https://marx-ad-assets.s3.amazonaws.com/ads/128049983718741/1296043619184552/932e24a1.jpg" alt="Qurist Meta ad: &quot;Low sleep score? Fix it tonight&quot; -- sleep gummies with 5,000+ happy sleepers" style="flex:0 0 32%;height:auto;border-radius:8px">
<img src="https://marx-ad-assets.s3.amazonaws.com/ads/128049983718741/2095647684413086/f265e3e9.jpg" alt="Qurist Meta ad: &quot;Thousands already sleep better with these&quot; -- sleep 3x deeper tonight" style="flex:0 0 32%;height:auto;border-radius:8px">
<img src="https://marx-ad-assets.s3.amazonaws.com/ads/128049983718741/2188152551733422/788f4fcd.jpg" alt="Qurist Meta ad: &quot;You didn&#39;t forget how to sleep at 30. Your body changed the rules.&quot;" style="flex:0 0 32%;height:auto;border-radius:8px">
</div>

Three Qurist ads for the same sleep product.

- "Low sleep score? Fix it tonight" -- **live 96 days**
- "Thousands already sleep better with these" -- **live 37 days**
- "You didn't forget how to sleep at 30. Your body changed the rules." -- **ran 4 days**

The third one is, by any copywriting standard, the best line of the three. It intercepts a specific moment, reframes a fear as a mechanism, and doesn't sound like a supplement ad. It's the one I'd have bet on.

It ran four days. The device-metric hook -- "low sleep score", which works because the viewer's watch told them the number this morning -- is still running at 96.

That's not a lesson about which hook is better. It's a lesson about how little a pre-launch opinion, human or machine, is worth against a live auction.

## Exhibit D: 236 days and 13, same brand, same line

<div style="display:flex;align-items:flex-start;gap:10px;flex-wrap:wrap;margin:8px 0">
<img src="https://marx-ad-assets.s3.ap-south-1.amazonaws.com/444025482768886/1254046316160353/thumbnail_0.jpeg" alt="Plix Meta ad: build their perfect box -- any three kids products for Rs 999" style="flex:0 0 46%;max-width:320px;height:auto;border-radius:10px">
<img src="https://marx-ad-assets.s3.ap-south-1.amazonaws.com/444025482768886/708581992000357/thumbnail_0.jpeg" alt="Plix Meta ad: Plix Kids bubble splash body wash, gentle bubbles and deep nourishment" style="flex:0 0 46%;max-width:320px;height:auto;border-radius:10px">
</div>

Plix's kids range: the bundle-builder ran **236 days**, the single-product launch **13**.

This one has a readable explanation -- a bundle at a price point is an offer, a product introduction is an announcement -- and that's exactly the point. The explanation is available *afterwards*. Before launch it's one of six plausible stories, and the model has no way to pick between them either.

## The variance isn't in the creative

Put those four exhibits together and a pattern emerges that should end the prediction conversation.

<div style="font-family:-apple-system,Helvetica,Arial,sans-serif;background:#0f1218;border-radius:12px;padding:20px 22px;margin:14px 0;color:#e8eaf0">
<div style="font-size:11px;letter-spacing:.14em;text-transform:uppercase;color:#8b93a7;font-weight:700;margin-bottom:14px">Days live, within a single brand's account</div>
<div style="margin-bottom:13px"><div style="display:flex;justify-content:space-between;font-size:13.5px;color:#c3c9d6;margin-bottom:5px"><span>Foreplay -- same video, two ads</span><b style="color:#fff">4 -> 52</b></div><div style="height:9px;border-radius:99px;background:#222836;overflow:hidden"><div style="height:100%;width:24%;background:linear-gradient(90deg,#d1566b,#f0a04b);border-radius:99px"></div></div></div>
<div style="margin-bottom:13px"><div style="display:flex;justify-content:space-between;font-size:13.5px;color:#c3c9d6;margin-bottom:5px"><span>Ace Blend -- testimonial video x2</span><b style="color:#fff">6 -> 216</b></div><div style="height:9px;border-radius:99px;background:#222836;overflow:hidden"><div style="height:100%;width:100%;background:linear-gradient(90deg,#d1566b,#7fd1a8);border-radius:99px"></div></div></div>
<div style="margin-bottom:13px"><div style="display:flex;justify-content:space-between;font-size:13.5px;color:#c3c9d6;margin-bottom:5px"><span>Qurist -- one product, three hooks</span><b style="color:#fff">4 -> 96</b></div><div style="height:9px;border-radius:99px;background:#222836;overflow:hidden"><div style="height:100%;width:44%;background:linear-gradient(90deg,#d1566b,#f0a04b);border-radius:99px"></div></div></div>
<div><div style="display:flex;justify-content:space-between;font-size:13.5px;color:#c3c9d6;margin-bottom:5px"><span>Plix Kids -- bundle vs single product</span><b style="color:#fff">13 -> 236</b></div><div style="height:9px;border-radius:99px;background:#222836;overflow:hidden"><div style="height:100%;width:100%;background:linear-gradient(90deg,#d1566b,#7fd1a8);border-radius:99px"></div></div></div>
<div style="margin-top:15px;font-size:13px;color:#8b93a7;line-height:1.55">Each row is one brand, one product line, one account. The spread inside a single account is larger than the spread a scoring model is trying to detect between accounts.</div>
</div>

A model has to beat that. It has to separate ads whose measurable attributes are identical, in accounts where same-creative outcomes already differ by an order of magnitude. That's a hard problem, and "82/100" is not a solution to it -- it's a confident restatement of the category average.

## How to audit a scoring vendor in one meeting

You don't need a data team. Five questions, asked in the demo:

1. **"What was the model trained on, and does it include losers?"** If the training set is only ads that ran long enough to be noticed, the model learnt survivorship, not performance.
2. **"Show me the holdout."** Ask for performance on ads the model never saw, in *my* category. Not a case study -- a confusion matrix, or at least a hit rate against a coin flip.
3. **"Does the score change if I swap the brand logo?"** If a known brand's logo moves the number, you're being sold a brand-equity lookup.
4. **"What's the score's accuracy on ads from one account?"** Cross-account accuracy is mostly base rates. Within-account is the only number that matters, and it's the one nobody reports.
5. **"What would make this score wrong?"** A vendor who can't answer this hasn't tested it. A score that can't be wrong isn't a measurement.

Ask these politely. The answers are usually honest and usually revealing -- most product teams know exactly where the model is weak, and almost nobody asks.

## What this category is actually good at

Rejecting prediction isn't rejecting the tools. The useful half is real, and it's the half that doesn't require a forecast:

- **Duplication detection.** Decode five variants and see that four share a hook type and an emotion. That's a measurable fact, available before launch, and it's the most common way a test budget gets wasted.
- **Account memory.** "This hook ran here in March and lasted nine days." Not a prediction -- a lookup. Teams forget; accounts don't.
- **Category saturation.** Counting how many competitors are already running your argument is arithmetic, not AI, and it kills more bad ads than any score.
- **Fatigue detection.** A decaying creative has a measurable shape. CTR, CPM, ROAS and frequency, weighted, moving over time. This is the one place where modelling genuinely earns its keep -- because you're describing a curve that already exists rather than one that doesn't yet.
- **Coverage gaps.** What the account has never tested is a fact about history, not a forecast.

Note what every item has in common: it's checkable. Someone can come back in six weeks and tell you it was wrong. A predicted score mostly can't be falsified, which is why it survives.

If you want the full pre-launch version of this -- the six checks and the scorecard -- it's in [pre-launch creative scoring](/blog/pre-launch-creative-scoring).

## Our position, including the inconvenient part

Marxx tested predictive scoring. Creative embeddings against outcomes, the standard approach. Once brand and objective base rates were controlled for, the signal did not hold. We didn't ship it.

That cost us a demo slide. It's the right call, and the exhibits above are why: in a world where the same video runs 52 days and 4 days in the same account, a number predicting which creative wins is theatre.

What Marxx does instead is report what is actually knowable -- what ran, what is decaying, what competitors are doing, and what the account has never tested -- decoding creative across hook type, offer type, ad type, emotion, social proof type, authenticity type and intended audience, over the full creative history of the connected account rather than the last thirty days.

**[Book a demo](/book-a-demo)** and we'll decode your last quarter live -- including the duplicates you didn't know you'd shipped.

## FAQs

### Do AI ad prediction tools actually work?

Not for ranking creative by future performance. The public record is full of near-identical ads with wildly different outcomes inside the same account -- which is the condition any predictive score would have to overcome, and generally doesn't.

### Why do scoring tools seem accurate in demos?

Because demos use ads that already ran. Scored retrospectively, a model that encodes brand size and category will look sharp. Ask for a holdout on unseen ads in your own category and the picture changes.

### Is there any creative signal at all?

Some. It's just thin relative to everything else driving the outcome -- audience, placement, auction conditions, budget decisions. Thin enough that acting on the score costs more than ignoring it.

### What should I use instead of a predicted score?

Duplication checks, account history lookups, category saturation counts and fatigue detection. All of them are facts rather than forecasts, and all of them can be shown to be wrong.

### Is longevity a good proxy for performance?

It's the best public one, with limits. An ad that has run 200 days is almost certainly working. A 4-day ad might have been killed for reasons unrelated to the creative. Treat long runs as strong evidence and short runs as weak evidence.

### Should I stop testing creative then?

The opposite. If you can't know in advance, the return comes from making tests cheap, fast and genuinely different from each other -- and from writing the kill rule down before you launch.

---
- [More Creative strategy articles](https://marxx.ai/posts/category/creative-strategy)
- [All articles](https://marxx.ai/posts)