Skip to content
Make Madly.
Creative testing

AI ad creative testing: more concepts, less waste

How to use AI to generate more ad concepts and test them with a small budget: concept batches, naming, thresholds, sample size cautions and decisions.

Published
Published
Last updated
Updated
Reading time
8 min read

Short answer

AI makes it cheap to produce many ad variations, but testing is limited by budget, not by ideas. Test a few concepts that differ in angle, give each enough spend to produce a readable result, decide your thresholds before launch, and judge winners on profit rather than on clicks.

Key takeaways

  • Generate widely, but test only a handful of concepts that genuinely differ.
  • Set the budget, the window and the decision rule before launching.
  • Change one meaningful thing at a time where possible.
  • Treat small differences on small samples as noise.

The bottleneck has moved

Before AI tools, the scarce resource in creative testing was production. Making ten ads took days. Now it takes an afternoon, so founders are tempted to launch all ten. The scarce resource is now spend. Each concept needs enough budget to reach a point where you can read it, and a small store cannot afford that for ten concepts at once. The sensible response is to separate generating from testing: generate widely, then test narrowly.

What counts as a concept

A concept is a distinct reason to buy, expressed in a distinct way. A concept is not a different background colour, a reworded headline with the same meaning, or a different crop. Those are variations. If you test ten variations of one concept you learn very little about what your market responds to. If you test four concepts, you learn which message deserves further investment.

Useful concept axes include the problem addressed, the benefit emphasised, the proof offered, the format used and the audience situation depicted. The ad hooks guide organises examples by angle, and the design ad variants for testing guide explains what to vary within a concept once you have a promising one.

A testing loop that works for small budgets

1. Generate a wide batch

Use AI to produce fifteen to thirty angle statements from your brief. Do not judge them on elegance. Group near-duplicates and delete anything untrue or off brand.

2. Shortlist three to five

Choose concepts that differ from each other. Include at least one that you are fairly sure will work and at least one that feels slightly risky. This gives you a baseline and a chance of learning something.

3. Produce each concept properly

Give each shortlisted concept the same production effort. If one concept has a polished video and another a rushed image, you are testing production quality as much as the idea. Use AI for first drafts, then edit each to the same standard.

4. Name everything

Use consistent names such as concept, angle, format and date. Clear naming makes results readable later and prevents the common situation where nobody remembers what "Ad 7 final" was.

5. Set the rule before you launch

Write down the budget per concept, the length of the window and the decision rule. For instance: each concept gets a fixed daily budget for a set number of days, and a concept moves forward only if its cost per purchase is below your break-even threshold and it has at least a minimum number of purchases. Deciding afterwards invites you to rationalise whichever result you hoped for.

6. Read the result against profit

Compare each concept's return with your break-even ROAS. The break-even ROAS guide explains how to find it. Click-through rate and cost per click are useful diagnostics, but a concept with cheap clicks that do not buy is not a winner.

Sizing the test budget

A rough method: work out your target cost per purchase, then multiply by the number of purchases you want to see per concept before deciding. If your break-even cost per purchase is 15 pounds and you want about ten purchases per concept, that is around 150 pounds per concept before you can say much. Four concepts need about 600 pounds. If that is too much, test fewer concepts. The creative testing budget calculator and ad budget planner do this arithmetic, and the allocate ad testing budget guide discusses how to split spend.

Ten purchases is not a statistical guarantee. It is a pragmatic threshold for a small store, and you should be modest about what it proves.

Reading differences honestly

Random variation is larger than most founders expect. Suppose concept A shows seven purchases and concept B shows five from the same spend. That gap could easily arise by chance. The A/B test significance calculator gives a rough sense of whether a difference is likely to be real, though it makes assumptions that ad platforms often violate, such as independent random assignment. Use it as a caution against over-reading small gaps, not as a verdict.

The A/B testing ads guide goes deeper on the structure of a fair comparison. In short: keep audiences and budgets comparable, avoid editing live tests, and give each variant time.

A hypothetical batch

Imagine a store selling merino running socks. The founder generates twenty-four angle statements with AI and groups them into five themes: blister prevention, odour control over long runs, packing light for travel races, comfort when cold, and gifting for runners. Odour control and blister prevention both depend on claims that need support, so the founder checks what evidence exists. The blister claim cannot be supported by any testing the brand has done, so it is rewritten as a statement about seam placement, which is a verifiable product fact.

Four concepts proceed: seam placement, odour control as described by real reviews, packing light, and gifting. Each receives the same daily budget for ten days. At the end, gifting and packing light are the only concepts with enough purchases to read, and gifting is above break-even ROAS. Seam placement has few purchases and is inconclusive, not failed. The founder moves gifting into a second round with new hooks and keeps seam placement for a retest with a stronger opening. This is an invented example, but the pattern of "winner, loser and inconclusive" is common, and treating inconclusive as its own category is important.

What to do with winners and losers

A winner earns a second round, not a victory lap. Test hooks, formats and audiences within the winning concept, and watch for creative fatigue as frequency rises. A loser may have failed because of the idea, the execution, the audience or the offer, and you often cannot tell which. Before discarding a concept, ask whether the execution gave it a fair chance. An inconclusive result might simply need a longer window or a stronger opening.

The pause or scale ad test guide helps with the next decision, and the AI ads pillar shows how testing fits the wider workflow.

Platform behaviour you cannot control

Ad platforms optimise delivery. If you place several creatives in one ad set, the system may give most spend to one early and starve the others, which can look like a result when it is partly an artefact of delivery. Some founders respond by separating concepts into different ad sets or campaigns to force more even spend, at the cost of a more fragmented budget. Neither approach is universally right. Check the platform's current documentation on how it distributes delivery and learning, and be explicit about which method you used when comparing results. Our Meta learning phase guide describes how Meta frames this.

Attribution and scope cautions

Platform-reported purchases depend on attribution settings, and some purchases are credited to ads that did not cause them. Comparing concepts inside the same account reduces this problem, since all are affected similarly, but it does not remove it. Comparing against store revenue or against a time period before the test can give a different picture, so avoid reading too much into a single dashboard.

Finally, a result describes one product, one offer and one period. A concept that wins in the week before a holiday may behave differently in a quiet month. Record the context and retest important findings.

Keep a testing log

A simple table prevents repeated mistakes. For each test record the date, concept, angle, format, budget, purchases, cost per purchase, ROAS, decision and notes about anything unusual such as a stock-out, a price change or a promotion. After a quarter you will see patterns that no single test shows, such as which proof types consistently help, and AI can summarise the log back to you as a starting point, provided you check its summary against the raw numbers.

Using AI to review results, not just make ads

AI can also help once results arrive. Export your testing log and ask a model to group concepts by angle, list patterns, and suggest follow-up hypotheses. Treat the output as a prompt for thinking. Ask it to cite the rows it relied on, and check the arithmetic yourself, since language models can miscalculate or invent patterns that are not in the data. A useful question is: which concepts had enough purchases to read, and what do those have in common? Another is: which inconclusive concepts might deserve a retest and why? Never paste customer personal data into a tool without checking its terms.

Concept/control/budget/decision matrix

Illustrative worked example, not a measured customer result. Replace assumptions with checked facts.

ConceptControlBudgetDecision rule
Handle demoSame mug, audience, offer, destination and dates£8/day × 5 days£40 cap; wait chosen attribution delay
Size questionSame controls; change opening only£8/day × 5 daysCompare purchase contribution; inconclusive if too few events
Gift contentsExploratory separate concept, not randomized causal A/B£8/day × 5 daysUse to prioritize next controlled test

Common questions

How many ads should I test at once?
Fewer than you can generate. For a small budget, three to five clearly different concepts is a reasonable limit.
How long should a creative test run?
Long enough to reach your pre-set purchase threshold and to cover normal weekday variation. Fixed rules matter more than a precise number of days.
Is click-through rate a good way to pick winners?
It is a useful diagnostic, but it can mislead. Cheap clicks that do not buy are not wins, so judge on profit-linked results.
Can AI tell me which creative will win?
Some tools offer scores or predictions. Treat them as opinions based on patterns, and let real results from your own store decide.
What if no concept beats break-even?
Check the offer, price, landing page and audience before blaming the creative. Then test a new set of concepts, not more variations of the same one.
Should I turn off a concept that is inconclusive?
Not necessarily. Inconclusive means insufficient data. You can retest with a stronger execution or more budget, or deprioritise it.

Sources and further reading

From reading to making

Make the next idea count.

Bring your product and your context. Explore a direction, then decide what belongs in the final ad.

Create account