A/B test on manual versus AI generated ad creative for seller acquisition
Say an operator running seller acquisition ads wants to know whether AI generated creative is actually as good as agencies are pitching it. A clean way to test it: same offer, same landing page, same audiences, Facebook and Instagram, media split evenly over 90 days. Group A is a handful of manual creatives shot and written by hand. Group B is a much larger batch of AI generated variants, images and copy, produced in a fraction of the time. A plausible result set looks like this. Group A: $2,100 spend, 71 form fills, $29.58 per fill, 9 conversations that reached a price discussion, 1 contract. Group B: $2,100 spend, 94 form fills, $22.34 per fill, 7 conversations to price, 0 contracts. So B wins on cost per lead by roughly 24 percent and loses on everything downstream. With a sample that small the contract gap alone isn't conclusive. What does deserve attention is conversation quality. Leads from AI generated creative sometimes arrive with a different idea of what they signed up for, particularly when the imagery shows a property that does not exist. The open question in a case like this is whether the downstream gap is a creative honesty problem, fixable by using AI only for copy and layout on real photos, or whether cheap volume simply drags in weaker intent regardless of how the creative was produced. Attribution across a long lead-to-contract window also complicates any read, since some of Group A's contract may trace back to earlier spend outside the test window. The decision that follows is usually between running a third group with real photos and AI copy, or abandoning cost per lead as the optimization target altogether in favor of cost per conversation, which requires rebuilding tracking most operators have been putting off.