Answer first: synthetic testing belongs before approval, not in place of it
A production buyer can use a synthetic audience as a fast diagnostic layer: test whether the proposition is understood, compare messages, identify segment differences and decide which version deserves another edit. The result should inform the next review. It should not be presented as proof that real customers will buy, that a regulated claim is substantiated or that the campaign has worked.
That distinction is the value of PitPat's The Wanderers. Scary Robots produced one 40-second television commercial and six 10-second sponsorship bumpers featuring AI-generated talking dogs. The work aired on UK channels including Dave, Gold, Alibi and Drama. During production, Electric Twin's synthetic-audience platform was used to test scripts, messages and edits rather than waiting for a single research round after the creative was largely fixed.
The buyer question is therefore larger than whether the dogs look convincing. It is whether a producer can add rapid simulated feedback without confusing iteration evidence, client approval, real-human research and in-market performance. Those four states need different labels and different owners.
What the published case confirms
The public case reports a synthetic panel of 1,141 UK dog owners across six demographic groups. It covered current customers, lapsed customers, category sceptics and prospective buyers, and answered more than 150 questions about tracker attitudes, adoption barriers and brand perception. Seven pieces of creative were tested: the six short bumpers and the long television spot.
The clearest reported creative change concerned the safety message. Recall rose from 51% when finding a lost dog was implied through a GPS map graphic to 74% when the dogs said the point aloud. The testing also identified 'no subscription fee' as the strongest single message and found different priorities by segment: lapsed users and prospects leaned toward finding a lost dog, while sceptics responded more to activity and health tracking.
Those figures describe results inside the reported test. They do not establish campaign reach, attention, brand lift, sales or return on media spend. No public comparison with a matched real-person PitPat panel, no full questionnaire, no confidence interval and no post-launch outcome are included in the launch coverage. The responsible reading is specific: the test changed creative decisions before transmission.

The useful change is continuous research during production
Conventional research often arrives at a few expensive moments: concept selection, a rough cut or a finished ad. By then, production commitments make feedback harder to act on. A fast simulated panel can move research questions closer to writing and editorial, when changing a line, information order or visual cue is still cheap.
That does not mean every synthetic preference should steer the film. The director and agency still own tone, surprise, performance and the coherence of the piece. The client owns the brand and business decision. A producer should define which questions the model may advise on—clarity, recall, objections, segment differences—and which remain human judgments, such as taste, reputational risk, humour, emotional truth and whether the work is worth making.
Record each tested asset with a version ID, the exact question, audience definition, result, decision owner and action taken. If the result changes the script or cut, preserve the before-and-after versions. Otherwise 'the audience preferred it' becomes an untraceable approval story rather than evidence.
Source data is part of the production brief
A synthetic audience is only as useful as the real data, modelling choices and evaluation around it. Electric Twin says its audiences are built from customer or high-quality real-world data and evaluated against held-out survey answers before use. Its published accuracy page reports up to 96% on 1-MAE and 92% on its stricter NDAM measure, while also stating that performance depends on audience data and that some questions should go to humans.
Those are platform-level claims, not a published accuracy score for the PitPat panel. Before commissioning, ask for the provenance and recency of seed data, sample coverage, segment construction, questions used for calibration, hold-out result for this audience, known weak areas, model and configuration version, security terms, retention policy and whether client data trains any general system.
The brief should also name the population the system cannot represent. Dog owners willing to complete prior surveys may differ from first-time buyers, people with low digital confidence or households whose objections were never captured. Fast answers do not remove sampling bias; they let a team repeat it more quickly unless the boundary is explicit.
Keep real people at the high-consequence gates
Electric Twin's own accuracy guidance says regulated claims, clinical studies and legally mandated consumer testing require real respondents, and that synthetic audiences should work alongside surveys rather than replace a research programme. That is the right production boundary. The same caution applies when a result could materially alter health, safety, vulnerable-audience, pricing or substantiation decisions.
For an ordinary commercial, use human research where surprise, offence, ambiguity or lived experience matters and whenever the synthetic panel is thin or untested. Ask real participants to validate the leading route and the biggest dissent, not simply to rubber-stamp the model's winner. Then compare the two result sets and investigate disagreement.
PitPat describes its products as grounded in dog science and promises not to guess or make things up. That brand position makes the boundary especially important: the campaign may use AI-generated dogs and simulated response data, while product, safety and performance claims still need evidence appropriate to the claim.

A practical approval ladder for the next campaign
At brief stage, define the business question, target population, protected claims and the decisions synthetic testing may influence. At model-readiness stage, approve the source-data note and audience evaluation. At creative stage, test named versions against pre-agreed questions rather than fishing for a favourable answer. At review stage, have the director, agency and client record what changed and why.
Add a human-validation gate for the winning route and any important dissent. Legal, compliance and accessibility reviewers should receive the actual proposed master, not a summary of simulated sentiment. At delivery, package the brief, audience definition, test register, version history, human findings, approvals, disclosure decision and final masters. That record belongs with the production, not inside a vendor dashboard.
After launch, measure real behaviour: completed views, message recall, brand lift, qualified visits, purchases or the outcome named in the brief. Compare those signals with the synthetic forecast. The purpose is not merely to score the campaign; it is to learn where the audience model was useful, where it failed and whether it should influence the next commission.
What this case does—and does not—prove
The Wanderers confirms that synthetic audience testing can run alongside an AI-assisted television production and can produce decisions concrete enough to change spoken messaging and the edit. It also shows an emerging studio model in which creative development, generation, research and versioning sit inside one tighter loop.
It does not yet prove that the campaign outperformed conventional creative, that 1,141 simulated respondents equal 1,141 recruited dog owners or that the publicised recall lift translated into sales. Those would require disclosed real-world validation and campaign results. Keeping that uncertainty visible makes the case more useful, not less.
The production opportunity is controlled iteration. A well-designed AI production workflow can make more versions testable while keeping sources, owners, approvals and delivery states legible. The gain is not permission to outsource judgment. It is earlier evidence, cheaper corrections and a clearer route from creative hypothesis to real-world result.
Build
Need a repeatable AI production workflow?
Mike designs the tools, review loops, and publishing systems that make it usable.
Launching a business of your own? Founder Launch OS connects the brand, offer, website and visual campaign in one guided Codex or Claude Code workspace.
