Case Study

How a global home appliance major tested two TVCs with 60 AI-moderated interviews

Two finished films, 30 in-depth interviews each, four metros, five fixed evaluation dimensions, and a spoken explanation behind every rating. Both studies ran in parallel and closed in under a week. This is how the study was built, and what that design lets you see that a standard ad test does not.

By Raunak Kochar, Founder, InquiSight · Published

Case study TVC testing AI-moderated research

The ask

A global home appliance major had two finished TVCs for two products in the same portfolio, one higher-ticket and one lower-ticket. Both were ready to air. The brand wanted a consumer read before the media money was committed.

The study

60 AI-moderated in-depth interviews, 30 per film, across Delhi, Mumbai, Bangalore, and Hyderabad. 15 to 30 minutes each, five fixed dimensions, spoken follow-ups behind every rating. Both films fielded in parallel.

The result

A diagnostic on each film, plus a comparison across the two that neither individual report contained. Brief to topline in under a week, in time to act on the edits.

Ad testing in India usually arrives too late to change anything. The film is shot, the media is booked, and the research becomes a scorecard rather than a decision tool. This brand wanted the opposite: a real consumer read on two finished films while there was still time to re-cut them. That meant the study had to be genuinely qualitative, run in four cities, cover two films, and close in days. To respect client confidentiality we have anonymised the brand, the products, and the findings. What follows is the method, which is the part worth copying.

The question behind the brief

The brief was the one every marketer writes: "Does the ad work?" That is not a researchable question. It is five, and they have to be asked in a fixed order, because the answer to each one contaminates the next:

  • Initial impressions. What comes back spontaneously after a single exposure, before anyone is prompted with anything?
  • Comprehension. What message does the viewer actually take away, in their own words?
  • Creative evaluation. Does the story and the casting earn attention, and does the product demonstration feel achievable at home?
  • Brand image. What kind of brand does the film imply, and who does the viewer assume it is for?
  • Purchase intent. Did the film move anyone, and what will they do next?

Ask about the tagline before you ask what they remember, and you will never learn whether the tagline was remembered. Ask about intent without knowing where the viewer started, and you cannot tell persuasion from selection. The sequence is the instrument.

How we set up the study

1. Real buyers, screened on behaviour not claims

We recruited SEC A1 to A2 adults aged 25 to 50, split across two age bands (25 to 39 and 40 to 49), who were the appliance decision-maker in their household and already owned an appliance worth 15,000 rupees or more. Four metros: Delhi, Mumbai, Bangalore, and Hyderabad. Ownership was the important screen. Someone who has actually spent that much on an appliance evaluates an appliance ad differently from someone who says they would.

2. In-depth interviews, not a survey with a video attached

Each respondent went through a 15 to 30 minute conversation, not a click-through. Ad testing usually degrades into a survey because depth is expensive to field. That constraint is what AI moderation removes, and it is the reason this study could be 60 real interviews instead of 300 shallow responses.

3. The same five dimensions, in the same order, for both films

Both films were evaluated on identical dimensions with identically worded probes. This sounds like housekeeping. It is the single decision that made the cross-film comparison possible later, and that comparison turned out to be the most useful part of the deliverable.

4. Spontaneous before prompted, without exception

The moderator asked for unaided playback first: what do you remember, what was it saying, what stayed with you. Only after that did it introduce anything specific. Every experienced human moderator knows to do this. Across 60 interviews in four cities, the order never slipped once, because it could not.

5. Purchase intent captured twice, on the same person

We recorded an intent band at screening, before the respondent saw anything, and again after exposure. Without that baseline, a strong post-exposure number tells you who you recruited, not what the film did.

6. A spoken "why" behind every rating

After each rating, the AI asked the respondent to explain it, out loud, in whichever language they were comfortable in, and probed the answer. Every number in the report carried real consumer language behind it, tagged to the exact question that produced it.

7. Both films in the field at the same time

Two independent samples, two instruments, one fielding window. Sequential fielding would have doubled the timeline and introduced a gap in which nothing else about the market stayed still.

What this design makes visible

The specific findings belong to the client. The categories of finding do not, and they are the reason to build a study this way rather than the fast way.

Whether a line is doing any work, or the benefit is doing it for you

Unaided recall is the only honest test of a tagline. Prompted recall flatters every line ever written, because recognition is not memory. What spontaneous playback reveals is which words the viewer reaches for on their own. Often those words are the product benefit, phrased in the viewer's own idiom, rather than the line that was written and scored and signed off. A memory slot gets filled either way. If the line does not claim it, the benefit will, and you would rather know that before the media buy than after.

The difference between an ad that converts and an ad that looks like it converted

Measuring intent on both sides of the exposure splits the audience into groups that behave very differently. Viewers who arrive on the fence and move up are the persuasion story. Viewers who arrive already convinced and move down are the story nobody reports, because a single post-exposure average absorbs them completely. A film can lift the middle and erode the top in the same 30 seconds, and only a before and after design on the same respondent will show it.

Which category the viewer files the product into

This one only comes out of open conversation. A viewer does not evaluate an ad in a vacuum; they slot the product into a category they already have, then judge it against the rules of that category and ask the questions that category invites. When the questions coming back are about a use case the film never showed, the problem is rarely the execution. It is that the ad landed in a mental category the brand did not write for. No rating scale reaches that. A spoken follow-up gets there in one probe.

What a second film tells you about the first

Running two films on identical dimensions produced observations that neither individual report contained, because an absence only reads as an absence when something comparable does not have it. Two films from the same brand, the same team, and the same creative approach will differ in small structural ways that are invisible in isolation and obvious side by side. If you have two pieces of creative in flight, test them together. The comparison is close to free and it is where the sharpest recommendations came from.

Why this needed AI moderation

What the study neededTraditional routeHow it ran on InquiSight
60 in-depth interviews across 4 metros, on 2 filmsTwo studies, field teams in each city, several weeksBoth films recruited and fielded in parallel in one window
Identical probe order in every single interviewDepends on moderator discipline and varies across moderators and citiesSame sequence, same wording, all 60 interviews
Spontaneous recall captured before any promptingEasy to lose in a long day of fieldworkStructurally enforced, not left to judgement
Intent measured at screening and again after exposureExtra design and analysis cost, often droppedBuilt into the screener and joined to the respondent
A spoken explanation behind every rating, in the viewer's languageRealistic only in 8 to 12 depth interviewsAll 60 respondents, tagged to the exact question
An answer before the media commitment4 to 6 weeks per studyBrief to topline in under a week, for both films

The honest summary: this is a study with the discipline of a quantitative instrument and the depth of qualitative interviewing. That combination is expensive and slow to do the traditional way, which is why most ad testing quietly gives up one or the other. AI moderation is what makes keeping both routine.

A note on honesty in reporting: at 30 interviews per film, city-level cells were 6 to 9 respondents. Every city and cohort cut was labelled directional wherever it appeared, and every quote in the report was matched to that respondent's own rating on the question it illustrated. Speed should never come at the cost of rigour.

What you can copy from this study

  • Fix the question order before you write a single probe. Spontaneous first, prompted second, intent last. Sequence is the instrument, not the questionnaire.
  • Baseline intent at screening. A post-exposure number on its own describes your sample, not your film.
  • Test unaided, then decide about the line. If viewers hand you the benefit instead of the tagline, the line has not lost yet, but it is not on screen long enough to win.
  • Ask which category they put it in, not just what they thought. The questions a viewer asks reveal the shelf they mentally placed you on.
  • Test two pieces of creative together whenever you have them. Side by side on identical dimensions, gaps surface that a single-film report will never contain.
  • Run it early enough to matter. Research that lands after the media booking is a report card. The same study a week earlier is a decision.

Where InquiSight can help

If you have a film, a cutdown, or a set of creative routes in flight, we can run this exact study on your consumers, in your markets and languages, and get you a diagnostic while the edit is still open. You can share your brief here or book a demo. For a study built the same way on a different question, read our celebrity casting case study, and for the method itself, our guide to AI-moderated interviews vs traditional research agencies.

Have a film going on air soon?

Test it with real buyers in your key markets before the media commitment. Spontaneous recall, intent measured on both sides of the exposure, and a spoken explanation behind every rating, in days rather than weeks.