← All questionsResponsive Search Ads

How do I evaluate the performance of a specific headline combination inside an RSA?

A quick primer

There's a limitation here worth naming up front, since a lot of advertisers don't realize it: the combinations report only shows impression counts for each combination of headlines and descriptions — no CTR or conversion data at the combination level. In other words, seeing "this specific pairing of three headlines and two descriptions drove X conversions" is simply not possible — that data doesn't exist anywhere in the interface in any form. The asset report gives metrics for individual headlines and descriptions, but not for their pairings, and those metrics are best used as a directional indicator only, since an asset's performance depends on what it was combined with.

This isn't an interface gap — it's a function of scale: with 15 headlines and 4 descriptions, the number of possible combinations runs into the thousands, and Google simply doesn't collect a statistically meaningful sample for every single combination separately. The system as a whole learns to favor better-performing combinations, but it doesn't disclose the exact criterion (presumably CTR) it's optimizing toward. The official, and effectively the only supported, way to get an honest read on specific copy variants is not to parse the combinations report, but to run a proper test through Ad variations (account/campaign-wide find-and-replace) or Custom experiments, which split traffic between the original and the modified version and let you compare results on a controlled sample — instead of guessing from after-the-fact impression data.

What to check before you touch anything

  • Whether you're trying to draw a conclusion about a specific combination's effectiveness directly from its impression count in the combinations report — that's simply not a valid inference, since impressions there aren't equivalent to performance, and there's no click/conversion data at the combination level.
  • Whether there's enough data at the individual-asset level (not combination level) to draw conclusions — compare headlines against an impression volume comparable to the rest of the group, not against isolated cases.
  • Whether there's an actual measurable problem behind the suspicion of an "underperforming combination" (a real drop in group-level CTR/conversions), or whether it's an assumption without numeric backing.
  • Whether an A/B test through Ad variations/Custom experiments has already run for this campaign — part of the answer may already exist in prior test history, and a new run may be unnecessary.
  • Whether the ad group has enough traffic to run a new test within a reasonable timeframe — on low-frequency groups, an experiment can drag on without reaching statistical significance.

Possible approaches

  • To test a specific hypothesis (a different call to action, a different way of phrasing the USP) — use Ad variations with find-and-replace: it gives you a controlled split of traffic and real click/conversion metrics on the modified version, not just impressions.
  • To test several distinctly different messages at once — run multiple RSAs in one ad group with partial pinning on the key differences, and compare them at the ad level (not the combination level), relying on Google's built-in comparison of ads within a group.
  • If traffic volume is limited — instead of testing one specific combination, rely on ad-group-level conversion metrics (CTR, conversions, cost per conversion), since Google recommends that level of aggregation as more reliable than analyzing individual assets or combinations.
  • Don't treat "this combination shows up more often" as proof of its effectiveness — the frequency likely reflects the algorithm's CTR preference, and doesn't automatically align with the business goal (conversions, ROAS/CPA), especially if the business isn't optimizing for clicks.
  • If a hypothesis is confirmed through an Ad variations test — pin only the winning variant, targeted, rather than leaving pins in place "just in case" on every tested variant.