A quick primer
Here's the thing to understand: a headline's effectiveness isn't a property of the text itself — it's the outcome of that text's interaction with the context of a specific ad group, and that context is genuinely different across ad groups even when the headline text is identical. Google's official account-organization guidance is built on exactly this principle: for each ad group, it's recommended to pick a narrow theme and use keywords relevant to that theme, including at least one of those keywords in the ad's headline — since a user is more likely to find an ad relevant when it explicitly mentions the term they searched for. If the same headline is used in two ad groups with different semantics, it will, by definition, match one group's queries more precisely than the other's, simply because the surrounding keywords and real search queries differ.
Google's own Ad Strength documentation directly ties headline relevance to the specific ad group's keywords: it recommends including keyword text in headlines and descriptions to improve relevance to that specific set of keywords — meaning relevance is scored relative to a given group, not in the abstract. A separate factor is how Google determines which ad group should even answer a given query in the first place: relevance is determined by the meaning of the search query, all of the ad group's keywords, and the landing pages within that group — so even with an identical headline, different ad groups can end up receiving structurally different queries, a different competitive auction environment, and different landing pages the click leads to. Finally, the RSA mechanism itself assembles a headline into combinations with the other headlines and descriptions in that specific group's pool — and since the pool differs between groups, the same headline ends up as part of different pairings, which also affects the end result.
What to check before you touch anything
- How closely the semantics (keywords) of the two groups where the same headline is being compared actually overlap — a narrow overlap already partly explains a performance gap.
- Whether the ads in those groups lead to the same landing page or different ones — a different page means a different post-click experience regardless of the ad's own text.
- Whether the actual search queries triggering the ad in each group are comparable — via the search terms report, not just the formally set keywords.
- Whether the rest of the headline/description pool in each group's ad is comparable — if it differs substantially, the shared headline ends up in different combinational contexts.
- Whether traffic volume and accumulated data are comparable across the two groups — on very different traffic levels, the comparison may be skewed by statistical noise.
Possible approaches
- If the difference is explained by different group semantics — don't carry conclusions about a headline's effectiveness mechanically from one group to another; test and evaluate it separately in each relevant context.
- If a headline consistently underperforms specifically in a broad or mixed-semantics group — consider narrowing that group's composition rather than replacing the headline itself, since the cause may be structural, not textual.
- If the landing pages differ between the compared groups — factor that in as a separate variable when interpreting a conversion gap, rather than attributing the whole difference to the headline.
- If the groups are comparable on semantics and landing page, and the gap is still significant — then it's worth looking at the rest of each ad's asset pool: the difference may come down to which other headlines/descriptions the shared one is most often paired with in each group.
- Use this effect deliberately: if the same headline performs well in one semantic niche and poorly in another, that's itself a useful diagnostic signal about which phrasing is closer to the language and intent of that specific cluster of queries.