[SEO Guide] The Last Post on Information-Gain in SEO you will ever need

Vada Booker

New member
Joined
Aug 22, 2026
Messages
1
Points
0
A couple days ago Ahrefs made a post about Info-gain (I made a thread about it) based on a 3 year old Google patent and their recent findings. It got me thinking and I decided to pen this post for my own personal blog. Halfway through I realised this forum would have better readership anyway and I don't believe in gatekeeping content. So here goes... hope you like it.

After this week's August spam update finished rolling out, I went back through a bunch of SERPs where the sites all look "good" individually.

Nice design. 2,000 words. Tables. FAQs. Author boxes. Entities everywhere.

Then you open the first five results and realize they are basically the same article written five different ways.

That is the problem I think a lot of people misunderstand when they talk about information gain.

First, info-gain is not "write something unique"

There is actually a Google patent family around what it calls an "information gain score".

The basic idea in the patent is pretty simple.

If a user has already consumed document A, how much additional information would they get from document B?

If document B mostly contains information the user already saw in document A, the information gain is low.

If document B adds information that was not present before, the information gain is higher.

The patent even talks about comparing semantic representations of documents rather than simply looking for matching words.

That distinction matters.

Changing


Most experts recommend updating WordPress plugins regularly.

into


You should regularly keep your WordPress plugins updated.

is not information gain.

Different sentence. Same information.

Also, before somebody turns this into another SEO myth, a Google patent does not prove Google is currently using that exact patent as a ranking factor.

What makes the concept useful is that Google's public documentation keeps pointing in the same general direction.

Google literally asks publishers whether their content

  • provides original information, reporting, research or analysis
  • provides information beyond the obvious
  • avoids simply rewriting other sources
  • provides substantial additional value and originality
  • provides substantial value compared with other pages in the search results

Google's scaled content abuse policy also specifically talks about large amounts of unoriginal content providing little or no value.

It even gives scraping search results and combining material from different pages without adding value as examples.

So I would not call this week's spam update an "information gain update".

Google never said that.

But if your SEO process is scrape SERP -> summarise competitors -> ask AI to make a better article -> publish 5,000 pages, you are sitting very close to exactly the type of content Google keeps telling people not to make.




PHASE 1 - Mapping the Baseline and Spotting Gaps

1. Map the SERP before you write


Take one query you actually want to rank for.

Not 100 keywords. One.

Pull the first 8 to 10 proper organic results.

Ignore Reddit, YouTube, shopping boxes etc if they are clearly serving another intent.

Now make a basic claim inventory. You are not looking at headings yet. You are looking at information.

For example, say the query is

best VPS for WordPress

You might find that almost everybody says -

  • NVMe is faster than SATA
  • Cloudflare helps performance
  • 2 GB RAM is enough for a small WordPress site
  • LiteSpeed is good for WordPress
  • managed VPS is easier for beginners
  • DigitalOcean, Vultr and Hetzner are cheap
Those are now the SERP's commodity layer.

You can still mention them if they are necessary.

But don't fool yourself into thinking you created a better page because you explained each one with another 300 words.

2. Find what the SERP does NOT answer

Now look for questions the existing pages leave open. This is where the page starts becoming worth publishing.

Maybe nobody tested -

  • TTFB from India, US and Europe
  • WooCommerce under 50 concurrent users
  • what actually happens when the server runs out of RAM
  • real monthly cost after backups, IPv4 and control panel fees
  • migration downtime
  • support response time
  • performance with the exact same WordPress installation
There is your gap.

Notice none of those require you to discover a new law of physics.

Information gain can simply be new useful information for this particular searcher.

3. Don't copy the competitors' outline and look for originality afterwards

This is a classic trap.

What happens is... SEO tools tell you the competitors use


H2A
H2B
H2C
H2D

So you also end up using


H2A
H2B
H2C
H2D

Then, as your personal differentiator, you will add H2 E and tell yourself you have information gain. Sory, but you donot.

You have already anchored the entire article around the existing consensus.

Instead, map the SERP to understand what it covers and then build the outline around the user's unanswered questions, not around the competitors' heading structure.

Sometimes your page should actually be shorter than the pages ranking above you.

4. Don't delete necessary information just because competitors already mentioned it

There is an opposite mistake here too. If every result explains something, that can mean the information is necessary to satisfy the query.

You don't need to remove it just to be different.

The goal is not


Say nothing anybody else has ever said.

The goal is


After reading the existing results, would this page still teach the user something useful?

Think of the page as having two layers.

Required information - The user needs this for the page to make sense.

Incremental information - This is why your page deserves to exist in addition to the others.

You normally need both.




PHASE 2 - How to Actually Create Real Information Gain

5. Stop thinking that information gain means "new facts" only


There are several ways a page can add real value.

Original data
You measured something yourself. Traffic data, prices, latency, conversion rates, survey results, benchmarks, failure rates etc.

Original testing
You actually tried the thing. Same site deployed on five hosts. Same prompt tested against six models. Same outreach email sent to 500 prospects.

First-hand experience
Things you only normally learn after doing the job. Setup problems. Support problems. Hidden charges. Things that broke. Things that looked good on paper and were useless in practice.

A better method
Everyone explains what to do. You explain exactly how to do it.

A failure case
This is massively underused. Ten articles say method X works. You tested X and found the exact situation where it does not. That is useful information.

A constraint
"Use X" is commodity advice. "Use X unless Y is true because Z happens" is much more useful.

A comparison nobody made
Not another feature table copied from pricing pages. Actually compare the thing users are deciding between.

Better synthesis
Sometimes all the facts already exist individually but nobody connected them properly. Putting A + B + C together and explaining what they mean for the user's decision can itself create value.

Primary evidence
Screenshots. Logs. Invoices. Benchmark output. Code. Photographs. Test methodology. Before and after data.

This is why first-hand experience is difficult to fake properly.

You can ask ChatGPT to say it tested 14 VPS companies but producing the 14 invoices, test environment, raw results and screenshots is another story.

6. Replace vague claims with useful specificity

Compare these two statements


Caching significantly improves website speed.

and


On our test WooCommerce install, enabling full-page caching reduced median TTFB from 684 ms to 211 ms across 500 requests.

The second version is useful not because it is longer.

It gives

  • a measurable result
  • a test condition
  • a comparison
  • something the reader can evaluate
That is the kind of density you should be chasing.

7. Show how you know

Google's own guidance for reviews recommends things like

  • evidence of your own experience
  • visuals
  • quantitative measurements
  • original research
  • explaining what separates one product from another
  • discussing advantages and disadvantages based on experience
You can apply the exact same thinking outside review content.

If you claim


We tested 200 expired domains and this footprint was the strongest predictor of spam.

show something.

Your methodology. Sample size. What you measured. An anonymised screenshot. A table. What failed.

Otherwise it is just another SEO guy saying "based on our testing".

8. Negative results are information gain too

This deserves attention because almost nobody publishes failures.

Say every guide recommends changing X.

You changed X on 20 sites and nothing happened.

That is useful.

Explain

  • what you changed
  • what you expected
  • what happened
  • how long you tested
  • what other variables were controlled
  • where you think the method might still work
A result does not have to support your original theory to be valuable.

9. Give AI models NEW INPUT if you want new output

This is how you actually use AI properly.

Google doesn't ban AI content. The problem is using AI to rewrite the existing SERP over and over.

Don't just give the model competitor pages.

Feed it things competitors don't have

  • your test results
  • customer interviews
  • support tickets
  • sales calls
  • survey responses
  • internal data
  • screenshots
  • technical logs
  • product documentation
  • your failed experiments
  • expert notes
Then use AI to organise, analyse and explain that material.

Now the model has a chance of producing something beyond a generic SERP remix.




PHASE 3 - False Signals and Traps to Avoid

10. Don't use quote-searches as an "information gain test"


I see this suggested sometimes.

Write a unique sentence. Search it inside quotation marks. Nobody else has it. Congratulations, apparently you have information gain.

No.

I can publish


Blue penguins improve WordPress rankings when Mercury enters retrograde.

and probably own the exact phrase.

It still adds nothing useful.

Lexical uniqueness and informational novelty are not the same thing. What matters is whether the underlying claim, evidence, method or conclusion gives the user something they did not already have.

11. Search your CLAIM, not your wording

If you think you found something new, search around the actual claim.

Different wording. Different entities. Different date ranges. Different sources.

You are trying to establish whether the information is genuinely uncommon in the current result set, not whether you invented a sentence nobody typed before.

And even then, "nobody else said this" isn't enough.

The information also needs to be useful and supportable.

12. Information gain does not mean inventing controversial takes

Another easy way to fool yourself.

The entire SERP says A.

You publish B.

Technically different. Still wrong.

Novel misinformation is not valuable information.

If your new conclusion contradicts the consensus, the burden of proof goes up. Show why.

13. A table is not automatically information gain

Taking the same specifications from five company websites and putting them into a table can improve usability.

That is useful presentation. But don't confuse presentation gain with information gain.

A table becomes much stronger when it contains something you actually produced like look at this for a Hosting Review Site.


ProviderAdvertised RAMActual Idle RAMTTFBSupport ReplyFull Monthly Cost
-------


Now the table contains information the pricing pages themselves do not have.

14. Screenshots are not automatically proof either

A screenshot of Ahrefs showing DR 71 doesn't prove your theory about rankings.

Evidence has to support the claim being made.

SEO case studies are full of graphs where something went up after somebody made a tweak, and the tweak automatically gets the credit.

If you cannot isolate causation, say so. Correlation with an honest explanation is still useful. Just don't turn it into false certainty.

15. Don't optimise for word count

Google has explicitly stated for years that it does not have a preferred word count.

Yet people still do


Competitors average 2,341 words so ours should be 3,000.

Why?

If the first 1,200 words answer the query and the remaining 1,800 repeat what everyone else says, the extra length isn't helping the user.

A 900-word page with original measurements can contain more useful information than a 4,000-word SERP rewrite.

Information gain is density of new useful information, not article length.




PHASE 4 - Editing, Layout and Content Auditing

16. Build a "commodity map" for your draft


This is what most content writers should actually start doing.

For every major claim in the draft, give it a simple score

0 = almost every ranking page already says this
1 = same information but explained slightly better
2 = adds useful specificity or an example missing from most results
3 = new test, evidence, data, method, experience or conclusion

If your 2,500-word article is mostly 0s and 1s, adding another 1,000 words won't fix it.

You just made a longer commodity article.

17. Test every paragraph against the SERP

Go through the draft. Ask this for each section

If somebody already read the first three good results, what do they gain from this paragraph?

If the answer is nothing, decide whether the paragraph is

  • necessary context
  • support for something original
  • or simply filler
If it is filler, cut it.

This is why those 200-word AI introductions are so useless.

Someone searches


how to migrate WordPress to Cloudflare

and gets


"In today's rapidly evolving digital landscape, website speed and security have become increasingly important..."

Nobody gained anything. You spent 40 words moving the user zero centimetres forward.

18. Put the extra value where the user can actually find it

Don't bury your one useful test 2,800 words down the page.

If the reason your article deserves the click is


We tested all 11 providers from France and measured actual TTFB

say that early.

Your title might say


11 WordPress VPS Hosts Tested From France With Real TTFB Data

instead of


Best WordPress VPS Hosting

Top 11 Providers

Same query target. Completely different reason to click.




PHASE 5 - Maintenance and Hard Realities

19. Don't assume CTR automatically proves information gain


If impressions exist and clicks are poor, the title or snippet may not be communicating the difference.

Or your ranking is bad. Or the SERP layout changed. Or your brand is weaker. Or Google is rewriting your snippet.

Don't turn one Search Console metric into a fake SEO law. Use it as a diagnostic clue.

20. Don't diagnose algo hits from broken reporting

Whenever Google rolls out large core or spam updates, check Search Console status pages for logging anomalies before panicking.

Google occasionally has data logging errors (like reporting drops in Generative AI or Search performance reports) during major updates.

Make sure a drop is an actual ranking hit before you start tearing your site apart.

21. Re-check the SERP after publishing

SERPs change. Competitors update pages. Google changes which intent it prefers.

Your supposedly original data can become commodity information six months later because everybody copied or cited it.

Revisit important pages periodically and ask

  • what does everybody now know?
  • what changed since publication?
  • what new questions are searchers asking?
  • what part of our page is still genuinely useful?
  • what new evidence can we add?
That is a much better content refresh than changing the year in your title tag.

22. If competitors copy your useful bit, go one level deeper

Say you publish actual pricing including hidden fees. Three months later, competitors copy your numbers.

The numbers themselves no longer differentiate you.

So add the methodology. Historical price changes - Failure cases - Support response tests - Regional differences - The next unanswered question.

Information gain is not a one-time checkbox.

23. Some queries simply have very little room left

This is worth accepting.

There are SERPs where five genuinely excellent pages already answer basically everything.

You have no data. No experience. No unique access. No better method. No tool. No different angle. No new evidence.

Maybe don't publish another article.

That is probably the cheapest information-gain optimisation you can make

Skip the page.




The Step-by-Step Workflow

Before writing a money page, run this process

  1. Pull the top 8 to 10 relevant organic results.
  2. Extract the major factual claims from each.
  3. Mark the claims repeated across most results (commodity layer).
  4. List the questions those pages leave unanswered (the gap).
  5. Choose at least one gap you can genuinely fill with proof.
  6. Build the outline around user needs, not competitor H2 structures.
  7. Write the required consensus information as tightly as possible.
  8. Spend most of your effort on the parts only you can contribute.
  9. Score each section from 0 to 3 for additional value.
  10. Remove filler that exists only to inflate word count.
  11. Put your strongest original value near the top of the page.
  12. Make the title communicate that specific value.
  13. Publish your raw methodology and evidence.
  14. Revisit the SERP later to see what has become commodity information.




The Easiest Test

Forget patents. Forget NLP scores.

Forget whether your SEO tool says you covered 94% of the recommended entities.

Open the top three good results. Read them.

Then open yours.

Ask one question

"What does the user know now that they did not know after reading those three?"

If you cannot point to anything concrete, you don't have an information-gain problem.

You have a reason-for-this-page-to-exist problem.

And rewriting the same SERP in nicer prose is not fixing it. It is still a paraphrase.

Which page of yours currently ranks in the top 10 but would fail that test?




Will post the reading material as a reply.
 
Top