{"slug":"phrase-based-indexing","title":"Phrase-Based Indexing — Content Optimization for Phrase Authority","tags":["seo","phrase-based-indexing","on-page","content-optimization","semantic-seo"],"agent_summary":"Applies Google's Phrase-Based Indexing patents to content optimization: identifies independent phrases, related phrases, and phrase co-occurrence requirements that build topical authority within a document. Covers phrase completeness scoring, related phrase injection, and the phrase authority model.","trigger_phrases":["phrase-based indexing","phrase based SEO","phrase optimization","semantic phrase SEO","phrase authority","related phrase injection","phrase co-occurrence","independent phrases SEO"],"runnable":true,"markdown":"\n## The Phrase-Based Indexing Patents\n\nGoogle's Phrase-Based Indexing patents (US7925655, US8090723, US8166045, multiple continuations) establish that Google indexes not just words but entire meaningful phrases. The algorithm:\n\n1. Identifies \"independent phrases\" — meaningful units that occur naturally across the web\n2. Maps \"related phrases\" — phrases that frequently co-occur with each other\n3. Uses phrase co-occurrence to assess document quality and topical authority\n4. Ranks documents higher when they contain not just the target phrase but the expected co-occurring phrases\n\nThe implication: writing naturally about a topic requires including the phrases that co-occur with the main topic in authoritative sources. Thin content that contains the keyword but omits its natural co-occurring phrases ranks poorly because it doesn't match Google's expected phrase pattern for the topic.\n\n## Independent Phrases vs Related Phrases\n\n**Independent phrases:** Meaningful phrases that stand alone. \"roof replacement cost\" is an independent phrase — it conveys a complete concept that users search for directly.\n\n**Related phrases:** Phrases that statistically co-occur with the independent phrase in authoritative documents. For \"roof replacement cost,\" related phrases include: \"asphalt shingles,\" \"labor costs,\" \"square footage,\" \"permit fees,\" \"material type,\" \"lifetime expectancy,\" \"insurance coverage.\"\n\nA page about roof replacement cost that omits multiple related phrases signals thin topical coverage. The algorithm interprets missing co-occurring phrases as evidence that the content doesn't fully cover the topic.\n\n## Identifying Required Related Phrases\n\n**Method 1: Google SERP extraction**\n\nSearch the target phrase. Read the top 3-5 ranking pages. Highlight every phrase that appears in at least 3 of the 5 pages. These are empirically confirmed related phrases for that topic.\n\n**Method 2: Google's Related Searches**\n\nScroll to the bottom of SERP for the target query. Google's \"Related searches\" section shows semantically connected phrases. These are derived from user query behavior — exactly the related phrase signal the algorithm uses.\n\n**Method 3: PAA extraction**\n\nPAA questions for the target query reveal the sub-topics Google considers related. Each PAA question should be treated as a related phrase cluster that the content should address.\n\n**Method 4: NLP analysis via Google Natural Language API**\n\nSubmit the top-ranking content to Google's NLP API (available via DataForSEO or directly). Extract entities and salience scores. High-salience entities that appear in competing content but not in your content = related phrase gaps.\n\n## Phrase Completeness Scoring\n\nA page achieves \"phrase completeness\" when it contains:\n- The target independent phrase: in H1, first paragraph, 2-3 times in body\n- Related phrases coverage: 80%+ of empirically identified related phrases present\n- Sub-topic coverage: at least one paragraph addressing each major sub-topic the related phrases represent\n\n**Phrase completeness check (per page):**\n1. List all related phrases identified via SERP extraction\n2. Scan page content for each phrase\n3. Calculate coverage ratio: phrases present / phrases required\n4. Target: 80%+ coverage for highly competitive queries, 70%+ for moderate competition\n\n## Implementation: Related Phrase Injection\n\nWhen optimizing existing content for phrase completeness:\n\n**Step 1: Gap identification**\nList related phrases found in competitor content but missing from yours.\n\n**Step 2: Natural integration (not stuffing)**\nFor each missing related phrase, identify the section where it logically belongs. Write 1-2 sentences that naturally incorporate the phrase while adding substantive information. Never force phrases into unrelated sections.\n\n**Step 3: Phrase density check**\nRelated phrases should appear 1-2 times in a standard-length article (1,500-2,500 words). The independent target phrase: 3-5 times including H1 and meta. Phrase stuffing (8+ repetitions) is detectable and penalized.\n\n## Phrase-Based Spam Detection\n\nThe patent also describes how Google identifies spam through phrase analysis: pages with high keyword density but low related phrase diversity signal keyword stuffing. Legitimate topical coverage naturally includes diverse related phrases because the topic requires them.\n\nSigns a page is phrase-deficient (spam risk):\n- Target keyword appears 10+ times in 500 words\n- Related phrase coverage below 40%\n- Same phrases repeated across multiple sections without new context\n- No statistical, factual, or process information despite claiming to cover the topic\n\n## The Phrase Authority Model for Topic Clusters\n\nIn a topic cluster, phrase authority distributes across the cluster. The pillar page carries all major independent phrases for the topic. Spoke pages carry independent phrases for sub-topics plus the pillar's primary phrase with reduced frequency.\n\n**Cluster phrase map:**\n- Pillar page: 100% independent phrase coverage + 80% related phrase coverage\n- Spoke pages: 100% coverage of spoke's target phrase + 40-60% coverage of pillar's related phrases\n\nThis structure signals to Google that the entire cluster is topically coherent — each spoke strengthens the pillar's phrase authority.\n\n## Cross-References\n\n- `koray-microsemantics` — predicate-noun phrase optimization that extends phrase-based indexing principles\n- `koray-query-semantics` — query-to-phrase mapping for heading construction\n- `entity-clouds-seo` — entity signals that co-occur with phrase clusters\n- `topic-cluster` — cluster architecture where phrase authority distributes across pillar and spokes\n\n#seo-sop #seo #phrase-based-indexing #on-page #content-optimization #semantic-seo\n","html":"<h2>The Phrase-Based Indexing Patents</h2>\n<p>Google's Phrase-Based Indexing patents (US7925655, US8090723, US8166045, multiple continuations) establish that Google indexes not just words but entire meaningful phrases. The algorithm:</p>\n<ol>\n<li>Identifies \"independent phrases\" — meaningful units that occur naturally across the web</li>\n<li>Maps \"related phrases\" — phrases that frequently co-occur with each other</li>\n<li>Uses phrase co-occurrence to assess document quality and topical authority</li>\n<li>Ranks documents higher when they contain not just the target phrase but the expected co-occurring phrases</li>\n</ol>\n<p>The implication: writing naturally about a topic requires including the phrases that co-occur with the main topic in authoritative sources. Thin content that contains the keyword but omits its natural co-occurring phrases ranks poorly because it doesn't match Google's expected phrase pattern for the topic.</p>\n<h2>Independent Phrases vs Related Phrases</h2>\n<p><strong>Independent phrases:</strong> Meaningful phrases that stand alone. \"roof replacement cost\" is an independent phrase — it conveys a complete concept that users search for directly.</p>\n<p><strong>Related phrases:</strong> Phrases that statistically co-occur with the independent phrase in authoritative documents. For \"roof replacement cost,\" related phrases include: \"asphalt shingles,\" \"labor costs,\" \"square footage,\" \"permit fees,\" \"material type,\" \"lifetime expectancy,\" \"insurance coverage.\"</p>\n<p>A page about roof replacement cost that omits multiple related phrases signals thin topical coverage. The algorithm interprets missing co-occurring phrases as evidence that the content doesn't fully cover the topic.</p>\n<h2>Identifying Required Related Phrases</h2>\n<p><strong>Method 1: Google SERP extraction</strong></p>\n<p>Search the target phrase. Read the top 3-5 ranking pages. Highlight every phrase that appears in at least 3 of the 5 pages. These are empirically confirmed related phrases for that topic.</p>\n<p><strong>Method 2: Google's Related Searches</strong></p>\n<p>Scroll to the bottom of SERP for the target query. Google's \"Related searches\" section shows semantically connected phrases. These are derived from user query behavior — exactly the related phrase signal the algorithm uses.</p>\n<p><strong>Method 3: PAA extraction</strong></p>\n<p>PAA questions for the target query reveal the sub-topics Google considers related. Each PAA question should be treated as a related phrase cluster that the content should address.</p>\n<p><strong>Method 4: NLP analysis via Google Natural Language API</strong></p>\n<p>Submit the top-ranking content to Google's NLP API (available via DataForSEO or directly). Extract entities and salience scores. High-salience entities that appear in competing content but not in your content = related phrase gaps.</p>\n<h2>Phrase Completeness Scoring</h2>\n<p>A page achieves \"phrase completeness\" when it contains:</p>\n<ul>\n<li>The target independent phrase: in H1, first paragraph, 2-3 times in body</li>\n<li>Related phrases coverage: 80%+ of empirically identified related phrases present</li>\n<li>Sub-topic coverage: at least one paragraph addressing each major sub-topic the related phrases represent</li>\n</ul>\n<p><strong>Phrase completeness check (per page):</strong></p>\n<ol>\n<li>List all related phrases identified via SERP extraction</li>\n<li>Scan page content for each phrase</li>\n<li>Calculate coverage ratio: phrases present / phrases required</li>\n<li>Target: 80%+ coverage for highly competitive queries, 70%+ for moderate competition</li>\n</ol>\n<h2>Implementation: Related Phrase Injection</h2>\n<p>When optimizing existing content for phrase completeness:</p>\n<p><strong>Step 1: Gap identification</strong>\nList related phrases found in competitor content but missing from yours.</p>\n<p><strong>Step 2: Natural integration (not stuffing)</strong>\nFor each missing related phrase, identify the section where it logically belongs. Write 1-2 sentences that naturally incorporate the phrase while adding substantive information. Never force phrases into unrelated sections.</p>\n<p><strong>Step 3: Phrase density check</strong>\nRelated phrases should appear 1-2 times in a standard-length article (1,500-2,500 words). The independent target phrase: 3-5 times including H1 and meta. Phrase stuffing (8+ repetitions) is detectable and penalized.</p>\n<h2>Phrase-Based Spam Detection</h2>\n<p>The patent also describes how Google identifies spam through phrase analysis: pages with high keyword density but low related phrase diversity signal keyword stuffing. Legitimate topical coverage naturally includes diverse related phrases because the topic requires them.</p>\n<p>Signs a page is phrase-deficient (spam risk):</p>\n<ul>\n<li>Target keyword appears 10+ times in 500 words</li>\n<li>Related phrase coverage below 40%</li>\n<li>Same phrases repeated across multiple sections without new context</li>\n<li>No statistical, factual, or process information despite claiming to cover the topic</li>\n</ul>\n<h2>The Phrase Authority Model for Topic Clusters</h2>\n<p>In a topic cluster, phrase authority distributes across the cluster. The pillar page carries all major independent phrases for the topic. Spoke pages carry independent phrases for sub-topics plus the pillar's primary phrase with reduced frequency.</p>\n<p><strong>Cluster phrase map:</strong></p>\n<ul>\n<li>Pillar page: 100% independent phrase coverage + 80% related phrase coverage</li>\n<li>Spoke pages: 100% coverage of spoke's target phrase + 40-60% coverage of pillar's related phrases</li>\n</ul>\n<p>This structure signals to Google that the entire cluster is topically coherent — each spoke strengthens the pillar's phrase authority.</p>\n<h2>Cross-References</h2>\n<ul>\n<li><code>koray-microsemantics</code> — predicate-noun phrase optimization that extends phrase-based indexing principles</li>\n<li><code>koray-query-semantics</code> — query-to-phrase mapping for heading construction</li>\n<li><code>entity-clouds-seo</code> — entity signals that co-occur with phrase clusters</li>\n<li><code>topic-cluster</code> — cluster architecture where phrase authority distributes across pillar and spokes</li>\n</ul>\n<p>#seo-sop #seo #phrase-based-indexing #on-page #content-optimization #semantic-seo</p>\n"}