Retrieval Chunk Boundaries
A retrieval system does not read your page, it reads a piece of it. A section that names the subject at the top and puts the qualifier three headings later produces a chunk that either misleads or gets discarded. The usual reaction is to shorten the page, which throws away the depth that made it worth citing, when the fix is to change where the boundaries fall rather than how much you say.
FORMAT
retrieval-chunk-repair.md
WHEN TO REACH FOR THIS
Use when a long page is plainly relevant to a question but assistants quote shorter pages elsewhere, because your answer is split across sections that never travel together.
The skill file
retrieval-chunk-repair.md
---
name: retrieval-chunk-repair
description: Use when a long page is plainly relevant to a question but assistants quote shorter pages elsewhere, because your answer is split across sections that never travel together.
---
# Retrieval Chunk Boundaries
A retrieval system does not read your page, it reads a piece of it. A section that names the subject at the top and puts the qualifier three headings later produces a chunk that either misleads or gets discarded. The usual reaction is to shorten the page, which throws away the depth that made it worth citing, when the fix is to change where the boundaries fall rather than how much you say.
## What you need first
- A page of roughly 1,500 words or more that ranks but is never quoted
- The heading outline as rendered in the HTML, not as it was planned
- The specific question each section is meant to answer, one per section
## Method
1. Cut the page at its H2s and read each piece as though it arrived with nothing else attached. Whatever stops making sense in isolation is what a retriever will hand to a model.
2. Restate the subject inside each section instead of leaning on the H1. A section opening "the tool does this" is unusable once it travels alone, so the product, concept or method has to be named again.
3. Move qualifiers, exceptions and conditions into the same section as the claim they modify. A caveat two sections below the claim is a separate document as far as retrieval is concerned, and the quote will go out without it.
4. Hold each section to one question. Two questions under a single heading means whichever chunk is retrieved carries half an answer to each.
5. Give every table, code block and figure a lead-in sentence stating what it shows, since a stripped table is a grid of numbers with no subject attached.
6. Break any section running past roughly 300 words with a subheading naming the sub-question, so the split happens where you chose rather than in the middle of an argument.
7. Check that no answer exists only inside an image. Alt text is often dropped in extraction, and a figure rendered as a picture does not travel at all.
## What this produces
A revised page in which every H2 section is independently readable, with subject, qualifiers and supporting figures inside the same section as the claim.
## Where this goes wrong
- Tuning to a specific chunk size quoted in a blog post - the boundaries differ per system and change without notice, so section-level self-sufficiency is the only durable version of this
- Splitting into so many fragments that the argument disappears and the page stops being useful to the human reader you still need
- Leaving the definitive answer under a heading like Conclusion, which is the section least likely to be retrieved for a question
- Repeating the subject so mechanically that the prose reads as machine-written, which costs more than the citation is worth
---
From the QuQi skill library - https://www.quqi.io/skills/retrieval-chunk-repair
Free to download · no account, no email
What you need first
-
A page of roughly 1,500 words or more that ranks but is never quoted
-
The heading outline as rendered in the HTML, not as it was planned
-
The specific question each section is meant to answer, one per section
Method
-
01
Cut the page at its H2s and read each piece as though it arrived with nothing else attached. Whatever stops making sense in isolation is what a retriever will hand to a model.
-
02
Restate the subject inside each section instead of leaning on the H1. A section opening "the tool does this" is unusable once it travels alone, so the product, concept or method has to be named again.
-
03
Move qualifiers, exceptions and conditions into the same section as the claim they modify. A caveat two sections below the claim is a separate document as far as retrieval is concerned, and the quote will go out without it.
-
04
Hold each section to one question. Two questions under a single heading means whichever chunk is retrieved carries half an answer to each.
-
05
Give every table, code block and figure a lead-in sentence stating what it shows, since a stripped table is a grid of numbers with no subject attached.
-
06
Break any section running past roughly 300 words with a subheading naming the sub-question, so the split happens where you chose rather than in the middle of an argument.
-
07
Check that no answer exists only inside an image. Alt text is often dropped in extraction, and a figure rendered as a picture does not travel at all.
What this produces
A revised page in which every H2 section is independently readable, with subject, qualifiers and supporting figures inside the same section as the claim.
Where this goes wrong
-
Tuning to a specific chunk size quoted in a blog post - the boundaries differ per system and change without notice, so section-level self-sufficiency is the only durable version of this
-
Splitting into so many fragments that the argument disappears and the page stops being useful to the human reader you still need
-
Leaving the definitive answer under a heading like Conclusion, which is the section least likely to be retrieved for a question
-
Repeating the subject so mechanically that the prose reads as machine-written, which costs more than the citation is worth
Use this skill in your own AI
The download is a plain markdown file with the name and trigger in its frontmatter. Where an assistant supports skills it can load itself, that frontmatter is what it reads to decide this one applies.
Claude Code
Save it as ~/.claude/skills/retrieval-chunk-repair/SKILL.md and Claude loads it on its own when what you are doing matches the trigger line. Put it in .claude/skills inside a project instead if the whole team should have it.
Claude
Upload the file in the skills section of your settings. Once it is there it applies itself in any conversation where the trigger fits, so you do not have to remember it exists.
ChatGPT
There is no skills format to install into, so paste the file contents into a Project instruction or a Custom GPT instead. It then applies to every chat in that project rather than only the one you paste it into.
Anything else
Paste the markdown into the chat before your question. It works in any assistant, it just has to be pasted again each time.
Questions about this skill
When do I re-section the page rather than shorten it?
When a long page is clearly relevant, ranks, and shorter pages elsewhere get quoted instead. Cutting it down is the usual reaction, and it throws away the depth that made the page worth citing. Retrieval hands a model a piece of your page rather than the page, so the fix is where the boundaries fall, not how much you say.
What do I need in hand before starting?
A page of roughly 1,500 words or more, the heading outline as actually rendered in the HTML rather than as planned, and the single question each section is meant to answer. Working from the planned outline is the trap, because headings drift during writing and it is the rendered H2s that decide where the page gets cut.
What do I end up with, and which part gets used?
A page where every H2 section stands alone, naming its subject again instead of leaning on the H1, with qualifiers and supporting figures moved inside the section holding the claim they modify. The relocated caveats are the part that pays. A qualifier two sections below its claim is a separate document to a retriever, and the quote goes out without it.
What ruins this most often?
Tuning to a chunk size quoted in a blog post. Boundaries differ per system and change without notice, so section-level self-sufficiency is the only durable version of this. The costly variant is over-splitting: fragment far enough and the argument disappears, so you lose the human reader you still need while chasing a citation.
More in AI search (GEO)