QuQi

Retrieval Chunk Boundaries

A retrieval system does not read your page, it reads a piece of it. A section that names the subject at the top and puts the qualifier three headings later produces a chunk that either misleads or gets discarded. The usual reaction is to shorten the page, which throws away the depth that made it worth citing, when the fix is to change where the boundaries fall rather than how much you say.

Obter o ficheiro da competência Deixe os agentes tratar disso
CATEGORIA
Pesquisa com IA (GEO)
FORMATO
retrieval-chunk-repair.md
PASSOS
7
PREÇO
Grátis — sem conta
QUANDO USAR ISTO

Use when a long page is plainly relevant to a question but assistants quote shorter pages elsewhere, because your answer is split across sections that never travel together.

O ficheiro da competência

retrieval-chunk-repair.md
---
name: retrieval-chunk-repair
description: Use when a long page is plainly relevant to a question but assistants quote shorter pages elsewhere, because your answer is split across sections that never travel together.
---

# Retrieval Chunk Boundaries

A retrieval system does not read your page, it reads a piece of it. A section that names the subject at the top and puts the qualifier three headings later produces a chunk that either misleads or gets discarded. The usual reaction is to shorten the page, which throws away the depth that made it worth citing, when the fix is to change where the boundaries fall rather than how much you say.

## O que precisa primeiro

- A page of roughly 1,500 words or more that ranks but is never quoted
- The heading outline as rendered in the HTML, not as it was planned
- The specific question each section is meant to answer, one per section

## Método

1. Cut the page at its H2s and read each piece as though it arrived with nothing else attached. Whatever stops making sense in isolation is what a retriever will hand to a model.
2. Restate the subject inside each section instead of leaning on the H1. A section opening "the tool does this" is unusable once it travels alone, so the product, concept or method has to be named again.
3. Move qualifiers, exceptions and conditions into the same section as the claim they modify. A caveat two sections below the claim is a separate document as far as retrieval is concerned, and the quote will go out without it.
4. Hold each section to one question. Two questions under a single heading means whichever chunk is retrieved carries half an answer to each.
5. Give every table, code block and figure a lead-in sentence stating what it shows, since a stripped table is a grid of numbers with no subject attached.
6. Break any section running past roughly 300 words with a subheading naming the sub-question, so the split happens where you chose rather than in the middle of an argument.
7. Check that no answer exists only inside an image. Alt text is often dropped in extraction, and a figure rendered as a picture does not travel at all.

## O que isto produz

A revised page in which every H2 section is independently readable, with subject, qualifiers and supporting figures inside the same section as the claim.

## Onde isto corre mal

- Tuning to a specific chunk size quoted in a blog post - the boundaries differ per system and change without notice, so section-level self-sufficiency is the only durable version of this
- Splitting into so many fragments that the argument disappears and the page stops being useful to the human reader you still need
- Leaving the definitive answer under a heading like Conclusion, which is the section least likely to be retrieved for a question
- Repeating the subject so mechanically that the prose reads as machine-written, which costs more than the citation is worth

---

Da biblioteca de competências da QuQi - https://www.quqi.io/pt/skills/retrieval-chunk-repair
Transferência gratuita · sem conta, sem e-mail

O que precisa primeiro

  • A page of roughly 1,500 words or more that ranks but is never quoted
  • The heading outline as rendered in the HTML, not as it was planned
  • The specific question each section is meant to answer, one per section

Método

  1. 01 Cut the page at its H2s and read each piece as though it arrived with nothing else attached. Whatever stops making sense in isolation is what a retriever will hand to a model.
  2. 02 Restate the subject inside each section instead of leaning on the H1. A section opening "the tool does this" is unusable once it travels alone, so the product, concept or method has to be named again.
  3. 03 Move qualifiers, exceptions and conditions into the same section as the claim they modify. A caveat two sections below the claim is a separate document as far as retrieval is concerned, and the quote will go out without it.
  4. 04 Hold each section to one question. Two questions under a single heading means whichever chunk is retrieved carries half an answer to each.
  5. 05 Give every table, code block and figure a lead-in sentence stating what it shows, since a stripped table is a grid of numbers with no subject attached.
  6. 06 Break any section running past roughly 300 words with a subheading naming the sub-question, so the split happens where you chose rather than in the middle of an argument.
  7. 07 Check that no answer exists only inside an image. Alt text is often dropped in extraction, and a figure rendered as a picture does not travel at all.

O que isto produz

A revised page in which every H2 section is independently readable, with subject, qualifiers and supporting figures inside the same section as the claim.

Onde isto corre mal

  • Tuning to a specific chunk size quoted in a blog post - the boundaries differ per system and change without notice, so section-level self-sufficiency is the only durable version of this
  • Splitting into so many fragments that the argument disappears and the page stops being useful to the human reader you still need
  • Leaving the definitive answer under a heading like Conclusion, which is the section least likely to be retrieved for a question
  • Repeating the subject so mechanically that the prose reads as machine-written, which costs more than the citation is worth

Use esta skill na sua própria IA

O ficheiro é markdown simples, com o nome e o gatilho no frontmatter. Quando um assistente consegue carregar skills sozinho, é esse frontmatter que lê para decidir que esta se aplica.

Claude Code Guarde-a como ~/.claude/skills/retrieval-chunk-repair/SKILL.md e o Claude carrega-a sozinho quando o que está a fazer corresponde ao gatilho. Coloque-a em .claude/skills dentro de um projeto se toda a equipa a deve ter.
Claude Carregue o ficheiro na secção de skills das suas definições. A partir daí aplica-se sozinho em qualquer conversa onde o gatilho encaixe, sem ter de se lembrar dele.
ChatGPT Não existe um formato de skills onde a instalar, por isso cole o conteúdo do ficheiro nas instruções de um Projeto ou de um GPT personalizado. Passa então a aplicar-se a todas as conversas desse projeto e não só àquela onde o colou.
Qualquer outro Cole o markdown na conversa antes da sua pergunta. Funciona em qualquer assistente, só tem de ser colado de novo de cada vez.

Mais em Pesquisa com IA (GEO)