QuQi

Retrieval Chunk Boundaries

A retrieval system does not read your page, it reads a piece of it. A section that names the subject at the top and puts the qualifier three headings later produces a chunk that either misleads or gets discarded. The usual reaction is to shorten the page, which throws away the depth that made it worth citing, when the fix is to change where the boundaries fall rather than how much you say.

Consigue el archivo de la habilidad Deja que lo lleven los agentes
CATEGORÍA
Búsqueda con IA (GEO)
FORMATO
retrieval-chunk-repair.md
PASOS
7
PRECIO
Gratis, sin cuenta
CUÁNDO USAR ESTO

Use when a long page is plainly relevant to a question but assistants quote shorter pages elsewhere, because your answer is split across sections that never travel together.

El archivo de la habilidad

retrieval-chunk-repair.md
---
name: retrieval-chunk-repair
description: Use when a long page is plainly relevant to a question but assistants quote shorter pages elsewhere, because your answer is split across sections that never travel together.
---

# Retrieval Chunk Boundaries

A retrieval system does not read your page, it reads a piece of it. A section that names the subject at the top and puts the qualifier three headings later produces a chunk that either misleads or gets discarded. The usual reaction is to shorten the page, which throws away the depth that made it worth citing, when the fix is to change where the boundaries fall rather than how much you say.

## Qué necesitas antes

- A page of roughly 1,500 words or more that ranks but is never quoted
- The heading outline as rendered in the HTML, not as it was planned
- The specific question each section is meant to answer, one per section

## Método

1. Cut the page at its H2s and read each piece as though it arrived with nothing else attached. Whatever stops making sense in isolation is what a retriever will hand to a model.
2. Restate the subject inside each section instead of leaning on the H1. A section opening "the tool does this" is unusable once it travels alone, so the product, concept or method has to be named again.
3. Move qualifiers, exceptions and conditions into the same section as the claim they modify. A caveat two sections below the claim is a separate document as far as retrieval is concerned, and the quote will go out without it.
4. Hold each section to one question. Two questions under a single heading means whichever chunk is retrieved carries half an answer to each.
5. Give every table, code block and figure a lead-in sentence stating what it shows, since a stripped table is a grid of numbers with no subject attached.
6. Break any section running past roughly 300 words with a subheading naming the sub-question, so the split happens where you chose rather than in the middle of an argument.
7. Check that no answer exists only inside an image. Alt text is often dropped in extraction, and a figure rendered as a picture does not travel at all.

## Qué produce esto

A revised page in which every H2 section is independently readable, with subject, qualifiers and supporting figures inside the same section as the claim.

## Dónde falla esto

- Tuning to a specific chunk size quoted in a blog post - the boundaries differ per system and change without notice, so section-level self-sufficiency is the only durable version of this
- Splitting into so many fragments that the argument disappears and the page stops being useful to the human reader you still need
- Leaving the definitive answer under a heading like Conclusion, which is the section least likely to be retrieved for a question
- Repeating the subject so mechanically that the prose reads as machine-written, which costs more than the citation is worth

---

De la biblioteca de habilidades de QuQi - https://www.quqi.io/es/skills/retrieval-chunk-repair
Descarga gratis · sin cuenta, sin correo

Qué necesitas antes

  • A page of roughly 1,500 words or more that ranks but is never quoted
  • The heading outline as rendered in the HTML, not as it was planned
  • The specific question each section is meant to answer, one per section

Método

  1. 01 Cut the page at its H2s and read each piece as though it arrived with nothing else attached. Whatever stops making sense in isolation is what a retriever will hand to a model.
  2. 02 Restate the subject inside each section instead of leaning on the H1. A section opening "the tool does this" is unusable once it travels alone, so the product, concept or method has to be named again.
  3. 03 Move qualifiers, exceptions and conditions into the same section as the claim they modify. A caveat two sections below the claim is a separate document as far as retrieval is concerned, and the quote will go out without it.
  4. 04 Hold each section to one question. Two questions under a single heading means whichever chunk is retrieved carries half an answer to each.
  5. 05 Give every table, code block and figure a lead-in sentence stating what it shows, since a stripped table is a grid of numbers with no subject attached.
  6. 06 Break any section running past roughly 300 words with a subheading naming the sub-question, so the split happens where you chose rather than in the middle of an argument.
  7. 07 Check that no answer exists only inside an image. Alt text is often dropped in extraction, and a figure rendered as a picture does not travel at all.

Qué produce esto

A revised page in which every H2 section is independently readable, with subject, qualifiers and supporting figures inside the same section as the claim.

Dónde falla esto

  • Tuning to a specific chunk size quoted in a blog post - the boundaries differ per system and change without notice, so section-level self-sufficiency is the only durable version of this
  • Splitting into so many fragments that the argument disappears and the page stops being useful to the human reader you still need
  • Leaving the definitive answer under a heading like Conclusion, which is the section least likely to be retrieved for a question
  • Repeating the subject so mechanically that the prose reads as machine-written, which costs more than the citation is worth

Usa esta skill en tu propia IA

El archivo es markdown simple, con el nombre y el disparador en su frontmatter. Cuando un asistente sabe cargar skills por su cuenta, es ese frontmatter lo que lee para decidir que esta le aplica.

Claude Code Guárdala como ~/.claude/skills/retrieval-chunk-repair/SKILL.md y Claude la carga solo cuando lo que haces coincide con el disparador. Ponla en .claude/skills dentro de un proyecto si la debe tener todo el equipo.
Claude Sube el archivo en la sección de skills de tus ajustes. Una vez ahí se aplica solo en cualquier conversación donde encaje el disparador, sin que tengas que acordarte.
ChatGPT No hay un formato de skills donde instalarla, así que pega el contenido del archivo en las instrucciones de un Proyecto o de un GPT personalizado. Así se aplica a todos los chats de ese proyecto y no solo a aquel donde lo pegaste.
Cualquier otro Pega el markdown en el chat antes de tu pregunta. Funciona en cualquier asistente, solo hay que volver a pegarlo cada vez.

Más en Búsqueda con IA (GEO)