QuQi

Machine Translation Quality Gate

Machine translation is publishable in many cases and damaging in a few, and the failures cluster predictably around product names, units, negations, legal claims and calls to action. Reviewing a sample and approving the batch fails because the sample comes from the middle of the distribution while the damage sits in the tail. The guidance here has moved: untouched machine translation was once named as spam outright, and current guidance judges the page on whether it helps the reader rather than on how it was produced, so the useful question is which pages need a human, not whether the tool is permitted.

Get the skill file Let the agents run it
CATEGORY
International SEO
FORMAT
machine-translation-quality-gate.md
STEPS
7
PRICE
Free - no account
WHEN TO REACH FOR THIS

Use when localised pages are being produced by machine translation at volume and you have to decide which of them need a human before they are published.

The skill file

machine-translation-quality-gate.md
---
name: machine-translation-quality-gate
description: Use when localised pages are being produced by machine translation at volume and you have to decide which of them need a human before they are published.
---

# Machine Translation Quality Gate

Machine translation is publishable in many cases and damaging in a few, and the failures cluster predictably around product names, units, negations, legal claims and calls to action. Reviewing a sample and approving the batch fails because the sample comes from the middle of the distribution while the damage sits in the tail. The guidance here has moved: untouched machine translation was once named as spam outright, and current guidance judges the page on whether it helps the reader rather than on how it was produced, so the useful question is which pages need a human, not whether the tool is permitted.

## What you need first

- The page inventory to be translated, with sessions or revenue per page in the source language
- Machine output for at least thirty pages spanning the range, deliberately including the shortest and the most technical
- A locked glossary of terms that must never be translated: product names, feature names, plan names, legal entity names
- One reviewer per target language with the authority to reject a batch and hold a launch

## Method

1. Tier the inventory by consequence before you look at quality at all. Anything carrying a price, a legal claim, a safety instruction or a conversion action goes to human translation whatever the output looks like; deep informational pages can ship post-edited.
2. Load the glossary into the translation step rather than correcting it afterwards. Product and plan names turned into ordinary nouns break every mention across the site simultaneously, and that is the usual reason a market gets re-translated from scratch.
3. Sample from the tails, not the middle: the shortest pages, the longest, the most technical, and anything containing tables, units or interface strings. That is where the tool fails and where a random sample almost never lands.
4. Check negation, modality and numbers explicitly on every sampled page. A dropped negative, a may turned into a must, or a decimal comma read as a thousands separator produces fluent text that states something false, and a reviewer reading for fluency goes straight past it.
5. Read the translated title, first heading and opening paragraph against the local result page for the target term. The tool translates the source phrasing, not the phrasing the market searches with, so a page can be entirely accurate and still target nothing.
6. Publish one tier and hold the rest until it has been measured. Impressions and engagement per market for the published set tell you whether the tier boundary sits in the right place far better than any quality score does.
7. Record which pages shipped as raw output and which were post-edited. Six months on nobody remembers, and the review that follows a complaint otherwise has to start from nothing.

## What this produces

A tiered translation plan stating which pages get raw output, post-editing or full rewrite, with a per-language sign-off record naming what shipped untouched.

## Where this goes wrong

- Reviewing for fluency alone, which passes text that reads well and says the wrong thing
- Launching every locale at once, which removes any way of telling whether the quality tier or the market was the problem
- Letting the translation step run over interface strings, brand names and plan names because they were never put in the glossary
- Machine-translating user-generated content such as reviews, which manufactures a large volume of low-value pages and is the pattern most likely to be judged as scaled abuse

---

From the QuQi skill library - https://www.quqi.io/skills/machine-translation-quality-gate
Free to download · no account, no email

What you need first

  • The page inventory to be translated, with sessions or revenue per page in the source language
  • Machine output for at least thirty pages spanning the range, deliberately including the shortest and the most technical
  • A locked glossary of terms that must never be translated: product names, feature names, plan names, legal entity names
  • One reviewer per target language with the authority to reject a batch and hold a launch

Method

  1. 01 Tier the inventory by consequence before you look at quality at all. Anything carrying a price, a legal claim, a safety instruction or a conversion action goes to human translation whatever the output looks like; deep informational pages can ship post-edited.
  2. 02 Load the glossary into the translation step rather than correcting it afterwards. Product and plan names turned into ordinary nouns break every mention across the site simultaneously, and that is the usual reason a market gets re-translated from scratch.
  3. 03 Sample from the tails, not the middle: the shortest pages, the longest, the most technical, and anything containing tables, units or interface strings. That is where the tool fails and where a random sample almost never lands.
  4. 04 Check negation, modality and numbers explicitly on every sampled page. A dropped negative, a may turned into a must, or a decimal comma read as a thousands separator produces fluent text that states something false, and a reviewer reading for fluency goes straight past it.
  5. 05 Read the translated title, first heading and opening paragraph against the local result page for the target term. The tool translates the source phrasing, not the phrasing the market searches with, so a page can be entirely accurate and still target nothing.
  6. 06 Publish one tier and hold the rest until it has been measured. Impressions and engagement per market for the published set tell you whether the tier boundary sits in the right place far better than any quality score does.
  7. 07 Record which pages shipped as raw output and which were post-edited. Six months on nobody remembers, and the review that follows a complaint otherwise has to start from nothing.

What this produces

A tiered translation plan stating which pages get raw output, post-editing or full rewrite, with a per-language sign-off record naming what shipped untouched.

Where this goes wrong

  • Reviewing for fluency alone, which passes text that reads well and says the wrong thing
  • Launching every locale at once, which removes any way of telling whether the quality tier or the market was the problem
  • Letting the translation step run over interface strings, brand names and plan names because they were never put in the glossary
  • Machine-translating user-generated content such as reviews, which manufactures a large volume of low-value pages and is the pattern most likely to be judged as scaled abuse

Use this skill in your own AI

The download is a plain markdown file with the name and trigger in its frontmatter. Where an assistant supports skills it can load itself, that frontmatter is what it reads to decide this one applies.

Claude Code Save it as ~/.claude/skills/machine-translation-quality-gate/SKILL.md and Claude loads it on its own when what you are doing matches the trigger line. Put it in .claude/skills inside a project instead if the whole team should have it.
Claude Upload the file in the skills section of your settings. Once it is there it applies itself in any conversation where the trigger fits, so you do not have to remember it exists.
ChatGPT There is no skills format to install into, so paste the file contents into a Project instruction or a Custom GPT instead. It then applies to every chat in that project rather than only the one you paste it into.
Anything else Paste the markdown into the chat before your question. It works in any assistant, it just has to be pasted again each time.

Questions about this skill

When is a tiered gate better than reviewing a sample and approving the batch?

Use it when machine output is going out at volume and somebody has to decide which pages need a human first. Sampling and approving fails because the sample lands in the middle of the distribution while the damage sits in the tail. Current guidance judges a page on whether it helps the reader, not on how it was produced.

What do I need in hand before starting?

A page inventory with sessions or revenue per page, machine output for at least thirty pages spanning the range, a locked glossary of product, feature, plan and legal entity names, and one reviewer per language with authority to hold a launch. Without the glossary loaded into the translation step, plan names become ordinary nouns across every page at once.

What do I end up with, and which part of it gets used?

A tiered plan routing each page to raw output, post-editing or full rewrite, with a per-language record of what shipped untouched. The tier boundary is what gets used daily: anything carrying a price, a legal claim, a safety instruction or a conversion action goes to a human whatever the machine output happens to read like.

What ruins this most often?

Reviewing for fluency. Fluent text stating the opposite of the source passes every read: a dropped negative, a may turned into a must, a decimal comma read as a thousands separator. Check negation, modality and numbers explicitly. Machine-translating user reviews is the separate and worse case, manufacturing volume that is most likely to be judged as scaled abuse.

More in International SEO