---
name: machine-translation-quality-gate
description: Use when localised pages are being produced by machine translation at volume and you have to decide which of them need a human before they are published.
---

# Machine Translation Quality Gate

Machine translation is publishable in many cases and damaging in a few, and the failures cluster predictably around product names, units, negations, legal claims and calls to action. Reviewing a sample and approving the batch fails because the sample comes from the middle of the distribution while the damage sits in the tail. The guidance here has moved: untouched machine translation was once named as spam outright, and current guidance judges the page on whether it helps the reader rather than on how it was produced, so the useful question is which pages need a human, not whether the tool is permitted.

## What you need first

- The page inventory to be translated, with sessions or revenue per page in the source language
- Machine output for at least thirty pages spanning the range, deliberately including the shortest and the most technical
- A locked glossary of terms that must never be translated: product names, feature names, plan names, legal entity names
- One reviewer per target language with the authority to reject a batch and hold a launch

## Method

1. Tier the inventory by consequence before you look at quality at all. Anything carrying a price, a legal claim, a safety instruction or a conversion action goes to human translation whatever the output looks like; deep informational pages can ship post-edited.
2. Load the glossary into the translation step rather than correcting it afterwards. Product and plan names turned into ordinary nouns break every mention across the site simultaneously, and that is the usual reason a market gets re-translated from scratch.
3. Sample from the tails, not the middle: the shortest pages, the longest, the most technical, and anything containing tables, units or interface strings. That is where the tool fails and where a random sample almost never lands.
4. Check negation, modality and numbers explicitly on every sampled page. A dropped negative, a may turned into a must, or a decimal comma read as a thousands separator produces fluent text that states something false, and a reviewer reading for fluency goes straight past it.
5. Read the translated title, first heading and opening paragraph against the local result page for the target term. The tool translates the source phrasing, not the phrasing the market searches with, so a page can be entirely accurate and still target nothing.
6. Publish one tier and hold the rest until it has been measured. Impressions and engagement per market for the published set tell you whether the tier boundary sits in the right place far better than any quality score does.
7. Record which pages shipped as raw output and which were post-edited. Six months on nobody remembers, and the review that follows a complaint otherwise has to start from nothing.

## What this produces

A tiered translation plan stating which pages get raw output, post-editing or full rewrite, with a per-language sign-off record naming what shipped untouched.

## Where this goes wrong

- Reviewing for fluency alone, which passes text that reads well and says the wrong thing
- Launching every locale at once, which removes any way of telling whether the quality tier or the market was the problem
- Letting the translation step run over interface strings, brand names and plan names because they were never put in the glossary
- Machine-translating user-generated content such as reviews, which manufactures a large volume of low-value pages and is the pattern most likely to be judged as scaled abuse

---

From the QuQi skill library - https://www.quqi.io/skills/machine-translation-quality-gate
