Translation cost in technical docs
Most technical-documentation teams underestimate their translation bill by a factor of three to five. The reason is structural, not the per-word rate.
Technical-documentation teams have a translation budget. Most of them think they understand where it comes from.
The conventional model is:
- Words translated × per-word rate = translation cost.
This model is correct as a unit-economics statement. It is misleading as a model for where to spend management attention. The two variables that actually drive the budget are upstream of this equation, not inside it. They are also the most addressable.
The variables that drive the bill
The per-word rate is the variable that gets negotiated. Procurement teams pour effort into the rate negotiation because it is the lever they can pull. The per-word rate is also the smallest of the four levers that matter:

The economic argument in one picture: many language renderings of the same content, but only one source component to translate and approve. The translation-cost line falls in proportion to the reuse the component model enables, which is the case the post works through quantitatively.
Source-content volume. The number of words in the source language that the documentation team chooses to publish. A 200,000-word product manual costs eight times as much to translate as a 25,000-word product manual, all else equal.
Number of target languages. The number of markets the content has to land in. A 200,000-word manual into 16 languages is a 3.2 million-word translation order. The multiplier is in the language count, not in the source manual.
Refresh frequency. How often the content set is republished. A 3.2 million-word translation order published annually is one cost; the same order published quarterly is roughly four times the cost, modulo TM savings on the parts that have not changed.
Per-word rate. The lever procurement focuses on. Usually 5-15% addressable through negotiation; rarely more without sacrificing quality.
The structural cost driver is the first three multiplied together. Source × languages × refresh = the volume the team is actually paying to translate. Per-word rate is the multiplier on that volume.
Where the source-content volume comes from
The source-content volume is the most addressable of the four levers, because most of it is duplicate content being translated separately for each document.
Consider a product family with three variants. Each variant has its own user manual. The three manuals share roughly 60-70% of their content: installation, safety, basic operation, troubleshooting. Each manual is, say, 80,000 words. The team publishes 240,000 words across three manuals. About 50,000 words per manual are variant-specific; the remaining 30,000 words per manual are duplicates of the same content in the other two manuals.
Under a document-based authoring model, those 30,000 duplicate words per manual are translated three times for every target language. With sixteen target languages, that is 30,000 × 3 × 16 = 1,440,000 words of translation that the team is paying for to deliver content the source team only wrote once.
Translation memory absorbs some of this. A mature TM with high fuzzy-match thresholds will catch verbatim duplicates and reduce the per-word price on them. It will not, however, eliminate the project-management overhead, the per-segment review cost, or the orchestration of running the translation across three documents. The structural waste remains; only the per-word cost on the duplicated segments comes down.
Under a component model, the 30,000 duplicate words exist as components shared across the three manuals. They are translated once per target language, regardless of how many manuals reference them. The 16-language translation order on the shared content drops from 1,440,000 words to 480,000 words.
That is the kind of arithmetic that shows up on the invoice.
What happens when source content actually changes
The translation cost in steady state, when the source content is stable, is one story. The translation cost when source content changes is a different story, and usually the more important one.
A safety statement in the user manual gets updated. Under a document-based model, the safety statement lives in every document that contains it: three manuals, the operator handbook, the maintenance guide, the regulator submission, the training pack, the spare-parts catalog. Eight documents, sixteen languages: 128 separate translation actions to refresh the same statement.
Under a component model, the safety statement is one component. Updating it refreshes one component in sixteen languages: 16 translation actions, run once. The 112-action saving on a single statement compounds across however many statements change in a refresh cycle. For a regulated-content team with a steady churn of regulatory updates, that compounding is where the CCMS pays for itself.
Where controlled-language sits in the math
Controlled-language specifications add a different lever to the same equation. The most prominent is ASD-STE100 (Simplified Technical English), originally developed for aerospace maintenance content and now applied across heavy industry, medical devices, and increasingly anywhere translation cost is a sustained concern.
A controlled-language specification constrains authors to a limited vocabulary (typically 850-900 approved words for general use) and a prescribed set of sentence structures. The translation-cost effect is twofold.
First, controlled-language source content has higher TM match rates because the phrasing is more deterministic. Two safety statements that mean the same thing in plain English (“Do not operate this equipment while wearing loose clothing” / “Operating the equipment in loose clothing is not permitted”) have very different TM matches. In controlled language, both authors are required to write the same sentence, the TM matches reach 100% on the second occurrence rather than 65%.
Second, the controlled vocabulary is faster for human translators to handle. Translators face fewer interpretation choices per segment because the source language is more constrained. Per-segment time drops; per-segment cost drops with it.
Both effects compound across a large content base. The arithmetic is the same as the component-content one: reduce the volume of distinct source content that has to be processed by a human translator. The controlled-language version applies that arithmetic at the sentence level rather than the paragraph level.
DxChecker enforces controlled-language rule sets at authoring time, so the source content meets the controlled-language specification as a structural property rather than as a review-stage cleanup.
Where the savings stop
The arithmetic above is the optimistic case. There are several places where the savings stop or reverse.
Truly variant-specific content. Some content is genuinely different per market: regulatory disclosures, market-specific safety language, certifications. That content has to be translated separately for each target language because the source itself is per-language. Components do not help; controlled language does not help. The only lever is to keep the source as short as the regulator allows.
Languages with low TM-matchable similarity. Languages that diverge significantly from the source (English to Japanese, English to Arabic) get less benefit from TM than languages closer to the source (English to German, English to French). The TM portion of the savings narrows; the component portion still stands.
Translation-memory misalignment after a refactor. A team migrating from a document-based model to a component-based model often has to do a one-time TM realignment so the new component-level source aligns with the existing segment-level TM. This is a one-off cost. Some teams under-budget it.
Small content bases. A team publishing a single product into two markets does not have a translation-cost problem worth solving structurally. The CCMS overhead exceeds the translation savings. The structural fix matters at scale; it is a wrong tool at small scale.
What to do with this
For a documentation lead estimating where to spend the next year’s tooling budget, the question is not “what is the per-word rate I am paying”, it is “what is the volume of source content the structure of my program is producing, and where would a component model or a controlled-language specification cut that volume”.
The answer is unique to the customer. The mechanism is universal: reduce the source-content volume; the translation invoice falls with it, regardless of the per-word rate negotiation.
For the broader structural argument, see the companion piece on component content management vs document management, that one walks the same architectural choice from the content-management angle. The Technical Documentation industry page covers how this plays out in production with sector-specific examples.
Frequently asked questions
Doesn't translation memory already eliminate this problem?
Partly. Translation memory absorbs verbatim duplicates well and near-duplicates moderately well, depending on the fuzzy-match threshold the customer is willing to accept. But translation memory operates after the source content is finalized and sent for translation; it does not change the volume of source content that has to be reviewed by a human translator. Even at high fuzzy-match rates the per-segment review cost is not zero, and the project-management overhead of running translation across hundreds of documents scales linearly with document count, not segment count.
What is the difference between TM-driven cost reduction and CCMS-driven cost reduction?
TM reduces the cost of translating a given source corpus. CCMS reduces the size of the source corpus. The two stack: a CCMS team with TM gets the CCMS savings on top of the TM savings, not instead of them. The CCMS savings come from translating one approved component once instead of translating the same content separately in every document that contains it. The TM savings come from matching fuzzy segments within whatever content does still need to be translated.
How does controlled language fit in?
Controlled-language specifications, most famously ASD-STE100 (Simplified Technical English), constrain authors to a limited vocabulary and prescribed sentence structures. The translation effect is twofold. First, controlled-language source content has higher TM match rates because the phrasing is more deterministic. Second, the controlled vocabulary is faster to translate per segment because translators face fewer interpretation choices. Both effects compound across a large content base.
Where does the math actually start to bite?
When the source content volume × number of target languages × refresh frequency produces a translation invoice that is meaningful relative to the documentation team's total cost. For a small team publishing one product into two languages, the math is not interesting. For a team publishing twenty product families across sixteen markets with quarterly revisions, the translation budget is often larger than the documentation budget. The CCMS-driven savings show up at that scale.
Is there a typical reduction figure?
Customer experience varies widely with content type, language pair, current toolchain, and translation-memory maturity. A conservative range is 30-50% translation cost reduction in the first refresh cycle after a component model is in production, on top of whatever TM savings the customer already had. Some customers report higher savings on second and subsequent refresh cycles, as the proportion of unchanged source content grows. None of the published numbers should be taken as a guarantee for a specific buyer; the structural mechanism is sound but the magnitude depends on the customer's content shape.