Short answer: test several approved models against representative content from the real Umbraco site. Compare meaning, terminology, formatting, language coverage, latency and total usage cost. Choose through a governed Umbraco.AI profile so editors use a known model and prompt rather than entering arbitrary credentials or instructions.
Is there one best AI model for translation?
No single model is best for every language pair, content type and organisation. A model that handles English-to-German product copy well may be less convincing for Japanese interface labels, Arabic right-to-left content or specialist legal wording. Models and prices also change, so a permanent recommendation based on one benchmark quickly becomes stale.
A useful selection process is repeatable: define what good translation means for the project, test a fixed content set, record the results and review the choice when models or requirements change.
What should you compare?
| Criterion | What to test | Why it matters |
|---|---|---|
| Meaning | Ambiguous phrases, references and calls to action | Fluent output can still change the intended claim |
| Language quality | Each real source and target pair | Capability varies by language and regional variant |
| Terminology | Brand names, protected terms and required translations | Consistency matters across pages and releases |
| Structure | Rich text, placeholders and short UI strings | Output must remain usable inside the CMS |
| Latency and reliability | Single pages and a representative batch | Bulk workflows amplify slow responses and rate limits |
| Cost | Input and output usage for a measured content sample | Model price alone does not reveal whole-site cost |
| Governance | Data region, retention, account controls and audit needs | Content may include confidential or personal information |
Build a realistic translation test set
Use a small, version-controlled sample containing the content editors actually publish:
- short headings and navigation labels;
- long-form editorial copy;
- rich text containing links and inline formatting;
- brand terminology and product names;
- numbers, dates and placeholders that must remain intact;
- ambiguous source phrases that require context;
- content from nested blocks; and
- every target language that materially affects the project.
Have a fluent reviewer score meaning, naturalness, terminology and the amount of editing required. The best model is usually the one that reduces review effort reliably—not necessarily the one that produces the most polished first sentence.
Can smaller, cheaper models translate Umbraco content well?
Yes. In our hands-on testing with Diplo Translator Omni, OpenAI's GPT-4.1 Mini and GPT-4.1 Nano both produced useful results for everyday website translation. Smaller models can be a sensible starting point for navigation, dictionary items, straightforward page copy and other high-volume first drafts, provided a fluent reviewer checks the output.
Treat those model names as a record of testing at the time rather than a permanent recommendation — provider line-ups move quickly, and both have since been superseded by newer small and nano tiers. The finding that matters has held: for everyday website content, the budget tier is usually good enough, and the money saved is better spent on review time than on a larger model. For a new project, the equivalents to evaluate are GPT-5 nano and GPT-5 mini, compared against the provider's current pricing using the same representative test set.
| Model class | Good starting use | What to watch |
|---|---|---|
| Small or nano | Dictionary items, interface labels and straightforward high-volume drafts | Nuance, context and terminology may require more review |
| Mini | A practical balance of translation quality, speed and cost for general site content | Test every important language pair and content type |
| Larger model | Nuanced campaign copy, difficult source material or specialist terminology | Higher cost does not remove the need for human review |
Which tier suits which language pair?
Model choice matters far more for some language pairs than others. Widely resourced European pairs are well served even by budget models; pairs with different scripts, honorific systems or word order expose weaker models quickly. Use this as a starting point for your own testing, not as a substitute for it.
| Language pair | Suggested starting tier | What to watch for |
|---|---|---|
| English ↔ French, German, Spanish, Italian, Dutch, Portuguese | Budget | Formal/informal address (tu/vous, du/Sie) and compound nouns |
| English ↔ Nordic languages, Polish, Czech | Budget to mid | Case endings and inflection after glossary substitution |
| English ↔ Japanese, Korean | Mid to premium | Politeness level consistency across a page; word order in short headings |
| English ↔ Chinese (Simplified or Traditional) | Mid to premium | Correct variant, punctuation width and spacing around inline markup |
| English ↔ Arabic, Hebrew | Mid to premium | Right-to-left rendering in templates and mixed LTR fragments such as URLs |
| Any pair not involving English | Mid to premium | Quality often drops when neither language is the model's strongest |
Right-to-left languages deserve particular care: the translation may be correct while the page still renders badly because the template lacks adir attribute. Verify the rendered page, not only the value in the backoffice. Costs by tier are set out inestimating AI translation costs.
Use profiles to govern the choice
Umbraco.AI chat profiles combine a provider, model and instruction set. A site can maintain a dependable default profile and, when useful, permit alternatives for faster drafts, higher-quality review work or a specific language pair.
How should the system prompt be written?
Keep the system prompt clear and narrow. State that the task is translation, require the original meaning and protected structure to remain intact, identify the intended audience or tone, and tell the model not to add commentary. Put stable terminology in glossary rules rather than growing one unmanageable prompt.
Common questions
Can editors choose different models?
Yes, when administrators intentionally allow multiple profiles. A smaller approved list is easier to support and produces more predictable results than unrestricted provider access.
Should the cheapest model translate the whole site?
Only if testing shows that its review cost remains acceptable. A cheaper request can be a false economy when editors must extensively repair terminology or meaning.
Should sensitive content be sent to an AI provider?
Not by default. Review the provider agreement, retention controls, data location and organisational policy, then exclude content that should not leave the application boundary.
Configure and test a provider
Follow the detailed Umbraco.AI configuration guide, review AI profiles and prompts, and use the safe AI translation workflow for the first production trial.