AI Workflows Outscored Human Translators In 4 Of 6 Content Types – China Benchmark Study

We carried out a benchmark examine of English-to-Chinese language localization, evaluating content material sorts, activity sorts and workflow fashions to check human efficiency towards machines.

Skilled human translators scored exterior the highest 5 on 4 of six content material sorts. On advertising copy, they completed tenth of 15, 22.2 factors behind one of the best AI workflow. The place people did win, the margin over one of the best machine workflow was 2.8 factors.

We measured 774 localized outputs throughout six content material sorts, seven workflow fashions (human, Chinese language LLM*, Chinese language LLMPE**, Western LLM*, Western LLMPE**, MT*, MTPE**), and three activity sorts (translation, transcreation, creation). Every output was scored on accuracy and consistency, fluency and language high quality, and elegance and cultural adaptation – every dimension weighted equally at one third of the overall. Scoring was carried out blind*** by Chinese language native-speaking skilled localizers.**** This measured localization high quality solely. No rating, visitors, or conversion information was collected.

Disclosure: The benchmark mentioned here’s a joint analysis challenge by EC Improvements, a localization providers supplier, and Jademond Digital, the place I’m a Accomplice and Director. Each companies promote providers associated to the practices this analysis evaluates.

People Misplaced 4 Of 6 Content material Varieties

Skilled human localization was examined towards fourteen machine and hybrid workflows throughout six content material sorts. It completed first on two of them. On the opposite 4, it completed exterior the highest 5 – seventh, seventh, ninth, and tenth out of 15.

Content material sort

Human rating

Human rank

Prime workflow

Informational

76.9

1st

Human (76.9)

website positioning

74.1

1st

Human (74.1)

Technical

64.8

seventh

PE-Qwen (79.6)

Product UI

63.0

seventh

PE-Qwen (73.1)

UGC

64.8

ninth

PE-Doubao / PE-Qwen (75.9)

Advertising

53.7

tenth

PE-Qwen (75.9)

Desk 1: Human localization completed first on two content material sorts and out of doors the highest 5 on 4. Ranks are competitors ranks out of the 15 workflows examined.

Human professionals gained two of six classes and completed tenth of 15 on advertising content material, 22.2 factors behind post-edited Qwen. PE-Qwen gained three classes outright, tied a fourth, and positioned second on the remaining two.

That is the precise form of the discovering, and it’s extra attention-grabbing than “people win website positioning.” Human experience is decisively higher at a particular factor: content material the place factual precision and terminological consistency dominate, and artistic latitude is close to zero. That describes informational and website positioning content material. It doesn’t describe advertising copy, UI strings, or social content material – and on these, paying skilled translation charges buys a worse consequence.

EC Improvements’ studying of the advertising result’s value repeating:

Human translators seem to over-correct the language, smoothing copy towards formal correctness and stripping out the up to date register that advertising content material depends upon. The identical intuition that makes a linguist wonderful at terminology self-discipline makes them a poor match for writing that should sound just like the web.

Nothing Right here Demonstrates A Rating Impact

We measured localization high quality. We didn’t monitor a single SERP place, and nobody ought to current a 22-point – or a 2.8-point – high quality hole as a rating end result. The revealed proof connecting content material high quality metrics to rating is thinner than the business typically concedes. Portent’s crawl of 756,297 rating pages discovered no correlation between readability and Google rating place (Portent, 2021), and particularly known as out the circularity of companies that promote content material high quality asserting that content material high quality drives rankings. That criticism lands on analysis funded by localization firms too. Deal with all the things on this article as an enter speculation to check, not a demonstrated rating impact.

The place People Nonetheless Win And By How Little

People gained two classes. Each wins are narrower than the headline suggests.

Right here is the complete per-model rating for website positioning content material – meta descriptions, headlines, keyword-carrying physique copy – from a dataset that has not been revealed at this granularity:

Rank

Workflow

Rating

1

Human skilled

74.1

2

PE-Doubao

71.3

2

PE-Qwen

71.3

4

PE-DeepSeek

66.7

4

PE-Gemini

66.7

4

PE-Google MT

66.7

7

PE-ChatGPT

65.7

7

Qwen (uncooked)

65.7

9

ChatGPT (uncooked)

62.0

10

DeepSeek (uncooked)

61.1

10

PE-Kimi

61.1

12

Doubao (uncooked)

59.3

12

Gemini (uncooked)

59.3

14

Kimi (uncooked)

56.5

15

Google MT (uncooked)

55.6

TABLE 2: How authors, instruments, and workflows compete on website positioning content material. “PE” stands for post-editing by a human skilled; “Google MT” is Google Translate. The human lead over one of the best AI workflow is 2.8 factors.

People nonetheless end first, however the hole to one of the best accessible AI workflow is simply 2.8 factors and given that every rating is a median of three equally weighted dimension scores, a sub-three-point distinction is contained in the vary the place I’d not wish to make a six-figure sourcing resolution with out replicating it alone content material.

Put up-edited Qwen or Doubao is roughly at parity with skilled human localization for website positioning content material and uncooked output from any mannequin will not be.

There may be additionally a discovering right here for anybody operating legacy infrastructure. Put up-edited Google Machine Translation (MT) scored 66.7 on website positioning. Forward of uncooked Qwen and forward of each uncooked mannequin within the examine. In case you have a mature MT pipeline with translation reminiscence and a longtime termbase, including a post-editing layer will get you additional than replatforming onto a uncooked LLM would.

Why The Class Common Hid All Of This

On website positioning content material, skilled human translation scored 74.1 out of 100. Uncooked Chinese language LLM output and uncooked Western LLM output each scored a median of 60.7. A 13.4-point hole in favor of people. That quantity is as near meaningless as a procurement enter, and understanding why is essentially the most helpful factor an website positioning can take from this dataset.

“Uncooked Chinese language LLM: 60.7” is a straight common of 4 fashions: Qwen, Doubao, DeepSeek, and Kimi. Here’s what these 4 truly scored on website positioning content material:

Mannequin

website positioning Content material rating

Qwen (uncooked)

65.7

DeepSeek (uncooked)

61.1

Doubao (uncooked)

59.3

Kimi (uncooked)

56.5

TABLE 3: Chinese language LLMs scoring unfold on website positioning Content material – Averaged: 60.7

A 9.2-point unfold. And website positioning is the tightest class within the examine. On technical content material, the identical 4 fashions span 22.3 factors, from Qwen at 70.4 to Kimi at 48.1. On advertising, 22.2 factors.

Mannequin

website positioning Content material rating

Gemini (uncooked)

59.3

ChatGPT (uncooked)

62.0

TABLE 4: Western LLMs scoring unfold on website positioning Content material – Averaged: 60.7

No person deploys “Chinese language LLM.” They deploy Qwen, or Doubao, or DeepSeek. The class common describes a mannequin that doesn’t exist.

CHART 1: Unfold throughout the “Chinese language LLM” class: The class common conceals variations of as much as 22.3 factors between the fashions it comprises. (Picture by writer, September 2026)
CHART 2: Unfold throughout the “Western LLM” class: The class common conceals variations of as much as 8.4 factors between the fashions it comprises. (Picture by writer, September 2026)

The 2 classes don’t behave the identical approach. Throughout all six content material sorts, the widest hole between the 2 Western fashions is 8.4 factors, on informational content material. Among the many Chinese language 4, it reaches 22.3, on technical. In case you have standardized on GPT or Gemini, “Western LLM” is a roughly sincere description of what you’re going to get. “Chinese language LLM” will not be an outline of something. It’s the common of a area whose finest and worst members sit greater than twenty factors aside, and the mannequin you truly deploy could possibly be at both finish.

Within the website positioning numbers, each classes common 60.7. Similar abstract statistic, completely totally different distribution behind it.

Gemini leads the uncooked area on user-generated content material at 67.6, the one content material sort within the examine the place a Western mannequin finishes forward of each Chinese language one (with out post-edit).

The identical downside reveals up from the opposite course. Kimi completed final among the many Chinese language 4 in all six content material sorts, and eradicating it, which is what any enterprise does the second it runs a two-week bake-off, strikes the “Chinese language LLM” determine by between 1.3 and 4.9 factors relying on the class. The class common will not be merely imprecise. it’s being dragged by a mannequin most consumers would eradicate in week one.

This isn’t a criticism of how the examine aggregated – equal-weight averaging throughout an outlined mannequin set is the proper solution to characterize a class. It’s a warning about how the ensuing quantity will get used. When you learn “Chinese language LLMs rating 60.7 on website positioning content material” and conclude Chinese language LLMs are unfit for website positioning work, you’ve drawn a conclusion the underlying information doesn’t assist.

Put up-Enhancing Is Not A Uniform High quality Layer

Put up-editing is often described as a top quality layer you both purchase otherwise you don’t. The class view reveals it’s nothing of the kind.

CHART 3: Put up-editing elevate by mannequin class, per content material sort. Put up-editing virtually at all times improves output, excluding advertising content material drafted by Western LLMs. The most important enchancment is on machine-translated UGC. (Picture by writer, September 2026)

On user-generated content material, a post-editing go over Google Translate output is value +30.6 factors. On advertising content material, the identical go over a Western LLM draft is value -0.9. The variable will not be how a lot enhancing you purchase. It’s whether or not the draft you hand the editor is shut sufficient to proper that enhancing improves it, or fallacious sufficient that the editor spends the finances preventing it.

That UGC determine – 33.3 uncooked to 63.9 post-edited – is the biggest single motion within the examine, and an unsurprising one. Uncooked machine translation of casual Chinese language social copy begins from near-unusable.

The identical unevenness holds mannequin by mannequin:

CHART 4: Put up-editing elevate by base mannequin averaged throughout all content material sorts. (Picture by writer, September 2026)
CHART 5: Put up-editing elevate by base mannequin on Technical content material: The identical editorial course of advantages some fashions greater than others. (Picture by writer, September 2026)
CHART 6: Put up-editing elevate by base mannequin on Advertising content material: Each ChatGPT and Kimi lose high quality after post-editing. (Picture by writer, September 2026)

Extra instructive than the averages are the three circumstances the place post-editing made output worse: PE-ChatGPT on advertising (-3.7), PE-Kimi on advertising (-2.8), and PE-ChatGPT on technical (-0.9). On technical content material, the unfold runs the complete width of the identical sample, from PE-ChatGPT at -0.9 to PE-Qwen at +9.3.

Put up-editing will not be a monotonic enchancment. Utilized with the fallacious register in thoughts – formalizing advertising copy, or enhancing technical content material with out area data – a human go destroys worth. This is identical over-correction impact seen within the human-only advertising scores.

Which factors on the factor I obtained fallacious myself. Mannequin selection issues greater than the enhancing layer, not much less. I’ve seen the alternative argued from the class averages, and I initially learn the information that approach too. Put up-editing elevate averages 5.6 factors. The unfold between one of the best and worst Chinese language mannequin runs as much as 22.3 factors inside a single content material sort. Selecting Qwen over Kimi is a bigger resolution than whether or not you post-edit in any respect.

Why Chinese language website positioning Adjustments The Calculus

All the pieces above measures content material high quality. For Chinese language website positioning particularly, content material high quality is never the binding constraint.

E-E-A-T is a Google framework and doesn’t switch cleanly to Baidu, which weights site-level signals – area historical past, ICP submitting standing, and internet hosting geography – extra closely than page-level language high quality. Hosting geography matters partly through latency: a website served from exterior the mainland is slower behind the border, and Baidu treats sluggish websites unfavorably. An ICP submitting, and mainland or Hong Kong internet hosting to go together with it, can do extra for visibility than the distinction between a 60.7 and a 74.1 translation rating.

That doesn’t make localization high quality irrelevant. It does imply the sequencing issues. In case your ICP submitting and internet hosting should not sorted, upgrading from post-edited Qwen to full human translation is an optimization layered on high of a constraint you haven’t eliminated – and it’s the costlier of the 2 fixes.

What To Truly Do

  • Cease evaluating “Chinese language LLMs” as a class. Consider Qwen towards Doubao towards DeepSeek by yourself content material. The within-category unfold is bigger than the human-versus-AI hole that will get all the eye.
  • For flagship website positioning and informational content material, use human translation or post-edited Qwen or Doubao. The distinction between them is sufficiently small that value and turnaround ought to resolve it.
  • For advertising, UI, technical, and social content material, cease paying for human translation. The info says you’re shopping for a worse consequence at the next worth. Put up-edited Qwen led all 4.
  • When you run a mature MT pipeline, add post-editing earlier than you take into account replatforming. PE-Google MT at 66.7 on website positioning beat each uncooked mannequin examined.
  • Construct the termbase. Terminology consistency is the place uncooked mannequin output fails most reliably on website positioning content material, and a glossary is the most affordable management for it. It additionally survives each mannequin change you’ll make within the subsequent two years.
  • Tag localized pages by workflow and watch what occurs. Inside two quarters, your own analytics will reply the rating query this examine couldn’t – in your verticals, your aggressive set.

When This Was Measured, And Why It Issues

A phrase on timing. The outputs had been produced in December 2025 and January 2026. Blind analysis ran till early March – Chinese language New 12 months sits in the course of that window and slows all the things in China down. Evaluation and report preparation took till the tip of Could, and the examine was revealed on June fifth.

Each mannequin on this examine was examined at its December 2025 model, and on this area a months-old snapshot is a historic doc. Qwen, Doubao, GPT, and Gemini have all shipped since. The precise rating above could not survive contact with the present releases.

The sturdy discovering will not be “use Qwen.” It’s that the variations between particular person fashions, in your particular content material sort, are massive sufficient to be value measuring your self. A benchmark that goes stale in months is an argument for operating your individual, not for ready on another person’s. A two-week inside bake-off throughout your three or 4 candidate fashions, on a consultant pattern of your individual content material, will inform you greater than any revealed benchmark, together with this one.

Run it. Then run it once more in six months.

What Was Measured:

* The LLM and MT methods examined had been GPT-5.2 (through ChatGPT), Gemini 3.0, Doubao 1.6, Qwen 3, Kimi K2, DeepSeek-V3.2, and Google Translate – all accessed by means of their internet interfaces.
** “PE” all through stands for post-editing by human professionals.
*** Scoring was achieved by skilled Chinese language localizers who didn’t know which textual content was produced by which workflow. The scorers had been totally different folks from those that produced the human and post-edited variations.
**** Take a look at window: localized outputs had been produced from December 2025 to January 2026. Blind analysis ran to early March 2026. Evaluation and report preparation continued by means of Could, and the examine was revealed in June 2026. Each mannequin model listed above is the model accessible in December 2025 and January 2026.

Full methodology and benchmark report. The per-model breakdowns on this article transcend what the revealed report comprises.

Extra Assets:


Featured Picture: Summit Artwork Creations/Shutterstock


#Workflows #Outscored #Human #Translators #Content material #Varieties #China #Benchmark #Examine

Leave a Reply

Your email address will not be published. Required fields are marked *