By Mahmoud Kamel
Artificial intelligence is changing translation in many ways, but perhaps one of the most interesting changes can be seen in a task that is becoming increasingly important for translators: Post-editing.
Post-editing means working on a machine-generated translation and improving it so that it becomes accurate, natural and appropriate for its purpose. With the development of ChatGPT and other large language models, an important question has emerged. Can artificial intelligence perform post-editing as effectively as human translators? And if AI can produce linguistically rich texts, what remains uniquely human in the translation process?
Chapter Seven of Translation Studies in the Age of Artificial Intelligence, edited by Sanjun Sun, Kanglong Liu and Riccardo Moratto, examines these questions in a study titled “Human Expertise vs AI Efficiency: A Comparative Analysis of Student and ChatGPT Post-Editing.” Written by Menglu Li and Dechao Li, the chapter compares student post-editing with ChatGPT-based post-editing in Chinese-to-English tourism translation.
Unlike the previous chapter, which focused on the experiences and concerns of professional translators using ChatGPT, this chapter looks closely at the actual language produced by human and AI post-editors. The researchers are interested not only in which version looks better on the surface, but also in the linguistic choices, strategies and thinking processes behind the two approaches.
A comparison between humans and ChatGPT
The researchers used a mixed-methods approach. They combined quantitative analysis of linguistic features with a qualitative analysis of students’ reflection reports.
The quantitative analysis examined several features, including lexical diversity, lexical density, sentence length ratio and noun-to-verb ratio. The qualitative part examined how students made decisions while post-editing, what problems they noticed and where they concentrated their efforts.
The study focused on Chinese-to-English tourism texts. Two Chinese tourism texts were selected: one introducing Dunhuang, consisting of 320 Chinese characters, and another describing the Taihang Mountains, consisting of 422 characters. Tourism texts were chosen because they occupy an interesting position between technical and literary writing. They must provide information, but they must also attract and persuade potential visitors. They, therefore, require both linguistic accuracy and cultural and communicative adaptation.
The study involved 17 undergraduate students from a university in Hong Kong. The researchers created two groups of post-edited texts. One consisted of translations post-edited by students, while the other consisted of versions post-edited by GPT-4.
Importantly, both groups worked from the same initial machine translations. This provided a common starting point and allowed the researchers to make a more direct comparison between human and AI post-editing. The resulting student post-editing corpus contained 34 texts and 9,521 tokens, while the GPT post-editing corpus contained the same number of texts but 10,407 tokens.
Where ChatGPT shows its strength
The quantitative results produced an interesting picture.
ChatGPT-based post-editing showed higher lexical diversity than student post-editing. The median lexical diversity score was 0.59 for GPTPE, compared with 0.56 for student post-editing. The difference was statistically significant.
The same pattern appeared in lexical density. GPTPE (ChatGPT Post Editing) had a median lexical density of 0.56, compared with 0.53 for student post-editing. This difference was also statistically significant.
In simple terms, the ChatGPT versions used a wider range of vocabulary and contained a higher proportion of content words. The researchers suggest that this may be related to the enormous linguistic resources available to large language models such as GPT-4.
This is one of the interesting characteristics of AI-assisted translation. A machine does not have the limited active vocabulary of an individual student translator. It can draw on enormous amounts of linguistic data and generate a wide range of lexical choices.
But this is where the study asks us to be careful.
More vocabulary does not automatically mean a better translation.
A text can be lexically rich and still fail to communicate appropriately with its readers. Translation quality also depends on context, coherence, cultural meaning and the purpose of the text. The researchers therefore argue that the higher lexical diversity and density of GPTPE should not be understood as evidence that AI has replaced human expertise.
What about sentence structure?
The comparison of sentence length produced a different result.
The median sentence length ratio was 0.22 for student post-editing and 0.20 for GPTPE. However, the difference was not statistically significant.
This means that, in these Chinese-to-English tourism texts, the human and AI post-editors showed broadly similar patterns in sentence structure.
The result is interesting because it reminds us that not every linguistic difference between humans and AI is equally important. Some features may be quite similar, while others reveal much clearer differences in the way the two approaches work.
The noun-to-verb ratio, for example, showed a statistically significant difference. The median ratio was 4.21 for student post-editing and 3.83 for GPTPE.
The researchers interpret the higher ratio in the student texts as evidence of greater nominalisation and stronger adaptation to English linguistic conventions. The lower ratio in the GPTPE texts suggests a more balanced use of nouns and verbs and, potentially, a closer structural relationship to the source text.
The human translator sees the reader
The most revealing findings, however, come from the qualitative analysis.
The students did not simply correct grammatical mistakes. They often tried to adapt the translation to the expectations of English-speaking readers.
This is particularly important in tourism translation.
A tourism text has at least two functions. It informs the reader, but it also tries to attract the reader. It describes a place while encouraging people to visit it.
The students appeared to understand this dual purpose.
They changed the tone, made expressions more natural and sometimes addressed the reader directly. One of the student versions transformed a relatively impersonal description into a more engaging text by using a second-person approach. Another used an inviting imperative and divided the information into shorter, more accessible segments.
This is where the difference between human and AI post-editing becomes especially interesting.
ChatGPT produced a fluent and grammatically accurate version, but it remained closer to the imagery and structure of the Chinese source. It retained expressions such as the metaphor of a “hundred-mile gallery.”
The students, by contrast, were more willing to reshape the text according to the expectations of the target reader.
The difference is not simply between correct English and incorrect English.
Both can produce grammatical English.
The deeper difference is between translating words and mediating communication.
When grammar is not enough
One of the clearest conclusions of the chapter is that grammatical accuracy is only one part of translation competence.
The analysis found that ChatGPT performed particularly well in grammatical accuracy. But the student translators showed greater sensitivity to culture-specific references and to the conventions of the tourism genre.
This distinction is important.
A tourism translator does not merely transfer information from Chinese into English. The translator has to think about what a tourist needs to know, how the information should be presented and how the language can make the destination attractive without distorting the original message.
The students showed awareness of this communicative purpose.
They were acting not simply as editors but as communicative mediators between the source text and the target audience.
This is perhaps the most significant human contribution identified by the study.
Why humans sometimes edit less
Another interesting finding concerns the amount of effort students invested in post-editing.
During light post-editing, students often kept parts of the original machine translation unchanged, even when those parts could have been modified. Their reflections suggest that they tended to concentrate their attention on the more difficult parts of the text.
In particular, they focused on passages containing cultural references rather than spending their time changing every lexical choice.
This may help explain why their final texts showed lower lexical diversity and density than the GPTPE texts.
The students were not necessarily trying to improve every word.
They were making decisions about where their attention was most needed.
This is a very human way of working.
A translator has limited time and limited cognitive resources. The translator therefore has to decide which problems matter most. A minor lexical variation may not require much attention, while a culturally sensitive expression may require considerable thought.
ChatGPT does not face the same cognitive limitations. It can generate a linguistically rich version without having to decide where to spend its limited attention in the same way a human translator does.
More words do not mean more understanding
This leads to one of the most important lessons of the chapter.
The GPTPE texts were stronger in lexical diversity and density. But the qualitative analysis showed that the students had a stronger awareness of context, communicative purpose and the needs of the target audience.
This means that quantitative linguistic superiority does not necessarily translate into qualitative appropriateness.
A text can contain more varied vocabulary and more information-rich language and still miss an important cultural association or fail to create the desired effect on the reader.
For translation, this distinction matters greatly.
The purpose of translation is not to demonstrate how many words a system can produce. The purpose is to communicate meaning effectively to a particular audience in a particular context.
In this respect, the study presents human expertise and AI efficiency as complementary rather than simply opposing forces. AI can contribute speed and linguistic variety, while human post-editors can concentrate on context, culture, pragmatics and audience expectations.
The future translator as a user-centred mediator
The findings have important implications for translator education.
If AI systems are becoming increasingly capable of producing fluent and lexically rich translations, translation students cannot be trained only to produce a first draft from scratch.
They also need to learn how to evaluate machine output.
They need to recognise when a sentence is grammatically correct but culturally inappropriate. They need to understand genre conventions, identify problems of meaning and connotation and adapt translations to the needs of target readers.
In other words, post-editing should not become a mechanical exercise in correcting mistakes.
It should become an intellectual exercise in making decisions.
The researchers therefore emphasise creativity, critical evaluation, problem-solving and decision-making in translator education. Students should learn how to use AI efficiently while continuing to develop the human skills that AI does not easily reproduce, particularly cultural adaptation and pragmatic refinement.
Learning when to intervene
There is another lesson here for translators who already use AI.
A common temptation is to assume that the more changes we make to a machine translation, the better the final text will become.
The study suggests something more subtle.
Human translators may be most valuable when they know where intervention is really necessary.
They can allow the machine to handle some of the linguistic work while directing their own attention towards the areas where human judgement has greater value.
This could lead to a new model of translation work.
Instead of spending equal effort on every sentence, the translator may increasingly become a selective editor, cultural mediator and quality controller.
AI produces possibilities.
The human translator decides which possibilities work.
The limits of the study
The authors are careful not to present their findings as universal.
The study involved only 17 undergraduate students from one Hong Kong university. The participants were novice translators, and the research focused specifically on Chinese-to-English tourism texts.
This means that the results cannot automatically be applied to professional translators, other language combinations or other types of texts such as legal, literary, medical or highly specialised translation.
The researchers therefore call for larger studies involving different language pairs, text types and levels of translator experience. They also suggest longitudinal research that could examine how post-editing skills develop over time.
This caution is important.
The study does not tell us that ChatGPT is better than humans, nor does it tell us that humans are always better than AI.
It shows something more interesting.
Different forms of intelligence produce different kinds of strengths.
Human expertise meets AI efficiency
Chapter Seven presents a picture of translation in which the relationship between human translators and AI is becoming more complex.
ChatGPT can produce texts with greater lexical diversity and density. It can generate language quickly and draw on an enormous linguistic resource.
Human post-editors, even at the student level, can bring something different. They can pay closer attention to the target reader, the communicative purpose of the text, cultural references and genre conventions. They can decide when a literal rendering is not enough and when a translation needs to be reshaped to communicate effectively.
The future described by this chapter is therefore not a simple competition between humans and machines.
It is a question of division of labour.
AI may take on more of the linguistic processing.
Humans may need to take greater responsibility for interpretation, adaptation and judgement.
For translation students, this means that learning English and translation techniques will no longer be enough. They will also need to learn how to work with AI, how to question its output and how to recognise the moments when human intervention matters most.
For professional translators, the message is equally important.
The value of the human translator may increasingly lie not in competing with the machine at everything the machine can do, but in doing what the machine still finds difficult.
A machine can produce a text.
A translator must decide what that text is trying to achieve, who it is speaking to and whether it has achieved its purpose.
That difference may become one of the defining features of translation in the age of artificial intelligence.
Dr Mahmoud Kamel
Professor at the Academy of Arts











