By Mahmoud Kamel
Artificial intelligence (AI) is changing the field of translation, not only because it can produce translations, but also because it is changing the way translators deal with these translations. One of the clearest examples of this change is post-editing, where the human translator starts with a machine-generated translation, then reviews it, corrects it and improves it instead of translating the whole text from the beginning.
With Generative AI systems such as ChatGPT becoming more capable of producing fluent translations, an important question comes to mind: What happens to the translator’s thinking when the first version of the translation has already been produced by AI? Does the translator simply correct the machine’s mistakes, or does working with an AI-generated text require a different kind of thinking and decision-making?
Chapter 9 of Translation Studies in the Age of Artificial Intelligence, written by Longhui Zou, Michael Carl and Jia Feng, looks at this question through an empirical study of student translators working with translations generated by GPT-4. The researchers were interested not only in the final translations produced by the students, but also in what the students did during the post-editing process.
They examined which errors the students corrected, which errors they left unchanged, what new errors they introduced and how their performance changed according to the type of post-editing task and the resources available to them. In this way, the chapter takes the discussion of AI and translation beyond the question of whether machine translation is good or bad and looks more closely at what happens when a human translator works with the machine.
Why post-editing matters
Post-editing is not a new activity in translation. Machine translation has been used in professional translation for many years, and translators have increasingly been asked to review and improve machine-generated texts. Generative AI, however, is changing the situation because systems such as GPT-4 can produce translations that are remarkably fluent and coherent.
This means that translators may spend less time producing the first version of a translation and more time checking, correcting and refining what the machine has already produced. But a fluent translation can still contain serious problems. A sentence may sound natural while giving the wrong meaning. A technical term may look acceptable but may not be suitable for the context. A machine may also leave out information, add information or choose a possible but incorrect meaning.
The translator, therefore, faces a slightly different question. Instead of asking only, “How do I translate this sentence?” the translator may now have to ask, “Can I trust this sentence?” This difference is at the heart of post-editing and shows why working with AI requires not only language skills, but also careful judgement.
Light post-editing and full post-editing
Zou, Carl and Feng look at two different approaches to post-editing: light post-editing and full post-editing. Light post-editing, or LPE, aims to produce a translation that is accurate and understandable without necessarily making it stylistically equal to a fully polished human translation. Full post-editing, or FPE, has a higher quality requirement. The translator is expected to revise the machine output more completely so that the final text meets the standards of a human translation, including grammar, syntax and style.
This difference is important because the mental effort of the translator can change according to the purpose of the task. If the aim is simply to make the machine translation accurate and understandable, the translator may leave some stylistic imperfections as they are. If the aim is to produce a fully polished translation, the translator has to look at the text more carefully and make decisions at different linguistic levels.
So, the same machine-generated sentence may receive different treatment depending on the instructions given to the translator. Post-editing is not only about finding mistakes and correcting them. It is also about deciding how much correction is needed and whether a change is really necessary.
Inside the experiment
To study these questions, the researchers conducted an experiment involving 30 student translators from a prestigious Chinese university. The participants had strong English and translation skills, and most of them had passed advanced translation or English examinations. However, their professional experience in translation and post-editing was relatively limited.
The students were asked to post-edit four English texts and their Chinese translations generated by GPT-4. The source texts, each around 250 words, were taken from American Translators Association certification examinations. They dealt with four different subjects: welfare, tourism, war and racism.
The researchers created four different post-editing conditions by combining two types of post-editing with two types of research resources. The students therefore worked under four conditions: light post-editing with a specialized termbase, light post-editing with internet search, full post-editing with a specialised termbase and full post-editing with internet search.
The participants worked in Trados Studio, a professional computer-assisted translation environment. The researchers also used keylogging and eye-tracking technology to study the translators’ behaviour, although the eye-tracking data were not analysed in this particular study. The final dataset contained 1,399 target segments produced by the participants.
What happens to an AI error?
The study did not simply ask whether a translation was good or bad. Instead, the researchers developed a detailed system for identifying different types of errors, including mistranslation, terminology, usage, grammar and omission or addition. They also divided the errors according to their seriousness into critical and minor errors.
This made it possible to see what happened to an error during post-editing. An error in the original GPT-4 translation could be corrected. It could be noticed but changed incorrectly. It could also be left as it was. There was another possibility which is particularly interesting: a student could introduce a new error into a passage that had originally been correct.
This is an important point because post-editing is not always a process in which the human simply improves what the machine has produced. Human intervention can sometimes make the translation worse. The translator is therefore not only correcting the machine. The translator is also making decisions, and these decisions can lead to both improvements and new problems.
GPT-4 was fluent, but not error-free
One of the important findings of the study concerns the errors in the original GPT-4 translations. The researchers found that 66 per cent of the errors were fluency errors, while 34 percent were accuracy errors. Minor usage errors were especially common.
This tells us something important about Generative AI. The problem is not necessarily that the machine produces translations that are obviously poor. Many of its problems are more difficult to notice. A sentence may look natural and easy to read while still containing a small linguistic or stylistic problem.
At the same time, the researchers found accuracy errors, including mistranslations, terminology problems and cases of omission or addition. This is why human evaluation remains important even when the machine output appears fluent. Fluency can create confidence, but confidence does not always mean accuracy.
Human translators made a significant difference
The student translators were generally successful in improving the GPT-4 output. The raw GPT-4 translation contained an average of 10.32 errors per segment. After post-editing, the students reduced the number of errors by at least 60 percent.
This is one of the clearest findings of the experiment. The students were not just correcting a few obvious mistakes. Their intervention made a substantial difference to the quality of the machine-generated translations.
But the researchers also found another interesting point. Although the students reduced the total number of errors, the distribution of the remaining errors changed. After post-editing, accuracy errors represented a larger proportion of the remaining problems, while fluency errors represented a smaller proportion.
This means that post-editing does not only reduce the number of errors. It can also change the type of errors that remain. When many of the more visible fluency problems have been corrected, questions related to meaning and accuracy become more important.
When more searching does not mean better translation
Another interesting finding concerns the research resources available to the students. The researchers compared the use of a specialized termbase with unrestricted internet search. One might naturally think that access to the whole internet would give translators an advantage because it provides much more information than a limited terminology database.
The results did not support this simple assumption. In general, student performance declined when they relied on internet search compared with the termbase conditions.
This may seem surprising. After all, the Internet contains a huge amount of linguistic and subject-related information. But having more information also means having more choices. When translators search the internet, they have to decide which sources can be trusted, which terminology is suitable, which information is relevant and which answer is actually useful for the translation problem.
The ability to search, therefore, is not the same as the ability to do good translation research. In the age of AI, this difference is becoming more important.
Dr Mahmoud Kamel
Professor at the Academy of Arts











