By Mahmoud Kamel
Artificial intelligence is entering another area of translation where language alone is not enough: audiovisual translation and subtitling.
A translated subtitle is not simply a sentence transferred from one language into another. It must work with the image, the movement, the voices, the humour and the rhythm of a scene. It must also be short enough to be read while the action continues on the screen.
This makes subtitling a particularly interesting test for Generative AI.
Chapter 8 of Translation Studies in the Age of Artificial Intelligence, by Zhuojia Chen and Zhiwei Wu, examines this question through a comparative case study involving human, ChatGPT and Google Gemini. The researchers focus on three difficult areas of subtitling: wordplay, multimodal humour and features of spoken language.
Why subtitling is different
Subtitling has always presented translators with problems that do not occur in ordinary written translation.
A subtitle has limited space and limited time. At the same time, the translator is working with several modes of communication at once. What a character says may have a completely different meaning when combined with a facial expression, a gesture or an object appearing on the screen.
Wordplay is one of the most difficult examples. A joke may depend on two words sounding alike or on one word having several meanings. Translating the dictionary meaning may preserve the information but destroy the joke.
Humour can be even more complicated when words and images work together. A character may say something funny while pointing at an object, making a facial expression or performing a particular gesture. The subtitle, therefore, has to communicate with not only what is said but also what is happening visually.
The third challenge is orality. Real speech contains repetitions, hesitations, incomplete sentences and interjections. These features can tell us something about a character and the way a conversation is unfolding. Removing all of them may make a subtitle grammatically cleaner but less natural and less faithful to the original performance.
A comparison between humans and AI
To investigate these issues, Chen and Wu selected a scene from the American sitcom Modern Family. They compared the existing subtitles produced by fansubbers (fans who create subtitles for audiovisual content) with subtitles generated by GPT-4-based ChatGPT and Google Gemini.
The researchers did not simply ask the AI systems to translate the dialogue. They experimented with different prompts, including prompts containing contextual information and prompts asking the AI to assume the persona of a fansubber or a professional subtitler.
This is an important part of the study because it raises another question about Generative AI: Does the way we ask the machine to translate affect the quality of the result?
The answer from this case study is clearly yes.
When one word has several meanings
One of the examples concerns the expression “Puppy Pound.” The word “pound” can refer to a unit of money or weight, but in the particular context of the scene it refers to a toy shelter for puppy dolls.
The problem becomes even more complicated because the dialogue later plays with the similarity between “Puppy Pound” and “Guppy Pond.”
This is not a simple vocabulary problem. The translator has to understand the situation, recognise the wordplay and then find a way to reproduce the humorous effect for another audience.
The study found that both human and AI translators faced difficulties here. Many of the AI versions concentrated on the literal or semantic meaning and failed to recreate the original wordplay effectively.
This is an important reminder that linguistic accuracy and translation creativity are not always the same thing.
When the picture becomes part of the joke
The researchers also examined multimodal humour.
In one example, a character points with his finger while speaking. The humour depends on the relationship between his words and the physical gesture.
The fansubbers were able to use the Chinese subtitle to make the relationship between the gesture and the dialogue clearer. Several AI-generated versions, however, understood the words without fully capturing the humorous connection between the words and the visual action.
This is one of the areas in which audiovisual translation exposes a fundamental limitation of purely textual processing.
The subtitle is not an isolated sentence. It is part of a scene.
ChatGPT and Gemini do not behave in exactly the same way
The study also found a difference between the two GenAI platforms.
Gemini generally performed better than ChatGPT in dealing with the subtitling challenges examined in the study, particularly wordplay and multimodal humour.
The authors suggest that Gemini’s ability to work with multimodal information may partly explain this difference. Since audiovisual translation involves both verbal and visual information, access to visual context can potentially improve the quality of the translation.
However, the researchers are careful not to turn this case study into a general claim about all AI systems. Their findings are based on one selected scene and are intended to generate hypotheses for future research rather than establish universal conclusions.
The importance of the prompt
Perhaps one of the most interesting findings concerns prompting.
Some prompts produced better subtitles when they included contextual information. A prompt that gave the AI more information about the scene could help it understand who was speaking, what the characters were doing and why a particular expression was humorous.
The researchers also found that a prompt using a fansubber persona performed better than one describing the system as a professional subtitler.
This does not mean that simply telling AI to “be a fansubber” will automatically produce creative subtitles. Rather, it suggests that the information and perspective supplied through the prompt can influence the way GenAI approaches a translation problem.
At the same time, more information was not always better. Some combined prompts produced poorer results, suggesting that excessive instructions can create information overload or conflicting requirements.
In other words, prompt engineering itself becomes part of the translator’s skill.
Where human creativity still matters
The most significant finding of the chapter may be found in the comparison between AI and fansubbers.
The fansubbers demonstrated a stronger ability to deal creatively with cultural and multimodal nuances. They did not simply translate the words. They considered the communicative effect of the entire scene.
The researchers even went one step further. After seeing the difficulties faced by both AI systems with wordplay, they experimented with additional prompts to see whether GenAI could be guided towards more creative solutions.
The results showed that AI could generate possibilities, but the human researchers still had to evaluate those possibilities and decide whether they actually worked.
This distinction is important.
The challenge in subtitling is not only to produce a grammatically correct sentence. It is to produce a subtitle that works for the viewer.
The problem of spoken language
The treatment of orality provides another revealing example.
The fansubbers were able to represent hesitation and repetition in ways that reflected the spoken performance. ChatGPT and Gemini generally preserved some hesitation and interjections, but they tended to omit repetition and could misrepresent incomplete utterances.
The authors, therefore, suggest that more specific instructions may be necessary if GenAI is expected to follow particular subtitling guidelines concerning spoken language.
Again, the issue is not simply whether the AI understands the words. It is whether it understands why the speaker is speaking in that particular way.
AI as a subtitling assistant
The findings of Chapter 8 do not suggest that Generative AI has no place in subtitling.
On the contrary, the study points towards a practical role for AI in the subtitling workflow. GenAI can produce drafts rapidly, offer alternative formulations and help translators explore possible solutions.
This can be particularly useful when speed is important. The authors note that Gemini, for example, may be a fit-for-purpose solution in situations where subtitles need to be produced quickly and only light post-editing is required.
But subtitles intended for high-quality publication, where humour, wordplay and the overall viewing experience are important, still require careful human checking and creative intervention.
The translator remains the final decision-maker
Chapter 8, therefore, adds another dimension to the discussion of AI and translation.
The future of audiovisual translation may not be a choice between humans and machines. It may involve a workflow in which AI produces possibilities while human subtitlers determine which possibilities actually communicate the meaning, humour and cultural effect of the original scene.
This is particularly important in audiovisual translation because the translator is not working with words alone. The translator is working with a complete communicative event involving language, image, sound, gesture, culture and emotion.
Generative AI is becoming increasingly capable of processing these elements. But the study by Chen and Wu shows that capability does not automatically produce creativity.
The machine can translate the dialogue.
The human subtitler must still ask a more difficult question:
Will the viewer laugh at the same moment?
That may be one of the most important questions for the future of AI-assisted audiovisual translation.
Dr Mahmoud Kamel
Professor at the Academy of Arts











