Cantonese AI Capability 2026: Research and Models
Cantonese AI capability has an academic framework and partial data, with systematic gaps for post-2025 models and models available in Hong Kong
Further reading: Grok vs Claude 2026: Benchmarks and HK Access
Cantonese is Hong Kong's main spoken language and an important component of Hong Kong written language, yet in academic research on AI language capability it has long occupied a relatively marginal position — most model evaluation centres on Mandarin or English, systematic test data for Cantonese is relatively scarce, and among the existing data, some of the best-performing models face access restrictions in Hong Kong.
Coverage of Existing Research
Systematic evaluation of LLM Cantonese capability currently relies mainly on three academic documents:
- arXiv:2408.16756 (August 2024): a systematic evaluation of mainstream LLMs on Cantonese tasks, covering four dimensions — factual generation, mathematical logic, complex reasoning and general knowledge.
- HKCanto-Eval (arXiv:2503.12440, March 2025): a benchmark designed specifically to evaluate LLMs on Cantonese comprehension and Hong Kong cultural awareness, focusing on Hong Kong-specific contexts.
- Measuring HK MMLU (arXiv:2505.02177, May 2025): an MMLU-style evaluation built on Hong Kong contexts.
Model Performance by Task Type (2024 Data)
According to the 2024 research:
- Mathematical logic: GPT-4, GPT-4o and Claude-3.5 perform best, followed by Mixtral-large-2, Llama-3.1-70b and GLM-4
- Complex reasoning: GPT-4 and GPT-4o are consistently best, followed closely by Qwen-2.5-72b and Claude-3.5
- General knowledge (MMLU-type): Qwen-2.5-72b performs consistently best
- Factual generation: Qwen-1.5-110b and Mixtral-large-2 lead
All tested models perform better in English than in Cantonese, making Cantonese a relatively weak area. Note that the data above is based on 2024 research and needs re-evaluation after the 2025–2026 model generation updates. Later-arriving models such as DeepSeek and Grok have not been evaluated on Cantonese specifically in the known studies above.
Distinguishing Three Types of Cantonese Text
Understanding AI Cantonese capability requires distinguishing three different forms of text:
- Written standard Chinese (formal Traditional Chinese): LLMs handle this best overall, but often confuse Hong Kong and Taiwan Traditional Chinese vocabulary (的士 vs 計程車, 地鐵 vs 捷運).
- Written Cantonese (containing Cantonese-specific characters such as 佢, 喺, 唔, 嘅, 囉 and 咁): handling varies by model, and some models, on receiving written Cantonese input, reply in formal written Chinese (Mandarin word order) instead of maintaining a Cantonese style.
- Romanised spoken Cantonese: current LLMs perform weakest here, and TTS intonation for Cantonese is often inaccurate.
Documented Common Errors
Mixed Traditional and Simplified characters: when Traditional Chinese is not explicitly requested, some models output a mix of Simplified characters. Mandarinised grammar: responding to Cantonese input with Mandarin sentence patterns and particles. Recognition of Cantonese-specific characters: characters such as 「霸」「𠮶」「嗰」「咋」 are error-prone in smaller models. Handling of Chinese–English code-mixing: large models are generally workable, but consistency varies.
The Availability Paradox Facing Hong Kong Users
Existing data shows the models with the strongest Cantonese capability (GPT-4o, Claude) face access restrictions in Hong Kong; the models freely available in Hong Kong (Grok, DeepSeek/Qwen) lack systematic Cantonese benchmark records; Gemini (Google) is available in Hong Kong and has some advantage in Chinese contexts, but Cantonese-specific data for it is likewise limited.
Summary
The research picture on Cantonese AI capability is: an academic framework exists, some data exists, but there are systematic gaps (especially in evaluation of post-2025 models and models available in Hong Kong). For actual selection, evaluation on a real test set of business conversations is recommended, in place of relying on a model's general language capability claims — "supports Chinese" does not mean "suits the Hong Kong Cantonese context".
For the engineering of bilingual document retrieval, see Your Document Is Half Chinese, Half English — This Is Where Most AI Systems Fall; for prompt strategy in Traditional Chinese writing, see AI-Assisted Traditional Chinese Writing: Common Errors, Prompt Strategy and Tool Selection; for embedding model testing in mixed-language contexts, see Embedding Model Selection for Production RAG: Four Evaluation Dimensions.
Levi is a Hong Kong-based independent AI engineer specialising in production LLM applications, RAG pipelines, and enterprise AI compliance architecture. Contact for a discussion of the topics covered here.
WhatsApp Free Initial Consultation → More enterprise case studies →Or email: support@hksoka.com