ChatGPT’s Kazakh-Language Push Shows Why AI Must Go Beyond English
Translated from English and summarized by DistantNews. Read the original for the full story.
At a glance
- A project involving the Qazaq Tili international association and OpenAI has created Kazakh-specific benchmarks for grammar, speech, proverbs, translation, children’s literature, safety and ethnography.
- Kazakhstan’s Kazakh Text Corpus contains 14 billion tokens from fields including education, science, law, medicine, history and media.
- The initiative aims to improve ChatGPT in Kazakh and reduce culturally inaccurate responses, reflecting a broader push for AI beyond English.
Teaching ChatGPT Kazakh involves more than translating English sentences. Kazakhstan is building language data and evaluation tools that test whether artificial intelligence understands the culture carried by the language.
A joint project between the Qazaq Tili international association and OpenAI has developed an AI Evaluation Benchmark Suite specifically for Kazakh. Created originally in Kazakh rather than translated from English, it assesses grammar, natural speech, proverbs and fixed expressions, academic and literary translation, children’s literature, safety and ethnography. A separate 500-question benchmark tests knowledge of Kazakh history, traditions and culture.
The project is also expanding the data available for Kazakh-language systems. The Kazakh Text Corpus has reached 14 billion tokens and covers different periods of language development and diaspora heritage. Its material spans education, science, technology, economics, law, medicine, history, ethnography, media and children’s content, using text, audio and images.
I think it is also going to be about context and culture. And I think we are definitely missing a trick. The problem is that most AIs are in English, but it should be in Navajo. It should be in hundreds of languages that we have. So it gets a much more deeper context.
Rauan Kenzhekhanuly, president of the Qazaq Tili international association, said more than 10 billion tokens were initially collected from archives, museums, libraries and other sources after the necessary permissions were obtained. The goal is to improve ChatGPT’s translation, information services and speech, text and audio recognition while reducing responses that do not fit Kazakh cultural realities.
Futurist and author Ron Immink said AI development must account for context and culture. “I think it is also going to be about context and culture. And I think we are definitely missing a trick. The problem is that most AIs are in English, but it should be in Navajo. It should be in hundreds of languages that we have. So it gets a much more deeper context,” he said. OpenAI education lead Valerie Focke called Kazakhstan “a pioneer and early mover” in AI adoption and said it was among the first countries to work with OpenAI on the future of education.
a pioneer and early mover
Originally published by The Astana Times in English. Translated, summarized, and contextualized automatically by DistantNews, with a note on how the source frames the story. Not individually reviewed before publishing. How this works.