Kakao AI Can Now Mimic Dialects, Emotions, and Intonations
Translated from Korean, summarized and contextualized by DistantNews.
At a glance
- Kakao has advanced its AI voice generation technology with the 'Kanana-o' model.
- Unlike previous AI that focused on reading text, 'Kanana-o' can reflect user-specified tones, emotions, and intonations.
- Users can request specific reading styles, such as fast-paced, low-pitched, sad, or even regional dialects like Gyeongsang.
Kakao has significantly upgraded its AI voice generation technology, introducing advancements to its proprietary AI model, 'Kanana-o.' The company announced on August 4 that the new capabilities allow for a more nuanced and personalized voice output than previously possible.
While earlier voice AI primarily focused on reading text in a natural-sounding manner, 'Kanana-o' goes a step further by incorporating user-defined speech characteristics. This means the AI can now mimic specific tones, emotional expressions, and intonations as requested by the user, offering a more dynamic and engaging audio experience.
Users can now instruct the AI to read text in various styles. Examples include requests for a fast-paced delivery, a low-pitched voice, or a sad emotional tone. Notably, the technology can also replicate regional accents, with users able to ask for readings in dialects such as the Gyeongsang dialect, showcasing a sophisticated understanding of linguistic variation.
Originally published by Chosun Ilbo in Korean. Translated, summarized, and contextualized by our editorial team with added local perspective. Read our editorial standards.