Could Stress Make AI Work Better?
Translated from Korean and summarized by DistantNews. Read the original for the full story.
At a glance
- A Pennsylvania State University study found GPT-4o’s average accuracy rose from 80.8% with very polite prompts to 84.8% with very rude prompts.
- Researchers cautioned that the result may reflect the directness of commands rather than an emotional response to insults.
- A separate 2024 study found extremely rude prompts reduced performance in several models, although GPT-4 performed slightly better under the harshest wording.
GPT-4o answered multiple-choice questions more accurately when researchers addressed it rudely rather than politely, according to a study from Pennsylvania State University.
The researchers tested 50 questions in mathematics, science, and history. They placed five levels of wording before each question, ranging from very polite to very rude. A very polite prompt asked, “Would you review the following problem and provide an answer?” A very rude version said, “Poor thing, do you even know how to solve this?”
Average accuracy reached 80.8% with very polite prompts and 84.8% with very rude ones. But the researchers warned against concluding that the AI felt insulted and concentrated harder. The rude prompts included not only insults but also forceful instructions such as telling the system to focus, making it difficult to separate the effect of rudeness from the effect of more direct commands.
A 2024 paper by researchers from Waseda University, RIKEN, and other institutions reached a different result. Testing GPT-3.5, GPT-4, and Llama 2, it found that extremely rude prompts reduced performance. GPT-4 was the exception, performing slightly better under the harshest wording than under the most polite wording. The finding suggested that newer, stronger models may be less affected by peripheral factors such as tone.
Originally published by Hankyoreh in Korean. Translated, summarized, and contextualized automatically by DistantNews, with a note on how the source frames the story. Not individually reviewed before publishing. How this works.