arXiv cs.LGOctober 2, 2026
Comedic Fool's Gold: Reward Exploits and Countermeasures in Conversational Humor
Excerpt
arXiv:2610.00197v1 Announce Type: cross Abstract: We investigate automated rewards for training language models in conversational humor, focusing on reward exploits and countermeasures. Two approaches aim to capture understandable surprise and predicted audience amusement. Controlled tests show that an embedding-based surprise reward accepts word-shuffled replies as readily as witty ones. A fluency filter detects the shuffles, but the combined reward also rejects some witty replies and fails fur