How to Trick AI: The Battle for Personality Privacy in the Age of Chatbots
How to Trick AI: Users' Strategies for Protecting Themselves from Automatic Personality Assessment
This paper investigates "How to Trick AI," specifically examining whether users can strategically manipulate their linguistic and conversational behavior to protect their privacy from automatic personality assessment (APA) chatbots. Using the Juji chatbot and the Big Five personality model (OCEAN), the study identifies 41 distinct user strategies for personality disguise and quantifies their efficacy in inducing falsified profiles.
TL;DR
Can you hide your soul from an algorithm? This CHI '20 study explores how users attempt to "trick" personality-assessment AI. While users can shift their AI-generated scores by nearly 10% through strategic lying and stylistic changes, the sheer mental effort required makes manual privacy protection nearly impossible in the long run.
Context: The Invisible Mirror
We are increasingly living in a world where our word choices, response latencies, and even use of punctuation serve as a "digital DNA." Automatic Personality Assessment (APA) systems, like the Juji chatbot used in this study, can analyze these cues to map us against the Big Five (OCEAN) traits: Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism.
The danger? These profiles are often used for Psychological Targeting—from manipulative political ads (á la Cambridge Analytica) to opaque job screening processes.
The "Tricking" Framework: 41 Strategies of Deception
The researchers identified a massive taxonomy of 41 strategies used by participants to mislead the AI. These can be categorized into four primary domains:
- Linguistic Syntax: Manipulating grammar, casing, and punctuation (e.g., using no periods to appear less "Conscientious").
- Word Choice & Style: Inserting specific keywords, using "filler" words, or lengthening sentences to mask true intent.
- Content Semantics: Reporting "opposite" behaviors or lying about personal interests.
- Conversation Dynamics: Deliberately slowing down or speeding up response times (latency) to appear more introverted or neurotic.
Figure: The gap between what users identify as a cue and what they can actually manipulate in real-time.
Methodology: The Interaction Experiment
The researchers employed a clever two-phase design:
- Phase 1 (Baseline): Participants chatted for 45 minutes in a realistic customer service scenario (booking a holiday, buying a backpack). The goal was to capture their "True" digital footprint.
- Phase 2 (The Trick): Participants were told the AI was harvesting their data and offered a financial incentive to make their second profile differ as much as possible from the first.
Figure: The experimental workflow from natural interaction to strategic manipulation.
Why Tricking AI is Harder Than You Think
The results revealed a fascinating psychological barrier. While participants successfully shifted their scores by roughly 8-10%, they found the process "exhausting and tedious."
- The Implicit Leakage: You might remember to avoid "swear words" to appear more Agreeable, but can you remember to use fewer personal pronouns ()? Most can't.
- Cognitive Load: Trying to maintain a "persona" while simultaneously solving a task (like booking a flight) leads to "fatigue effects."
- Accuracy Paradox: Interestingly, the study found that the AI itself wasn't perfectly accurate compared to traditional BFI-2 questionnaires, yet users still tended to "over-trust" the AI's assessment of them, sometimes questioning their own self-image when the AI disagreed.
Experimental Quantitative Results
The absolute difference achieved across the traits was significant but not transformative:
| Trait | Absolute Difference (%) |
|---|---|
| Openness | 7.57% |
| Conscientiousness | 8.53% |
| Extraversion | 8.25% |
| Agreeableness | 9.53% |
| Neuroticism | 7.98% |

Deep Insight: The Future of Privacy-Protective Design
The authors conclude that self-protection is not enough. If we cannot effectively "fake it" without mental burnout, the burden of protection must shift to designers and policymakers.
Three provocative directions for future tech:
- Machine Non-Readable Content: Developing systems that render text as distorted images for machines but clear text for humans (like "Text CAPTCHAs" for every post).
- Privacy Feedback Tools: Real-time writing assistants that warn you: "This sentence makes you sound highly neurotic; consider this alternative to protect your profile."
- Legislative Enforcement: Moving beyond "endless license agreements" to mandatory opt-in consent for any form of automatic personality inference.
Conclusion
This paper serves as a wake-up call. Personality assessment is no longer the domain of psychologists with clipboards; it is a silent, background process of the modern web. While we can "trick" the AI to a point, our distinct patterns of behavior are incredibly resilient. True privacy in the AI era may require not just better strategies, but a fundamental redesign of how we interact with machines.
