A critical look
Limits and pitfalls
Personality tests can describe patterns, but they cannot predict what you will do in a single situation. The most important limits are the Barnum effect, where vague descriptions feel accurate to everyone, the biases in self-reporting, and the way typologies turn sliding scales into boxes.
Know the limits of the frameworks
Any serious approach to understanding personality has to acknowledge the limits of its frameworks. This chapter gives you the tools to read your own results with a critical, but constructive, eye.
The Barnum effect: when everything feels true
In 1949, the psychologist Bertram Forer gave all his students an identical "personal" assessment, taken from an astrology book. Average rating: 4.26 out of 5 for accuracy.
Four psychological mechanisms
- Vague wording: "Sometimes" fits everyone
- Positivity bias: We accept flattering descriptions more readily
- Authority effect: A "scientific" source increases acceptance
- Confirmation bias: We look for evidence that supports the description
Implication: A test can feel remarkably accurate without measuring anything meaningful. Customer satisfaction is no guarantee of validity.
Typologies: popular, but limited
Typological personality tests have drawn substantial criticism:
- Test-retest reliability: Many people get a different type on retest
- Dichotomous categorization: Forces continuous traits into binary categories
- Lack of predictive validity: Does not reliably predict job performance
The nuance: Typologies can measure real differences, but with lower precision than dimensional models such as Big Five. They work as tools for reflection, not for hiring or clinical decisions.
The Enneagram: wisdom without full validation
Hook et al. (2021) found "mixed evidence for reliability and validity":
- Factor analyses typically find fewer than 9 factors
- Little research supports wings and lines of integration
- A limited number of empirical studies
Conclusion: Useful for self-reflection, but treat Enneagram results as hypotheses, not settled answers.
Big Five: the gold standard has limits too
- Self-report bias: We answer the way we want to come across
- Limited self-insight: Some aspects are easier for others to see
- Mood and context: Answers are affected by how the day is going
- Cultural limits: Developed primarily on WEIRD populations
- Statistical vs. practical: Explains 10-15% of the variance in life outcomes
Every test has weaknesses
Categories vs. continuum
Most people sit in the middle of a dimension, not at its extremes. The advantage of categorizing is that it communicates; the cost is a loss of nuance.
A static vs. a dynamic view
Tests give snapshots. But personality matures over time, is shaped by life events, and varies with context.
Selection bias
Much of the research is based on psychology students at Western universities – a group that is not representative.
- Treat results as hypotheses, not truths
- Pay attention to borderline cases (scores near 50%)
- Ask someone who knows you whether they recognize you in it
- Separate description (tendencies) from explanation (causes)
- Keep context in mind – you are different at work and with friends
The hybrid approach compensates
- Triangulation: Several frameworks pointing the same way increases confidence
- Different lenses: Each framework illuminates different aspects
- Concrete scores: Quantitative results are more specific than vague descriptions
- Cross-mapping: An internal consistency check reveals contradictions
The constructive position
The key is to avoid two pitfalls:
- Naive acceptance: "The test says I am X, so that is what I am"
- Cynical dismissal: "Personality tests are just horoscopes"
The constructive position lies in between: use the tests as tools for exploration, not as settled truths.