Methodology
How SELVET builds on the research
The questions are built on IPIP material and the Schwartz PVQ, answers in one framework update the scores in the others, and follow-up questions are inserted where the uncertainty is greatest. Below is how each of the 540 questions is actually used.
SELVET’s design principles
SELVET represents a new generation of personality testing that combines academic rigor with modern technology. Four core principles:
- Scientific grounding: Questions from validated instruments (IPIP, PVQ)
- Cross-framework mapping: Answers update scores across frameworks
- Adaptive precision: Follow-up questions where they are needed most
- Transparent methodology: You can see how the results are calculated
The item bank
SELVET uses 540 questions in total, across four frameworks:
| Framework | Count | Primary source |
|---|---|---|
| Big Five | 300 | IPIP material (our own item bank, written in Norwegian) |
| Thinking style | 100 | Developed in-house (Jung-based) |
| Motivation | 90 | Developed in-house (Enneagram-inspired) |
| Values | 50 | Schwartz PVQ (extended) |
How your scores are calculated
Big Five: domain and facet level
Scores are calculated across the 5 main domains, and across 30 facets (6 per domain) where the data supports it. Each facet is built from the facet codes on the questions themselves, not from their order in the test, and requires at least 4 answered items. A facet level is shown only once all six facets in a domain are measured — so the full version (300 questions) yields facets, while the short version (50) yields domains only. Reverse-scored questions are handled automatically, and raw scores are normalized to 0-100.
Thinking style: preference strength
For each dimension (E/I, S/N, T/F, J/P) both preference and strength are calculated. 50 = no preference, 100 = strong preference.
Motivation: ranking with a wing
All nine types are scored, the primary type is the highest score, and the wing is determined by comparing the neighboring types.
Values: a ranked hierarchy
Values are presented as a ranked list, centered on your personal average in order to show relative priorities.
SELVET starts with a prior (an initial probability) based on population averages. Each response updates not only its primary dimension, but also correlated dimensions in the other frameworks. The result is a posterior that reflects all available information.
Adaptive testing
When a score sits in the "uncertainty zone" (typically 45-55%), follow-up questions are activated, designed to discriminate sharply between the alternatives.
An example follow-up question for a borderline E/I:
"After a full day of social activity: do you feel mostly (A) energized and content, or (B) tired and ready for time alone?"
You can always choose to skip, but you are then told that the result is less certain in that area.
You can see the uncertainty in your result
SELVET reports not only point estimates, but confidence intervals as well. The Standard Error of Measurement is calculated from the number of questions and the reliability of the scale (Cronbach’s alpha).
Openness to Experience
Score: 78 (confidence interval: 73-83)
Interpretation: You score clearly above average.
A shareable personality code
SELVET generates a compact, shareable code that sums up the profile:
The code contains the thinking style (4 letters), the Enneagram type with its wing, and the two highest-ranked values. Designed to be short enough for social media, intriguing enough to spark curiosity, and comparable with other people’s codes.
AI-generated dashboard
The premium dashboard uses a large language model (LLM) to generate personalized insights. Prompt engineering ensures:
- Specificity: References concrete score combinations
- Contrastive statements: Includes what the profile does not indicate
- Action-oriented: Concrete suggestions to test
Quality assurance includes an automated check that the text refers to actual scores and avoids absolutes.
Continuous improvement
- Test-retest reliability: The same person, two weeks apart
- Convergent validity: Correlation with established instruments
- User feedback: Questions with a low "accurate" rating are flagged for revision
- A/B testing: New wordings are tested against existing ones