Assessment guide
Personality Test Reliability and Validity: What to Look For
A personality description can feel familiar and still leave important questions unanswered. Here is how to read the evidence behind an assessment and decide whether it fits your purpose.
What is the difference between reliability and validity?
Reliability concerns the consistency or precision of assessment results. Validity concerns whether the evidence supports the interpretation and use you want to make of those results. They are related questions, but one does not answer the other.
For example, imagine a questionnaire that repeatedly describes you as decisive. Consistency would be relevant to reliability. Whether that result supports choosing you to manage a particular project is a different question. The role may depend on experience, technical knowledge and other qualities the questionnaire does not assess.
Start by finishing this sentence: I want to use the result to… Reflect on how I communicate? Plan a team conversation? Predict a particular outcome? Your answer determines which evidence matters.
Which kind of consistency was measured?
Internal consistency examines how items within a scale relate to one another. Test-retest research examines results across separate administrations. For an assessment that assigns categories, agreement in the final category is another relevant question. Ask which of these a reported number describes.
Suppose two reports use the word “reliable.” One analyzes responses from a single sitting. The other follows people who took the inventory twice. Those reports provide different information, even if both contain impressive-looking coefficients.
A useful follow-up is: “Does this number tell me about the individual scales, the overall score, or whether people received the same final type?” Keeping the unit of analysis clear prevents a scale statistic from becoming a promise about every participant's result.
What does a TDF reliability finding tell us?
The TDF Pattern Inventory Research Manual reports internal-consistency analyses for its six pattern scales in an initial sample of 861 adult businesspeople and a follow-up sample of 333. It also reports a French-language sample of 129. Source and research context.
The coefficients provide evidence about how the items within each scale worked together in those samples. The manual explicitly states that test-retest reliability had not been directly assessed. We therefore cannot use those findings to promise that someone will receive the same pattern on another occasion.
Likewise, a coefficient of .90 would not mean “90% of people receive the correct personality type.” It is a statistic about consistency, with a meaning that depends on the method used. For the specific findings and their scope, see how TDF was developed and studied.
What would support the use you have in mind?
Comparisons with other assessments can help researchers examine what an inventory describes. They do not automatically establish that the inventory predicts success at work, improves relationships or produces better team decisions.
Consider the TDF and MBTI comparison. It records how two sets of classifications overlap. That is useful when asking whether the results are interchangeable. It does not measure whether a workshop using either instrument improves a team's performance.
For a team workshop, separate the evidence about the assessment from the evidence about the workshop. You can also set a practical follow-up: agree on a change in how the team runs meetings, try it, and review what happened. That review helps you judge the local usefulness of the activity. It is not a validation study of the inventory.
Be equally specific about the people studied. Findings from one language, occupation or setting may leave questions about another. Ask how the sample relates to the people who will use the assessment.
Five questions to ask before choosing an assessment
- What exactly is being described? Ask for a plain-language definition, with an example of what the result does and does not tell you.
- Which assessment version was studied? Check the language, response format and scoring approach as well as the name.
- Who took part? Look for a sample size and a description of the participants.
- What outcome was measured? Item consistency, repeatability, relationships with other inventories and practical outcomes answer different questions.
- Can I read the source? A useful citation identifies the report, authors and relevant section, with enough context to understand the claim.
Write the answers beside your intended use. If you still cannot connect a finding to that use, ask for clarification before treating the finding as a reason to choose the assessment.
What if a result simply feels right?
Recognition can be a productive starting point. Make it specific: which sentence describes something you actually do, and when? Then look for an example that does not fit. Both deserve a place in your interpretation.
A participant might find that asking for the wider context makes a recurring disagreement easier to discuss. That is a useful personal observation. It does not establish a general success rate for the assessment or prove every part of the description.
Use the result to ask a better question, then pay attention to the answer. The TDF results guide offers a short exercise for doing this without forcing your experience to match a label.
Sources and research context
General assessment concepts: American Educational Research Association, American Psychological Association and National Council on Measurement in Education, Standards for Educational and Psychological Testing, 2014, chapters 1 and 2. The distinctions between evidence for an interpretation, consistency and intended use inform this guide. The project and workshop examples are illustrations.
TDF findings: Boyd Spencer, Bill Roberts and David Farr, TDF Pattern Inventory Research Manual, second edition, copyright 2002, pp. 5–6. The English-language samples were studied in 1984 and 1989; the French-language sample in 1989. The follow-up English study used a different response format, discussed on p. 8. These are historical findings for the studied versions and samples, not a new validation of today's online inventory.
The manual is a TDF publication. Its findings should be considered with that source context. The research reference page identifies additional limitations and table issues.