Skip to main content

What makes a test valid?


 
What makes a test valid? is a tricky question. 



The short, and rather obnoxious response is, “nothing.” 




Like reliability, validity is a property of test scores
 rather than tests but more accurately, validity is an interpretation
of the scores.


But it is important to take the question seriously when test-takers and users are wondering how much confidence to place in a test score. As with many aspects of science, the answers can be simply stated but there is a complicated backstory.


Validity Traditions


For many, the traditional views of test score validity will be sufficient. Tests measure constructs. Scientific constructs are ideas that have features that can be measured like reading comprehension, dominance, short-term memory, and verbal intelligence.


Construct validity is not a single entity but rather the current state of knowledge about how a test instrument’s scores have functioned in many settings and in relation to criteria. Construct validity primarily includes findings from studies of content validity, convergent validity and discriminant validity.


Content validity is based on judgment analysis from experts who mostly agree that test items measure the construct (e.g., marital satisfaction).

The other types of validity are based on the concept of correlations with a criterion. Researchers ask participants to take a specific test X along with other tests Y and Z. Test X is the test of interest such as a new math achievement test. Test Y represents other similar tests such as other math tests. When test X and test Y yield similar scores we have evidence of convergent validity.


When test X and test Z yield dissimilar results such as a relationship between our test X math achievement and test Z vocabulary, we have evidence of discriminant validity—a math test ought not to measure vocabulary aside from the minimal vocabulary used in the instructions and word problems. The relationship between the tests is based on a statistic called the validity coefficient, which will vary anytime you have a group of people taking two tests—even the very same people will get different scores on two different testing dates.


Criterion validity compares test scores to some criterion. The relationship between depression test scores measuring depression today is called concurrent validity. The relationship between test scores today and some future measurable performance is predictive validity—for example, a pre-employment test may be correlated with supervisor ratings after six months on the job.
Aside from content validity, most traditional studies are looking at the strength of the relationship between one set of test scores and another.


Factor analysis is a complex correlational procedure that examines the underlying relationship among test items and how they relate to other test items. For example, a set of vocabulary items may be correlated with answers to questions about general knowledge and be considered a “verbal factor” when the two sets of items may be grouped as representing an underlying verbal factor. These abstract underlying factors are sometimes called latent variables or latent traits.


Read more about validity of surveys and tests in CREATING SURVEYS- Chapter 18.



Counselors, read more about validity of test scores in APPLIED STATISTICS: CONCEPTS FOR COUNSELORS- Chapter 20.















Related Post




Post Author

Geoffrey W. Sutton, Professor Emeritus of Psychology at Evangel University, holds a master’s degree in counseling and a PhD in psychology from the University of Missouri-Columbia. His postdoctoral work encompassed education and supervision in forensic and neuropsychology and psychopharmacology. As a licensed psychologist, he conducted clinical and neuropsychological evaluations and provided psychotherapy for patients in various settings, including schools, hospitals, and private offices. During his tenure as a professor, Dr. Sutton taught courses on psychotherapy, assessment, and research. He has authored over one hundred publications, including books, book chapters, and articles in peer-reviewed psychology journals. His website is https://suttong.com You can find Dr. Sutton's books on   AMAZON    and  GOOGLE. Many publications are free to download at ResearchGate   and Academia  

 

Comments

Popular posts from this blog

Academic Self-Efficacy Scale ASE

  Overview The  Academic Self-Efficacy Scale is an application of Self-Efficacy Theory   to examine the relationship between self-efficacy and academic performance using 8-items rated on a 7-point scale. The work of Chemers et al. (2001) has been widely cited. Format The 8-items are rated on a 7-point Likert-type scale ranging from 1 = Very Untrue to 7 = Very True. Sample Items 2. I know how to take notes. 6. I usually do very well in school and at academic tasks.   Reliability, Validity, and Other Research notes In the article describing the development and use of the ASE, the authors observed: “As predicted, academic self-efficacy was significantly and directly related to academic expectations and academic performance.” (Chemers et al., 2001, p. 61)   Sutton et al. (2011) reported Cronbach's alpha of .83 in their study of academic self-esteem and personal strengths. ASE was highly positively correlated with ACT scores (.24) and GPA (.39)....

Brief Self-Control Scale (BSCS)

  Assessment name:   Brief Self-Control Scale Scale overview: The Brief Self-Control Scale (BSCS) is a 13-item self-report measure of self-control, which is also called self-discipline and willpower.   Authors: June P. Tangney, Roy F. Baumeister, Angie L. Boone   Response Type: Items are rated on a five-point scale of self-evaluation from 1 = not at all like me to 5 = very much like me. Scale items Example items include “I am good at resisting temptation” and “I have a hard time breaking bad habits” (reverse coded). Psychometric properties In their original publication, Tangney and others (2004) documented high internal consistency and retest values. High scores representing self-control were associated with academic success and better relationships whereas low scores were correlated with personal problems and problem relationships. Availability: The BSCS is widely used and available in many languages. The original 36-item scale and the Brief...

Rotterdam Emotional Intelligence Scale (REIS)

  Scale name: Rotterdam Emotional Intelligence Scale (REIS) Scales overview : The Rotterdam Emotional Intelligence Scale (REIS) is a 28-item measure of emotional intelligence having four subscales. Authors: Pekaar, Keri A., Bakker, Arnold B., van der Linden, Dimitri, & Born, Marise Ph. Response Type: All items are rated on a 5-point Likert type scale ranging from 1 (totally disagree to 5 (totally agree). Subscales = 4   Following are the four subscales with a sample item. Self-focused emotion appraisal( 1. I always know how I feel.) Other-focused emotion appraisal ( 8. I am aware of the emotions of the people around me.) Self-focused emotion regulation ( 15. I am in control of my own emotions.) Other-focused emotion regulation ( 22. I can make someone else feel differently.)   Reliability and Validity “The results indicate that the REIS follows a four-factorial structure and can be reliably measured with 28 items. The REIS was strongly correl...