Sidor

Showing posts with label Faking. Show all posts
Showing posts with label Faking. Show all posts

Friday, December 21, 2012

Norms - a crucial issuse in testing



Test norms need to be specific to the user and the test context. This should be obvious, still is often ignored, perhaps due to the expenses involved. What happens if norms are not specific?

1. A very important aspect is that of faking. Faking is abundant in job applicants. If norms are collected from incumbents or, even worse, the population at large, test scores can be grossly misleading. The reason is that many applicants fake and the distribution of their test scores is shifted towards a higher mean than for incumbents who fake very little or not at all. As a consequence, test scores for applicants will be systematically overestimated. In a stanine scale, the error could easily be 2 or 3 steps. This problem could be greatly mitigated by using a correction procedure using one or several scales for measuring the tendency to respond in a socially desirable manner. In our data, about 95 % of the effect is eliminated this way. Note, however, that the correction model must be scale specific since scales are usually not equally vulnerable to distortion.

2. Test scores may be strongly dependent on the organizational context. In some contexts, independences is not a desired trait and people will on the average have low scores on this trait. Another example is perseverance in the face of failure. If failure is rarely obvious, test takers will report low perseverance. For reasons such as these, norms need to be specific to the organizations.

It is not excessively demanding to construct specific norms, given modern IT technology, and the sample size need to be only as small as 300, or even in some cases 120. The first step is to realize the importance of specific norms, of norms corrected for impression management if they are based on incumbents or the population at large, and the fact that the sample size can be fairly small. In our practice we work with such norms, but many Swedish test providers seem unaware of the issue and that the problems can be solved with relatively modest resources.

Wednesday, August 22, 2012

Successfully dealing with faking on a self-report personality test



Faking on self-report personality tests is common and a strong drawback of such tests. Many approaches have been tried to counteract this serious source of error, see e.g. a recent papers in the Journal of Applied Psychology (Bangerter, Roulin, & König, 2012; Fan, et al., 2012).

The UPP test (Sjöberg, 2010/2012) is a self-report personality test and as such it is vulnerable to faking in high-stakes testing situations. However, this test uses a simple but powerful methodology for correcting test scores for faking. It measures separately two social desirability (SD) dimensions, one overt (similar to the classical Crowne-Marlowe scale (Crowne & Marlowe, 1960)) and one covert. The covert scale uses items similar to conventional personality items but selected for their strong correlation with the overt scale. The two scales are highly correlated and give similar results when used to correct test scales for faking.

The correction procedure uses regression models where each test scale in turn is the dependent variable and the SD scales are independent variables. It is necessary to fit a new model for each test scale because the different scales are related to SD in different ways, correlations varying widely. The corrected test scales are the residuals in these regression models.

This procedure gives corrected test scales which correlate zero with SD. So far, so good, but does it also work? In other words, can it be validated on empirical data? One way to validated it is to study groups tested under different levels of involvement, from incumbents where test results have no consequences, to applicants where they do, and consequences are very important. In a recent study of applicants to the officers' training program in the Swedish Army, I had a chance to study this question, using the UPP test and its SD scales. (Previous studies had given similar results). Data were available for 5 groups:

A. Norm
B. Incumbents
C. Applicants (low consequences of test results)
D. Applicants (moderate consequences)
E. Applicants (high-stakes testing)

I expected increasing SD scale values in the order A - E. I also expected test scales to have the same rank order, if they were sensitive to SD, such as emotional stability. Finally, I expected the group differences in emotional stability to vanish if the test data were corrected for faking using the two SD scales (and a multiple regression model). For the results, see Figs. 1 and 2 below, and Table 1. 


Fig. 1. Means of SD scales


Fig. 2. Means of emotional stability before and after SD correction



Tabell 1. Mean values of emotional stability (standardized scales), uncorrected and corrected data, effect size and one-way ANOVA of group differences.
Grupp
Before correction
Corrected for SD
A. Norm
-0.25
-0.05
B. Incumbents
0.05
0.07
C. Applicants (low consequences of test results)
0.43
0.28
D. Applicants (moderate consequences)
0.56
0.06
E. Applicants (high-stakes testing)
0.73
0.11
Effect size (eta2)
0.147
0.006
One-way ANOVA
F(4,1638) = 70.693, p < 0.0005
F(4,1828) = 2.763, p = 0.026

Note that the effect size decreased to about 5 %.

In other work on leader effectiveness, using 360 degrees feedback as criterion, I found that the validities of the test scales increased after correction for SD according to the same method (Sjöberg, Bergman, Lornudd, & Sandahl, 2011), see Fig. 3. 

Fig. 3. Validities of uncorrected and corrected persnality scales


In conclusion, a simple method for correction for faking has been found to successfully remove about 95 % of the variance due to SD in test responses, and such a method increased the validity of the test scores against an external criterion. 

It is often argued that SD scales really measure "personality", such as need for approval, and not a tendency to distort responses. However, the present results strongly refute this view. It is very plausible that different levels of consequences of testing should lead to different levels of motivation for impression management, but unlikely that they should result in different levels of some personality dimension such as need for approval.

References

Bangerter, A., Roulin, N., & König, C. J. (2012). Personnel selection as a signaling game. [doi:10.1037/a0026078]. Journal of Applied Psychology, 97, 719-738.
Crowne, D. P., & Marlowe, D. (1960). A new scale of social desirability independent of psychopathology. Journal of Consulting and Clinical Psychology, 24, 349-354.
Fan, J., Gao, D., Carroll, S. A., Lopez, F. J., Tian, T. S., & Meng, H. (2012). Testing the efficacy of a new procedure for reducing faking on personality tests within selection contexts. [doi:10.1037/a0026655]. Journal of Applied Psychology, 97, 866-880.
Sjöberg, L. (2010/2012). A third generation personality test (SSE/EFI Working Paper Series in Business Administration No. 2010:3). Stockholm: Stockholm School of Economics.
Sjöberg, L., Bergman, D., Lornudd, C., & Sandahl, C. (2011). Sambandet mellan ett personlighetstest och 360-graders bedömningar av chefer i hälso- och sjukvården. (Relationship between a personality test and 360 degrees judgments of health care managers). Stockholm: Karolinska Institute, Institutionen för lärande, informatik, management och etik (LIME).

Tuesday, August 14, 2012

Validity of integrity tests


Traditionally the view has been that integrity tests (actually honesty tests) have very high validity, based on an early meta-analysis (Ones, Viswesvaran, & Schmidt, 1993). Some skeptical comments have pointed out that many of the studies in this meta-analysis came directly from reports from test vendors. Yet the high validity of integrity tests it has become an established truth, and a basis for an entire industry producing integrity tests, based on Schmidt and Hunter (1998) who wrote that the g-factor + integrity is the best basis for prediction of work performance. This is probably wrong.

A current and updated meta-analysis clearly shows that validities of integrity tests are not higher than 0.2, perhaps as low as 0.1 (Van Iddekinge, Roth, Raymark, & Odle-Dusseau, 2012a, 2012b), even if they are corrected for measurement error in criteria and range restriction in the test. The earlier estimates were at level 0.4, i.e. higher than the standard personality test. It appears now that the skeptics have been right: the high validities come from test providers' own information, independent research does not confirm therm. A rather high value of validity can be obtained with self-ratings of counterproductive behavior at work, but this is not very interesting.

This is an example of how early meta-analysis can result in errors. Van Iddekinge et al. have published a very  ambitious project. The result is clear. Integrity test seems not to have significant practical value. And then we have not even discussed that such tests can easily be faked..

References

One, DS, Viswesvaran, C., & Schmidt, FL (1993). Comprehensive meta-analysis of integrity test validities: findings and implications for personnel selection and theories of job performance. Journal of Applied Psychology Monograph, 78, 679-703.

Schmidt, F. L. & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 124, 262-274.

Van Iddekinge, CH, Roth, PL, Raymark, PH, & Odle-Dusseau, HN (2012a). The criterion-related validity of integrity tests: An updated meta-analysis. [Doi: 10.1037/a0021196]. Journal of Applied Psychology, 97 (3), 499-530.

Van Iddekinge, CH, Roth, PL, Raymark, PH, & Odle-Dusseau, HN (2012b). The critical role of the research question, inclusion criteria the, and transparency in meta-Analyses of integrity test research: A reply to Harris et al. (2012) and Ones, Viswesvaran, and Schmidt (2012). [Doi: 10.1037/a0026551]. Journal of Applied Psychology, 97 (3), 543-549.

Sunday, September 11, 2011

Hur många gånger kan man upprepa samma test?

Personlighetstest som OPQ och HPI används ofta, och har funnits länge på marknaden. Som ett resultat av detta möter man ofta jobbkandidater som redan testats en eller många gånger med ett eller flera av de vanliga testen. Det är då naturligt att fråga sig om testens värde urholkas av omfattande erfarenhet med dem. Såvitt bekant förekommer inga försök att registrera antalet testningar en person gjort med ett test.(Utom för UPP 2.0, som lanseras snart. UPP är ett nytt test, medan det finns hundratusentals personer i Sverige som tagit test som OPQ, HPI, Master, 16 PF, Myers-Briggs och Thomas/PPA).

Frågan om effekter av omtestning med personlighetstest har inte undersökts noga tidigare.  Hausknecht (2010) finner emellertid i en studie med Gordon-testet (ett Big Fivetest), mycket stora effekter, i förskönande riktning, hos dem som inte anställts efter första testningen. (Hos dem som fick jobbet var det inga effekter). Detta var fallet trots att dessa personer inte fick ingående återkoppling, något som troligen skulle ha ökat effekterna ytterligare eftersom den kan ge ledtrådar till den testade om vad i testresultatet som var mindre lyckat.

Särskilt låga värden vid första testningen förbättrades starkt vid nästa testning. Sambandet mellan de två testningarna var svagt, mycket svagare än vad man normalt får med upprepade personlighetstestningar. Det tyder också på att de testade har andra svarsstrategier andra gången, och troligen innebar detta i sin tur att testets validitet sjönk betydligt.

Den som vill använda ett av de vanliga testen bör därför tänka sig för noga. Hur stor är risken att personen som ska testas redan har tagit testet? Har han eller hon fått återkoppling? Detta är ju regel i Sverige. Kan man få reda på om tidigare testning gjorts? Ja, det går ju an att fråga, men många vet inte vilka test de tar eller har tagit, och de har inte fått någon skriftlig rapport med sig från testningen. Vissa testföretag hemlighåller t o m vilka test de använder.

Kan test bli "utslitna" och tappa den validitet de eventuellt kan ha haft? Javisst, om man under decennier har kört tiotusentals testningar per år i Sverige. Det behövs nya test, och systematiska frågor till den testade om tidigare erfarenheter av testning, för att komma tillrätta med detta stora problem.

Referens

Hausknecht, J. P. (2010). Candidate persistence and personality test practice effects: Implications for staffing system management. [doi:10.1111/j.1744-6570.2010.01171.x]. Personnel Psychology, 63(2), 299-324. Klicka här

Friday, April 22, 2011

SIOP 2011 comments: Soldier recruitment

A major project in the US Army involves a new personality test, Tapas. This is a Big Five test with a combination of ipsative and normative formats which should make it less vulnerable to faking, although this has yet to be proven. A very interesting finding was that mental ability, or g, was an important predictor for "can do" criteria, while personality was important for "will do" criteria (around 0.2, probably not corrected for measurement error and restriction of range). The debate on how to weight g and personality must clearly take into account what is to be predicted. It is striking how much more important personality, even constrained to the ineffective Big Five framework, is with regard to "will do" criteria. Just what personailty dimenions are important is a matter of concern. Traditional military psychology, based as it was on WW II experience, said emotional stability, if combat effectriveness was a criterion and studied in real-world applications (war). Current work emphasizes conscientiousness as it is a dominating dimension in civilian and and peacetime applications, and perhaps even peaceful, settings. Will conscientiousness really help in high-stakes and threatening situations?

The Swedish Government has recently decided to create a professional army, where soldiers will get a small but decent salary. (SEK 16 500 per month). So far, the program is hugely popular with some 22 000 applicants for about 2000 openings. This means a selection ratio of 10 % which should make screening testing very feasible and effective. Values, held to be very important and measured in another US Arny project, can be measured by the proxy dimensional of emotional inteligence (EI) (self report).We have a wealth of data showing this. The second Army project uses another new  personality test, GAT, (see earlier blog entry) but it sems to lack a reltionship to Tapas. No such studies were mentioned (and nobody asked). The GOT test is, by the way, kept secret and item formats and content are not disclosed, measring such things a spiritutal value sand justice seems to be a real challenge.

Check out these for the promise of a proxy measure of values:


Engelberg, E., & Sjöberg, L. (2005). Emotional intelligence and interpersonal skills. In R. D. Roberts & R. Schulze (Eds.), International handbook of emotional intelligence (pp. 289-308). Cambridge MA: Hogrefe.
Click here.

Engelberg, E., & Sjöberg, L. (2006). Money attitudes and emotional intelligence. Journal of Applied Social Psychology, 36(8), 2027-2047. Click here.

Engelberg, E., & Sjöberg, L. (2007). Money obsession, social adjustment, and economic risk perception. Journal of Socio-Economics, 36(5), 689-697. Click here.





Alternatively, write an e-mail to get reprints, write to lennartsjoberg@gmail.com

Saturday, April 16, 2011

Comments on SIOP 2011: Faking on personality tests

The issue of faking is alive and well. Several sessions at the 2011 SIOP are devoted to it. Nobody or very few deny that faking occurs and that it can affect the outcome of a test, sometimes severely so. It is also realized that faking greatly hurts the credibility of personality testing. Non-experts test users simply are convinced that the test takers often fake good in a high-stakes situation, such as when they apply for a very desirable job or admittance to a prestige school.

It is clear that faking reduces the validity of personality tests, if left uncorrected. The effect can be very substantial. Meta analyses of the validity of personality tests tend to be based on data from incumbents, since job performance (criterion) data cannot normally be obtained from all applicants, and applicant scores are only correlated about 0.5 with incumbent scores. Hence, faking makes the data used in meta analyses of doubtful relevance to the question of test validity.

The most important of the Big Five factors, conscientiousness, is the one most affected by faking. It is also clear that the group of fakers, while heterogeneous, may contain some people who are risky to hire. Ignoring faking comes with great risks for the test users.

What can be done?A powerful alternative is to measure social desirability (SD) and use and SD scale to correct other scales for faking, to the extent hat they correlate with SD (not all scales do and correlations vary strongly in the typical case).The procedure has been validated both in experimental and field work.

There are a few objections, however.

1. SD scales are said to measure "personality". It is somewhat unclear what this means and why it is an argument. SD scales to have correlates with many other dimensions and they also have a certain amount of consistency over time and situations. So what? They can still measure faking at any given time.

2. There are several SD scales and they do not measure the same thing. The best known scales do have high intercorrelations, however.

3. You cannot detect who is a faker. Well, you can to some extent, albeit not perfectly, but who said that psychometrics ever comes up with perfect solutions?

4. Some people fake bad. This can be detected, but is not a major problem. Few people fake bad in  a high-situations where they have applied for a desirable job.

Some commercial test suppliers and their agents try to solve the problem of faking by denying that it exists. This is not a credible statement. Since the future of personality testing is probably dependent on there being a solution to the faking problem - why not use the solution described here? It works.
Free counter and web stats