sci.psychology.personality
The science of personality.
The sci.* counterpart to the alt personality groups: trait models, psychometrics, twin studies and construct-validity argument held to citation standards — with researchers occasionally posting under their own names.
Big Five consolidation debates of the 1990s played out here in miniature.
On this page
- The paperwork, and the date on it
- The charter, and the instruments it named
- The lexical hypothesis: counting the words
- Five factors, and how they were arrived at
- The critics, the alternatives, and the person-situation argument
- Reliability, and its several meanings
- Validity, and why construct validity was the hard one
- Factor analysis and the rotation arguments
- Norms, and what a standardisation sample is
- Behaviour genetics: the designs and the published findings
- Behaviour genetics: the standing criticisms
- The instruments people argued about
- A scientific room on a public network
- After the gateway: the replication crisis
- The group's own record, and what is not established
- Scope and limits
The paperwork, and the date on it
sci.psychology.personality has an exact birthday, and the document that records it survives. The newgroup control message that created it was issued by David C. Lawrence from uunet.uu.net and is dated Monday 12 June 1995, 14:30:52 GMT. Its body runs to fifteen lines: an authorising sentence, a line for the newsgroups file, and the charter. The authorising sentence states the arithmetic, as such messages did, recording that the group passed its vote for creation by 122:19 as reported in news.announce.newgroups on 6 June 1995.
Behind that single message sit three months of paperwork, all of it preserved in the Usenet administrative archive. On 14 March 1995 John M. Grohol, writing from an address at Nova Southeastern University in Florida, posted a Request for Discussion proposing not one group but a reorganisation of the whole sci.psychology namespace: ten renamings and creations on a single ballot. The stated grievance was the ordinary grievance of a namespace that has outgrown its own name. Traffic had risen steadily; sci.psychology was, the proposal noted, one of the older science newsgroups and therefore had no charter specifying appropriate content; and it had become, in the proposers' words, a dumping ground for current events that may or may not have any relevance to the field of psychology. Kill files, the document added drily, were no longer useful in a group that busy.
Two Calls for Votes followed, on 12 and 23 May 1995, conducted by Michael Handler on behalf of the Usenet Volunteer Votetakers, the standing pool of neutrals who ran Big Eight ballots in that era. Votes had to reach the votetaker by 23:59:59 UTC on 2 June. The result was posted on 5 June, late enough in the United States that the control message dated the report to the sixth: 171 valid votes cast and one invalid. Eight of the ten proposals passed. The two that failed — a group for unrefereed journals and a marketplace group — failed not on the two-thirds test, which both cleared, but on the separate requirement of at least a hundred more yes votes than no. sci.psychology.personality passed 122 to 19, and the result posting records the minute at which it crossed the threshold: Tuesday 30 May 1995 at 22:10:00.
The reorganisation itself — its politics, its moderated groups, and the wider history of psychology on Usenet — belongs to the alt.psychology page in this directory and is not retold here. The later governance of the branch, including the attempt in 1996 to put a moderator on a sibling group, belongs to sci.psychology.psychotherapy. What this page owns is one line of that reorganisation's output, and the subject that line named.
The line is still in service. The newsgroups file distributed by the Internet Systems Consortium — the descriptions list a news administrator installs so that a reader sees more than a bare name — carries the group today exactly as the 1995 control message asked it to:
sci.psychology.personality All personality systems & measurement.
Six groups of the sci.psychology branch remain in that file: announce, misc, personality, psychotherapy, research and theory. The moderated consciousness group and the two moderated groups that carried electronic journals are no longer there. The group described on this page outlasted them in the file. The archive consulted for this article records the listings; it does not record why the branch thinned out in the way it did.
The charter, and the instruments it named
Big Eight groups came with a charter, and the charter was a public document that anybody in an argument could quote. This one was reproduced verbatim in the control message, culled — the message says — from the call for votes:
This newsgroup is for the discussion of all personality systems and measurement in psychology. Examples of appropriate topics for this newsgroup might be discussion of the psychodynamic conceptualization of personality, Dollard & Miller, the PAI-2, the MBTI, the MMPI-2, etc. This newsgroup would encourage discussion not only on personality systems, but also on personality disorders, as well as healthy personality systems.
It is a revealing paragraph, mostly for what is missing from it. The five-factor model, which by 1995 had been the organising controversy of the field for a decade and a half, is not mentioned. What is mentioned first is the psychodynamic conceptualisation of personality; what is mentioned second is Dollard and Miller, meaning John Dollard and Neal E. Miller of Yale's Institute of Human Relations, whose Personality and Psychotherapy: An Analysis in Terms of Learning, Thinking, and Culture had appeared from McGraw-Hill in 1950 and set out a learning-theory account of personality and neurosis. A charter drafted in 1995 that reaches for a forty-five-year-old book and passes over the live argument in silence is a charter written to be inclusive rather than to take sides, which for an unmoderated group is a sensible way to draft.
Then three instruments, named without comment: the PAI-2, the MBTI and the MMPI-2. The middle one, the Myers-Briggs Type Indicator, belongs to another page in this directory and is dealt with in one sentence below. The last, the second edition of the Minnesota Multiphasic Personality Inventory, had been published six years earlier and was among the most widely used instruments of its kind. The first is a small puzzle. The Personality Assessment Inventory was published by Leslie Morey in 1991, and a revised edition did not follow until 2007. What the drafters meant by the suffix in 1995 is not recorded, and the archive offers no clarification. It is the sort of small slip that survives in a founding document because nobody ever revises a Usenet charter.
The rationale attached to the proposal is more pointed than the charter, and it explains what kind of room the proposers thought they were building. The new group was to move scientific discussion of personality systems out of the general newsgroup, sci.psychology, and also out of what the document calls the lay-person newsgroup, alt.psychology.personality; and although the proposal expressly did not seek to remove that older group, it suggested that the new one would supersede it. The split between the popular and the scientific ends of the subject was therefore written into the paperwork before a single article was posted. This directory follows the same division: the popular end, where type systems, enthusiasm and folk taxonomy lived, belongs to alt.psychology.personality.
One word in the one-line description does the real work of distinguishing this group from its alt.* neighbour, and it is the last one. Measurement. That is what the rest of this article is largely about.
The lexical hypothesis: counting the words
The trait tradition that dominated personality research in the group's years rests on an argument about vocabulary, generally called the lexical hypothesis. Stated in its usual two parts, it holds that personality characteristics important to a group of people will end up encoded in that group's language, and that the more important a characteristic is, the more likely it is to be encoded as a single word. If both are true, then a dictionary is a record of what people have found worth noticing about one another, and the structure of the vocabulary is evidence about the structure of the thing described.
Francis Galton is generally credited as the first to propose sampling language for this purpose, in an 1884 paper called Measurement of Character. The counting began slowly. George E. Partridge listed roughly 750 English adjectives for mental states in 1910; M. L. Perkins estimated some 3,000 such terms in Webster's New International Dictionary in 1926; Ludwig Klages put the German stock at about 4,000 in 1929. In 1933 Franziska Baumgarten of the University of Bern published the first psycholexical classification proper, identifying 1,093 German terms from dictionaries and the characterology literature.
The study that mattered came three years later. In 1936 Gordon Allport of Harvard and Henry Odbert of Dartmouth worked through Webster's New International Dictionary — roughly 400,000 words — and extracted 17,953 terms used to describe personality or behaviour. They then sorted these into columns, of which the first, containing what they took to be observable and relatively permanent traits, held 4,504 adjectives. That figure of 4,504 is the raw material from which much of the subsequent trait taxonomy was cut, and the exercise is a fair illustration of what psychometric work involved before computers: a dictionary, three anonymous judges whose classifications agreed with one another only about forty-seven per cent of the time, and a great deal of hand sorting. Faced with that disagreement the authors published Odbert's classification and described their own result as arbitrary and unfinished.
The list was reworked repeatedly. Warren Norman, dissatisfied with the inherited categories, went back to Allport and Odbert's source, added terms from the 1961 edition of the dictionary, removed those that had fallen out of use, and worked a pool of roughly 40,000 candidate terms down to 2,797 by stripping out the archaic, the purely evaluative, the obscure, the dialect-bound, the merely physical and the only loosely personal. Each of those exclusions is a judgement, and each judgement shapes what the eventual factor analysis can possibly find. Critics of the lexical approach made exactly that point. The published objections included the charge that using verbal descriptors imports the sociability bias of ordinary language and a negativity bias in emotional terms, tilting the resulting dimensions before any statistics are run; that lay usage of these words is ambiguous; that the terms were not devised by experts for the purpose; that the mechanisms which put personality words into a language in the first place are not well understood; and that restricting the material to adjectives excludes ways of describing people that need a phrase or a paragraph.
Five factors, and how they were arrived at
The five-factor model was not discovered once. It was arrived at several times, by different people using different data and different methods, and the claim that these several arrivals were arrivals at the same place is itself an empirical proposition that had to be argued.
Raymond Cattell took Allport and Odbert's list in 1943 and cut it to about 160 terms by eliminating near-synonyms, added terms from other psychological categories, and arrived at 171. Factor analysis of these yielded some sixty clusters and a further handful of minor ones; the reduced set of about thirty-five terms that came out of this line of work is what later researchers took up. Working from that reduced set, and with Maurice Tatsuoka and Herbert Eber, he published the Sixteen Personality Factor Questionnaire in 1949. In July of that same year Donald Fiske of the University of Chicago, using 22 terms taken or adapted from Cattell's work, extracted five factors, which he named Social Adaptability, Emotional Control, Conformity, Inquiring Intellect and Confident Self-expression.
The version that stuck came out of the United States Air Force. In 1957 Ernest Tupes and Raymond Christal, research psychologists at Lackland Air Force Base, studied peer ratings of Air Force officers made against Cattell's terms; from 1958 they combined this with data from Cattell's and Fiske's earlier work and from a 1953 meta-analysis by John W. French of the Educational Testing Service. Their analysis produced five factors, which they labelled Surgency, Agreeableness, Dependability, Emotional Stability and Culture. Warren Norman replicated the structure in 1963 and renamed two of them, Surgency becoming extraversion or surgency and Dependability becoming conscientiousness. The names in current use are largely his.

Consolidation followed at a distance. At a 1980 symposium in Honolulu, Lewis Goldberg, Naomi Takemoto-Chock, Andrew Comrey and John M. Digman reviewed the personality instruments then available; in 1981 Digman and Takemoto-Chock reanalysed data from Cattell, Tupes, Norman, Fiske and Digman and reaffirmed a five-factor structure, with weak evidence for a sixth. In the same year Goldberg coined the term Big Five for the factors, and Digman's 1990 review pressed the case further. Meanwhile Paul Costa and Robert McCrae, at the National Institutes of Health, had been building an instrument from a different direction. Their 1978 book chapter described a three-factor Neuroticism-Extraversion-Openness model; the NEO Personality Inventory of 1985 measured those three; and in 1992 the revised NEO PI-R added agreeableness and conscientiousness, giving a commercially published instrument with five domain scores and six facets under each. Costa and McCrae took Eysenck's conception of extraversion rather than Jung's, a genealogy worth noting because the two are frequently conflated.
One footnote to the story is a claim of priority. When the fourth edition of the 16PF appeared in 1968, five global factors were derived from its sixteen — extraversion, independence, anxiety, self-control and tough-mindedness — and advocates of that instrument have since described these as the original big five. Whether two factor solutions obtained from different item pools by different rotations are the same factors is precisely the kind of question that made the field's arguments so protracted, and it is not a question that a correlation coefficient settles by itself.
The critics, the alternatives, and the person-situation argument
The group opened in June 1995. Three months earlier, in the March issue of Psychological Bulletin, one of the most substantial published attacks on the five-factor consensus of that decade had appeared, along with its rebuttals, in the same volume. Jack Block of the University of California, Berkeley, published A contrarian view of the five-factor approach to personality description at pages 187 to 215; Costa and McCrae replied under the title Solid ground in the wetlands of personality; Goldberg and Saucier replied separately; and Block's rejoinder, Going beyond the five factors given, closed the exchange at pages 226 to 229. Block's central methodological complaint was that treating factor analysis as the sole route to conceptualising personality is too narrow a paradigm, and he pointed to work specifying aspects of character that the five factors do not subsume. Anyone who wanted a reading list for an argument in this newsgroup in its first month had one, in one issue of one journal.
Hans Eysenck had made his objection three years earlier, in a 1992 paper in Personality and Individual Differences whose title left little to interpretation: Four ways five factors are not basic. Eysenck's own model had three dimensions rather than five — extraversion, neuroticism and psychoticism, with a lie or social-desirability scale alongside them — and he grounded them in a biological theory of temperament, with extraversion tied to variability in cortical arousal and neuroticism to individual differences in the limbic system. He had introduced extraversion and neuroticism as the two most important dimensions in Dimensions of Personality in 1947, adding the third later; a revised version of the Eysenck Personality Questionnaire, the EPQ-R, was described in 1985. Part of his objection to the five-factor model was that there is no universally agreed basis for choosing between factor solutions with different numbers of factors, which is a criticism of the method rather than of anybody's arithmetic.

A different order of objection came from Dan McAdams, whose 1995 paper in the Journal of Personality, What do we know when we know a person?, called the Big Five a psychology of the stranger: a set of traits that are relatively easy to read off someone you have just met, and which by construction leave out what is private, contextual or narrative about a life. The related and much-repeated complaint was that the model is not theory-driven at all, being a statistical account of which descriptors tend to cluster together, and that a taxonomy is not an explanation.
Older and larger than any of these was the person-situation debate, which had been running since Walter Mischel published Personality and Assessment in 1968. Reviewing the literature, Mischel reported that correlations between a trait measure and behaviour, or between behaviour in one situation and behaviour in another, rarely exceeded about .30 to .40, and argued that broad traits therefore explained much less than the field supposed. The trait side replied along several lines: that aggregated behaviour over time is far more predictable than any single act; that a correlation of .40 is not small; and, in work by David Funder and Daniel Ozer, that when the effects of situational variables reported by social psychologists are converted into the same metric, they fall into much the same .30 to .40 range, including in celebrated obedience research. The usual summary of where the argument came to rest is that both person and situation contribute, situations predicting behaviour in a particular moment and traits predicting patterns over time; William Fleeson and Erik Noftle later proposed a formal synthesis in which an individual has a characteristic mean level of a trait around which behaviour varies by situation. By 2009 personality and social psychologists generally agreed that both personal and situational variables are needed to account for behaviour.
Two structural alternatives should be recorded alongside the critiques. Cattell's sixteen factors never went away, and their advocates continued to argue that five domains discard information their instrument preserves. And in 2004, after the group's active life, Kibeom Lee and Michael Ashton published the HEXACO model, derived from lexical studies in several European and Asian languages, which retains recognisable versions of the five and adds a sixth, honesty-humility.
Reliability, and its several meanings
The group's charter named measurement, and measurement is where its subject matter differed most sharply from the alt.* personality rooms. The concepts below are set out here as concepts, as they stood in the literature of the period; nothing in this section is guidance about any instrument or any score.
Classical test theory treats an observed score as the sum of a true score and an error, and reliability as the proportion of observed variance attributable to the true score. The trouble is that reliability is not one question but at least four, and the four can give different answers about the same instrument. Stability over time is measured by re-testing the same people after an interval. Equivalence is measured by comparing two forms of the same instrument. Internal consistency asks whether the items within a single administration behave as though they are measuring something in common. Agreement between judges is a fourth, relevant wherever a human being scores a response. Each of these can be high while another is low, and a bare statement that an instrument is reliable does not say which was measured.
The statistical machinery accumulated over half a century. Charles Spearman, working in the first decade of the twentieth century, supplied both the earliest common factor analysis and the correction for attenuation, the formula that estimates what a correlation between two variables would be if both were measured without error. Split-half methods, with the Spearman-Brown adjustment for the shortening of the test, followed. In 1937 G. F. Kuder and M. W. Richardson published The theory of the estimation of test reliability, giving the formulas known thereafter as KR-20 and KR-21 for dichotomously scored items. Louis Guttman set out a related family of lower bounds in Psychometrika in 1945.
Then, in 1951, Lee Cronbach published coefficient alpha, and it swallowed the field. Alpha generalised the Kuder-Richardson approach to items with more than two response categories and was easy to compute and easy to report; by the 1990s it was near-universal in personality research, and a paper reporting a new scale without an alpha was unusual. Cronbach himself was unsentimental about the reason. In 1978 he remarked that his 1951 paper was so widely cited mostly because he had put a brand name on a commonplace coefficient. Melvin Novick and Charles Lewis had shown in 1967 that alpha equals reliability only when the parts of the test are essentially tau-equivalent, and the persistent misreading of alpha as an index of homogeneity or unidimensionality — a high alpha being taken to show that the items measure one thing — was already being corrected in the textbooks during the group's lifetime, and is still being corrected now.
Two lines of work went further. Generalizability theory, introduced by Cronbach with N. Rajaratnam and Goldine Gleser in 1963 and given book-length treatment with Harinder Nanda in 1972, abandoned the single reliability coefficient in favour of decomposing error into named facets — persons, items, occasions, raters, settings — and asking how much each contributes, so that a study can be designed to control the ones that matter. Item response theory, pioneered independently by Frederic Lord at the Educational Testing Service, the Danish mathematician Georg Rasch and the sociologist Paul Lazarsfeld, modelled the probability of a particular response as a function of the respondent's position on a latent dimension and the properties of the item. Its uptake was gated by arithmetic: it did not come into wide use until the late 1970s and 1980s, when personal computers put the necessary computation within reach of an ordinary researcher, which is a recurring theme in this history.
Validity, and why construct validity was the hard one
Reliability asks whether an instrument gives a stable answer. Validity asks whether the answer means anything, and it was the harder question by a long way.
Through the 1940s the literature had accumulated a clutter of validities — intrinsic, face, logical, empirical and others — with no agreement about which were distinct and which were useful. Between 1950 and 1954 the American Psychological Association's Committee on Psychological Tests worked on the problem, producing the Technical Recommendations for Psychological Tests and Diagnostic Techniques, issued as a preliminary proposal in 1952 and in full in 1954. Out of that work came the 1955 paper in Psychological Bulletin by Lee Cronbach and Paul Meehl on construct validity, which became the standard reference for this part of the subject.
Cronbach and Meehl's argument was that where a test is supposed to measure something not directly observable — anxiety, extraversion, conscientiousness — there is no criterion against which it can simply be checked, because the criterion would need validating too. What can be done instead is to specify a nomological network: a set of claims about how the construct relates to other constructs and to observable behaviour, from which predictions follow. Evidence for construct validity is then the accumulation of results fitting that network, and a failure can indict the test, the theory, or both, with no automatic way of telling which. Validation on this account is not a study but a programme, and never finishes. Meehl's later formulation of what makes a construct worth having — that the best is the one supporting the greatest number of inferences, most directly — states the pragmatic version of the same point.
In 1959 Donald Campbell and Donald Fiske added a practical tool that has been in use ever since, the multitrait-multimethod matrix. The design is simple to describe and awkward to satisfy: measure several traits by several methods, and inspect the resulting correlations. Measures of the same trait by different methods should agree, which is convergent validity; measures of different traits by the same method should not agree merely because they share a method, which is discriminant validity. A great deal of personality research in the following decades consisted of self-reports correlated with other self-reports, and the Campbell and Fiske design is the standing reproach to that practice: shared method variance can manufacture a nomological network out of nothing but questionnaire format.
The later movement in validity theory was towards unification. Samuel Messick argued that construct validity was not one type of validity among several but the whole of it, an integrated evaluative judgement about how far evidence and theory support particular inferences and actions based on test scores — a formulation that folds the consequences of testing into the validity question itself. That view had become the official one by the time this newsgroup was running. The Standards for Educational and Psychological Testing, published jointly by the American Educational Research Association, the American Psychological Association and the National Council on Measurement in Education, had grown from a forty-page pamphlet in 1966 into the 1985 edition, the first to carry the modern title, and were revised again in 1999 and 2014. The 1985 edition was the governing document for most of the period discussed here.
One consequence of that history deserves a plain statement, because it constrains what a page like this one can responsibly say. Under the unified view, validity is not a property an instrument possesses; it is a property of an inference drawn from a score, in a stated population, for a stated purpose. It follows that no general remark about an instrument amounts to a statement that it is fit for any particular use, and none is offered here.
Factor analysis and the rotation arguments
Nearly every structural claim in this literature was produced by factor analysis, and most of the disagreements about the claims were, on inspection, disagreements about the method.
Charles Spearman was the first psychologist to work with common factor analysis, in a 1904 paper in the American Journal of Psychology. Finding that schoolchildren's scores across unrelated subjects were positively correlated, he postulated a single general ability underlying them all. Multiple-factor analysis proper came from Louis Leon Thurstone in two papers of the early 1930s, summarised in his 1935 book The Vectors of Mind, which introduced communality, uniqueness and rotation, and which argued for what Thurstone called simple structure: a solution in which each variable loads highly on one factor and near zero on the rest.
Simple structure was needed because of an awkward mathematical fact. A factor solution is determined only up to rotation. The same data, fitted equally well, can be expressed by infinitely many sets of axes, and the loadings — and therefore the names the factors get, and therefore the theory — depend on which set is chosen. Rotation criteria are attempts to make that choice non-arbitrary rather than to avoid making it. Henry Felix Kaiser's varimax criterion, published in 1958, maximises the variance of the squared loadings and pulls an orthogonal solution towards simple structure. Orthogonal rotations keep the factors uncorrelated, which is tidy; oblique rotations let them correlate, which is often more plausible for personality traits but produces two matrices to read rather than one, and correlated factors invite the extraction of higher-order factors above them.
The second decision is how many factors to retain, and the era had three standard answers, in ascending order of respectability. The Kaiser criterion retains factors with eigenvalues above one, on the reasoning that an eigenvalue of one is the information in an average single item; it is known to over-extract, and its lasting influence owes a great deal to its having been the default in SPSS and most other statistical packages. Raymond Cattell's scree test, introduced in 1966, plots the eigenvalues in descending order and retains the factors above the point where the curve flattens into rubble, which requires a judgement of the eye. John L. Horn's parallel analysis, published in Psychometrika in 1965, was framed explicitly as a correction to the Kaiser rule: it compares the eigenvalues from the real data with those from random data of the same size, on the argument that sampling error inflates sample eigenvalues and causes the Kaiser rule to keep too many factors.

Three consequences follow, and all three were live in the personality literature of the 1990s. First, the number of factors is a decision, defensible or not, and different defensible decisions yield four, five, six or sixteen. Second, factor labels are interpretations imposed on columns of numbers, and a label such as conscientiousness carries connotations the arithmetic does not license. Third, exploratory and confirmatory factor analysis answer different questions: the first asks what structure the data suggest, the second asks how badly a stated structure fits, and evidence of the first kind is weaker than it looks when reported as though it were of the second. Eysenck's complaint of 1992, that there is no universally recognised basis for choosing among solutions with different numbers of factors, was directed at exactly this point, and it was not answered by producing another five-factor solution.
Norms, and what a standardisation sample is
Norming is the least glamorous part of psychometrics and the part where most of the practical trouble lives. A raw score on a personality inventory — so many items endorsed in a keyed direction — means nothing on its own. It acquires meaning only by comparison with a distribution of scores obtained from a defined sample of people, which is what a norm is.
The standard transformations are straightforward arithmetic. A z-score expresses a raw score as the number of standard deviations it lies above or below the mean of the reference distribution. A T-score is the same quantity rescaled to a mean of 50 and a standard deviation of 10, a convention adopted to avoid negative numbers and decimal points. Percentile ranks state the proportion of the reference sample scoring lower. None of these tell you anything about a person that the reference sample does not supply; they are all statements about position within a group.
Which is why the composition of the sample is the whole question. Two properties matter and are routinely confused. A sample can be large and unrepresentative, which buys precision about the wrong population. And norms age: the reference group is drawn at a moment, and the population it stands for changes underneath it. The clearest worked example from this period is the MMPI itself. The original instrument was published by the University of Minnesota Press in 1943, and its comparison group was drawn locally, in Minnesota, around the time the instrument was being built, rather than from any national sample. The restandardisation published in 1989 as the MMPI-2 was undertaken, in the language of the project, to develop a new set of normative data representing current population characteristics, and it enlarged the normative database considerably; the psychometric properties of the clinical scales themselves were left for later work, which arrived in 2003 with the Restructured Clinical scales and in 2008 with the MMPI-2 Restructured Form. A parallel adolescent version, the MMPI-A, appeared in 1992.
The Standards of 1985 and 1999 addressed all of this in the register of a professional code: describe the sample, describe the conditions of administration, say what population the norms are meant to represent, and do not let a test's documentation outlive its evidence. The general point, which is what belongs on this page, is that a norm is a description of a sample of people at a time and place, and that every number derived from it inherits those limits.
Behaviour genetics: the designs and the published findings
Twin and adoption research was central to personality psychology in the group's years, and it is also among the most politically contested areas in the discipline's history. What follows sets out the designs, reports what was published and by whom, and states the standing criticisms of the designs. It reaches no conclusion, and none should be read into it.
The classical twin design exploits a difference in relatedness. Monozygotic twins develop from a single fertilised egg and share their segregating genome; dizygotic twins develop from two and share, on average, half of it, the same as ordinary siblings. If a trait is more similar within monozygotic pairs than within dizygotic pairs, the design attributes the difference to genetic influence — on the stated assumption that the two kinds of pairs experience equally similar environments. The variance in a trait is then conventionally decomposed into three components: additive genetic influence, shared or common environment, and non-shared environment, the last of which also absorbs measurement error. Adoption designs approach the same question from the other side, comparing children with their biological and their rearing relatives. Twins reared apart, rare and correspondingly prized, combine the two.
The named studies of the period were substantial. John Loehlin and R. C. Nichols published an analysis of a large sample of twins identified through the National Merit Scholarship testing programme in 1976. Loehlin, with Joseph Horn and Lee Willerman, ran the Texas Adoption Project at the University of Texas at Austin. At the University of Minnesota, Thomas J. Bouchard Jr. directed the Minnesota Study of Twins Reared Apart, which assembled and assessed separated twin pairs over many years; the two most cited papers from that programme are Auke Tellegen, David T. Lykken, Bouchard, K. J. Wilcox, Nancy L. Segal and S. Rich, Personality similarity in twins reared apart and together, in the Journal of Personality and Social Psychology in June 1988, and Bouchard, Lykken, Matt McGue, Segal and Tellegen, Sources of human psychological differences: the Minnesota Study of Twins Reared Apart, in Science in October 1990. The Minnesota group also ran a separate twin registry alongside the reared-apart study.

Two findings from this literature became general property, and both are worth stating in the form their authors gave them rather than the form they acquired in circulation. The first is the magnitude of the similarity between separated identical twins, which Bouchard himself was at pains to characterise carefully: identical twins raised separately are, he said, about fifty per cent similar on average, which in his words defeats the widespread belief that identical twins are carbon copies, each remaining a unique individual. The second came from Robert Plomin and Denise Daniels, whose 1987 paper asked why children in the same family are so different from one another, and summarised evidence that the environmental influences which matter psychologically are the ones that make siblings different rather than the ones they share. That result cut against a good deal of received wisdom about family influence, and it was as unwelcome to some readers as the heritability estimates were to others.
Behaviour genetics: the standing criticisms
The criticisms of these designs were published contemporaneously with the findings and were of several kinds.
The first is aimed at the equal environments assumption. The twin design attributes the excess similarity of identical over fraternal pairs to genes, which requires that the two kinds of pair be treated no more similarly by their environments than each other. Critics argued that this is false, and that heritability estimates are inflated in consequence. A more moderate position, also on the record, concedes that the assumption is usually inaccurate but holds that the inaccuracy has only a modest effect on the estimates. Defenders pointed to studies of pairs whose parents were mistaken about their zygosity, reporting that children believed to be fraternal but in fact identical remained as concordant as pairs known to be identical.
The second is aimed at the representativeness of separated twins in particular. Pairs separated in infancy are separated by adoption, which makes their families of origin unlike ordinary twin families and their adoptive families unlike them in a different way, adoptive placements being screened and disproportionately childless. Volunteering for a study adds a further filter.
The third is statistical. Peter Schönemann attacked the heritability estimation methods developed in the 1970s and argued that what a twin study estimates need not reflect shared genes at all. His best-known demonstration was a reductio: applying the statistical models published by Loehlin and Nichols in 1976 to trivial questionnaire items yielded narrow heritabilities of .92 for men and .21 for women for having had one's back rubbed, and of 130 per cent for men and 103 per cent for women for wearing sunglasses after dark. A heritability above one is not a finding but a symptom. The response from within the field was that the approximate methods of the pre-computer era, adopted for tractability, had been abandoned since the 1980s in favour of structural equation modelling, under which such estimates are not obtainable. Broader critiques continued to be published long afterwards; Burt and Simons argued in 2014 that conclusions reached by the twin method are ambiguous or meaningless.
The fourth is not methodological. Twin research on psychological traits carries a long institutional history, and it was contested partly on that ground. The most-cited episode is the Burt affair: Sir Cyril Burt, who held the chair at University College London after Charles Spearman and whose students included both Raymond Cattell and Hans Eysenck, published twin data on the inheritance of intelligence that were discredited after his death in 1971. Leon Kamin, whose The Science and Politics of IQ appeared in 1974, noticed that Burt's correlation coefficients for identical and fraternal twins remained identical to three decimal places across papers even as the reported sample grew; Oliver Gillie, then medical correspondent of The Sunday Times, published in 1976 what is generally recorded as the first public accusation of fraud against Burt, reporting his own unsuccessful attempts to trace two women Burt had named as research collaborators; and Burt's own official biographer, Leslie Hearnshaw, concluded on examining the criticisms that most of Burt's post-war data were unreliable or fraudulent. Earl B. Hunt later argued that the lasting damage was less to the specific findings, which other studies broadly paralleled, than to the standing of the whole field and its funding. Funding was itself an issue: partial support for the Minnesota work came from a grant from the Pioneer Fund, an organisation whose history was raised repeatedly by critics of the research.
All of the above was in print, and in ordinary library circulation, while this newsgroup was running. What individual posters made of it is not recoverable from the surviving record, and no attempt is made here to guess.
The instruments people argued about
The charter named instruments, and instruments are what a group about measurement will argue about. The paragraphs below report what the research literature concluded and who concluded it. They are not descriptions of what any instrument is for, and nothing here should be read as saying that any of them is suitable for any use.
The Minnesota Multiphasic Personality Inventory was the landmark instrument of the period. Developed in the late 1930s and early 1940s by Starke Hathaway, a psychologist, and J. C. McKinley, a neurologist, and published in 1943, it was built by empirical criterion keying: items were selected for the clinical scales because groups with a given diagnosis endorsed them differently, not because a theory said they should. That made the instrument conspicuously atheoretical for its time, which its authors counted as a virtue and its critics as an evasion, and it produced scales whose content is heterogeneous by design. The subsequent history is a long argument with its own construction — the 1989 restandardisation, the Restructured Clinical scales of 2003, which their developers described as removing a non-specific demoralisation component held to impair discriminant validity, and the Restructured Form of 2008.
Alongside it stood the questionnaire tradition proper: Cattell's 16PF from 1949, Eysenck's questionnaires, the Personality Assessment Inventory published by Leslie Morey in 1991, and the NEO PI-R of 1992. These were commercial instruments, sold with manuals and restricted by their publishers, which had a specific effect on a public newsgroup: the items themselves were copyrighted and circulated under restriction, so arguments about them proceeded at one remove, from published psychometric properties rather than from the material. Lewis Goldberg's answer to that constraint was the International Personality Item Pool, assembled during the 1990s as a public-domain set of items keyed to the major factor models and free for anyone to use; he and colleagues set out the rationale in the Journal of Research in Personality in 2006, under a title arguing for the future of public-domain personality measures. For an online community, a public-domain item pool was the difference between discussing a literature and being able to run something.
The projective techniques were among the era's most contested instruments. Hermann Rorschach published Psychodiagnostik in 1921, after studying 300 patients and 100 controls and selecting ten inkblots from several hundred he had drawn; he died the following year, and the scoring systems that followed were other people's, principally those of Samuel Beck, Bruno Klopfer and, from the 1970s, John Exner, whose Comprehensive System attempted to put the scoring on a more statistically rigorous footing. The Thematic Apperception Test asked respondents to construct narratives around ambiguous scenes. Both attracted sustained psychometric criticism throughout the second half of the century, and the criticism was consolidated in 2000, when Scott Lilienfeld, James Wood and Howard Garb published a monograph-length review, The Scientific Status of Projective Techniques, in Psychological Science in the Public Interest — a journal whose stated purpose is to report to non-specialists what the psychological evidence supports. The review reported that the Comprehensive System scoring of the Rorschach performed poorly on the psychometric criteria described earlier in this article; that assessment has been contested by other researchers ever since, and the argument continues.
Behind all of the popular instruments sat a finding about the reader rather than the test. In 1947 Ross Stagner gave personality tests to a group of personnel managers and then handed each of them the same feedback, assembled not from their answers but from horoscopes and graphological material; more than half rated it accurate. The following year Bertram Forer ran the version everyone remembers, giving 39 students what each believed was an individual personality sketch and what was in fact one sketch, assembled from a newsstand astrology book, for all of them; the average accuracy rating was 4.30 out of 5, and Forer published it in 1949 as the fallacy of personal validation. Paul Meehl attached the name Barnum effect to the phenomenon in a 1956 essay. The finding does not distinguish good instruments from bad ones — it concerns how convincing vague feedback is to its recipient — and it recurs wherever personality description is delivered to the person described, which is why it appears in this directory again on the page about psychological astrology.
The Myers-Briggs Type Indicator, the popular instrument the charter named, and the one whose public standing and research standing diverged sharply, is covered in full, along with Jung's Psychological Types and the analytic tradition behind it, on the alt.psychology.jung page.
A scientific room on a public network
What did it mean, in practice, for a Usenet group to be held to citation standards? Less than the phrase suggests, and more than nothing.
It did not mean a rule. This group was unmoderated, and its charter, quoted in full above, is a list of admissible topics and nothing else: no requirement to document a claim, no procedure for anything, no sanction of any kind. The one group in the branch whose founding document did ask its authors to cite the literature or otherwise document their claims was the moderated sci.psychology.research, chartered a year earlier; that charter, and the branch's moderated groups generally, belong to the alt.psychology page linked at the head of this article. In an unmoderated group a citation is a rhetorical move, not a gate. Nobody could be stopped from posting without one.
What the norm rested on instead was the structure of the surrounding literature, which happened to suit the medium unusually well. Personality research of this period was published in a small number of identifiable journals, its claims came with authors, volumes and page numbers, and the dates did real work: an argument about whether the five-factor model was theory or taxonomy could be settled, at least as to what had actually been said, by anyone willing to walk to a library. A poster who wrote Block, Psychological Bulletin, 1995 had made a checkable statement, and a poster who wrote studies show had not. That asymmetry is the whole mechanism, and it needs no enforcement to operate: it operates on readers.
Against it ran the standing difficulty of any research group on a public network, which is structural and does not require anybody to behave badly. A Big Eight group propagated, unasked, to every server that carried the branch, which in the 1990s meant most universities and a good many companies. No credential could be checked in either direction, because none can be: a signature claiming a chair and a signature claiming nothing arrive in the same typeface. Nothing distinguishes an expert's considered summary from a confident paraphrase of a magazine article except the reader's ability to tell, which is precisely the ability under dispute in a group about measurement. And the subject was of a particular kind, in a way its neighbours in the sci.* hierarchy were not: everybody arrives at a discussion of personality holding what feels like data, namely themselves and the people they know. Oceanography has no equivalent constituency of readers with lifelong first-hand experience of the sea floor.
The archive added a permanence nobody had designed for. Deja News began archiving Usenet on the web in March 1995, three months before this group existed, so the group is one of the earlier scientific rooms whose entire life was, in principle, searchable by strangers from the day it opened. A remark made in an argument in 1996 was not ephemeral, and the professional participants in the branch were aware of it; the ethics of posting under one's own name in a public forum were being worked out in these years across the whole of academic Usenet, and this group inherited the general situation rather than creating its own.
One further structural point belongs here because the paperwork states it. The group was designed as the scientific counterpart to an existing lay group, and the proposal said so. That gave it something most Usenet groups lack: a stated audience defined by contrast. Whether the contrast held in the threads is not recoverable from the record, and this page does not claim it was.
After the gateway: the replication crisis
What follows postdates the newsgroup's active life and the news2mail gateway alike, and is included because no honest account of this literature can stop in 2004 and leave the reader with the field's confidence of that year.
Warnings had been on the record for decades. Concerns about the scarcity of direct replications in psychology were voiced in the late 1960s and early 1970s; the work of Paul Meehl, Jacob Cohen and of Amos Tversky and Daniel Kahneman in the 1960s and 1970s is now routinely read as an early statement of the problem; and the suspiciously high rate of positive findings in the journals was being commented on from 1962 onwards. None of it changed practice much.
What did was a sequence of events beginning in 2011. Daryl Bem published a series of experiments in the Journal of Personality and Social Psychology reporting evidence for precognition. The paper drew heavy methodological criticism, reanalysis found no evidence for the effect, and direct replications failed — but the aspect that unsettled the field, as later accounts emphasise, was that the procedures and statistical tools Bem had used were ordinary research practice. Later the same year Joseph Simmons, Leif Nelson and Uri Simonsohn published False-positive psychology: undisclosed flexibility in data collection and analysis allows presenting anything as significant in Psychological Science, demonstrating how much undisclosed analytic latitude was compatible with the field's conventions. Failures to replicate accumulated in several literatures, including social priming and ego depletion, and two much-cited replication attempts on John Bargh's 1996 elderly-priming study did not reproduce it.
The systematic study came in August 2015. The Open Science Collaboration, coordinated by Brian Nosek, repeated 100 studies drawn from three journals — the Journal of Personality and Social Psychology, Journal of Experimental Psychology: Learning, Memory, and Cognition, and Psychological Science. Of 97 original studies that had reported significant effects, 36 per cent replicated at the conventional threshold, with effect sizes averaging about half the original magnitude. A 2018 study in Nature Human Behaviour repeated 21 social and behavioural science papers from Nature and Science and reproduced about 62 per cent of them.
Two features of that record bear on this group's subject specifically. The first is that the flagship journal of personality and social psychology was one of the three sampled, so the replication crisis reached this literature directly rather than by analogy. The second is that the correlational, large-sample trait research described in this article differs in method from the small-sample experimental social psychology where most of the failures clustered, and commentators have disagreed since about how much of the crisis transfers. That disagreement is live and is not adjudicated here.
Related work in the same period tested the structural claims themselves outside the populations that had produced them. In 2013 Michael Gurven, Christopher von Rueden, Maxim Massenkoff, Hillard Kaplan and M. Lero Vie published a study in the Journal of Personality and Social Psychology asking how universal the Big Five is, using data from forager-farmers in the Bolivian Amazon; the expected five-factor structure did not emerge as it does in the college samples on which the model was largely built. The general point — that results obtained from undergraduates who volunteer for course credit do not always reproduce in other populations or other languages — had been made before, but this was a hard case.
By the time any of this happened the newsgroup was long past its active life, and none of it was argued out there. It is set down because a reader arriving at this page from a 1990s bibliography should know what happened next to the confidence with which some of these claims were being made.
The group's own record, and what is not established
The evidence for the group itself is administrative rather than conversational, and it is worth saying exactly what survives. There is the Request for Discussion of 14 March 1995, two Calls for Votes, the result posting of 5 June with its tallies and its list of the moments each proposal crossed the threshold, and the control messages in the Internet Systems Consortium's archive. There is the one-line entry in the newsgroups file, which reads today as it read in 1995. That is the group's paperwork, and it is complete.
The control archive also preserves something the paperwork alone would not show: how propagation actually worked. The creating message of 12 June 1995 was reissued verbatim on 19 June, on 12 July, and again on 31 January 1996 — the routine periodic re-broadcast that swept up servers which had missed or discarded the first one. Two further newgroup messages in the archive originate not from the Big Eight's control issuer but from individual sites, one in Japan and one in Texas, echoing the creation locally. Nothing about a Usenet group's existence was ever a single event; it was a state that had to be continually re-asserted across thousands of independently administered machines.
There is one anomaly. In November 2001 an rmgroup control message for sci.psychology.personality entered the archive, crossposted to sci.config, sci.control and sci.groups, whose entire body reads: please remove the bogus newsgroup sci.psychology.personality. It carries a nonsense sentence as a signature, and it did not come from the address that had issued the group's newgroup messages. Its stated ground was false on the face of the record — the group had been created by a published vote six years earlier and was in the newsgroups file — and it had no effect: the group was not removed and is listed by the Internet Systems Consortium to this day. Forged and mistaken control messages were a familiar feature of the period, and an rmgroup was in any case only ever advisory, each administrator deciding whether to honour it. The episode is recorded here because it is in the archive, and for no larger reason.
What the record does not establish is a longer list than what it does. There is no readership or subscriber figure for this group at any point in its life. There is no article count, for any year. There is no way to establish what proportion of its participants were academics, students, practitioners or readers with no professional connection at all, because the medium recorded no such thing. No thread is described on this page, no post is quoted, and no participant is named except those who signed the formal proposals or ran the ballot, all of whom did so in public documents. Nothing in the surviving administrative record supports any claim about the tone, volume or quality of what was posted, and none is made.
Scope and limits
This article draws on two bodies of evidence, and it is worth stating which claims rest on which. The group's own history — dates, tallies, charter, description line, control messages — comes from the Usenet administrative record archived by the Internet Systems Consortium, and every figure quoted appears in one of those documents. The charter and the newsgroups line are quoted exactly; the phrasing of the proposal's rationale is reported closely and can be checked against the March 1995 Request for Discussion in the same archive.
The history of the field is cited by author, journal and year throughout, so that a reader can check it, and because in this literature secondary paraphrase is unusually unreliable about dates. Where a date or a number is contested or was unavailable, it has been left out rather than approximated.
On the substance, this page is documentary and historical only. It describes what researchers argued and when, and attributes each position to the people who took it. It offers no clinical advice and no diagnostic guidance, interprets no instrument and no score, and states of no test that it is suitable for any purpose whatever. Where a body of work is contested — and the behaviour-genetic literature above is contested on methodological, statistical and political grounds simultaneously — the designs, the published findings and the published criticisms are set out beside one another, and the reader is left to consult the sources. The absence of a verdict in those sections is deliberate and is not an oversight.
Finally, a note for anyone who arrived by following an old link. This group is far better evidenced by its own paperwork than by anybody's index of it. It was voted into existence 122 to 19 out of 171 valid votes, its charter survives verbatim in the newsgroup-creation archives, and its description — all personality systems and measurement — still sits in the file that news administrators install today, three decades after a votetaker counted the ballots.
Reading sci.psychology.personality today
- Historical archive: Google Groups — sci.psychology.personality (coverage varies by group and era).
- Open in a newsreader:
news:sci.psychology.personality— the original site offered exactly this link, and it still works if your system has a newsreader registered for thenews:scheme. - Live access: point an NNTP newsreader at a modern server — see accessing Usenet today.
- The original news2mail e-mail subscription service ended in the mid-2000s and no longer operates.