Monday, February 04, 2008

Together We Can Change the World...

Well, the long anticipated integration between Pearson and former test publishing giant Harcourt Assessment, Inc. (also known at times in it's history as The Psychological Corporation and the testing division of Harcourt, Brace and Jovonovich) is complete! Senior Pearson leadership were in San Anotonio this week to meet with the new Pearson employees and to celebrate the lengthy process of DOJ approval.

In a previous press release, Pearson CEO, Marjorie Scardino, said:


"We have long admired these businesses. They bring new intellectual property, capabilities and skills to Pearson, and will enable us to accelerate our strategy of leading the personalisation of learning, both in the US and around the world. We know that their people share our commitment to education, and we look forward to welcoming them as colleagues."
For me personally, it was like "old home week" seeing many of the faces I've gotten to know over the years and reviewing the wonderful products, services and staff from Pearson's new "crown jewel."

Make no mistake, while challenges lie ahead in the education and assessment arenas, Pearson's goal to educate, inform and entertain our customers remains our primary motivator and this latest acquisition is another step toward accomplishing that goal.

Welcome aboard!

You can comment on this posting by emailing TrueScores@Pearson.com.

Monday, January 28, 2008

If He is Correct, Then I Must Have Been Wrrrrrrrong(?)

I have had the occasion to know Dr. James Popham for many years and in many contexts. You might recall that Dr. Popham was the keynote speaker at the last ACT-CASMA conference and that I had devoted some space to the conference in a previous post.

Now please understand, Dr. Popham has worked in measurement for many years and describes himself as a "reformed test builder," presumably implying some sort of 12 step program. Despite this, or at least as a prelude to this, Dr. Popham has been very influential in assessment. He was an expert witness in the landmark "Debra P." case in Florida, and was involved in the early days of teacher certification in Texas and elsewhere. He is also the author of numerous publications.

Over the years I have listened to Jim says some outrageous things. For those of you who know Jim, this is no surprise. He is quite a presenter and, I suspect, basks a little too much in the glow of his own outrageousness. However, many of the things I have heard him say (at the Florida Educational Research Association-FERA meeting, for example) were just plain incorrect. I won't bother you with the specifics as I am sure Dr. Popham would claim he is correct. Yet, it does put me in a quandary. Despite his recent statements, I actually have to agree with what Dr. Popham said at the ACT-CASMA conference back in November.

Jim's theme—one he has articulated in multiple venues—regarded what he calls "instructional sensitivity." Here are the basic tenets of his argument:

"A cornerstone of test-based educational accountability:
Higher scores indicate effective instruction; lower scores indicate the opposite."

"Almost all of today's accountability tests are unable to ascertain instructional quality. That is, they are instructionally insensitive."

"If accountability tests can't distinguish among varied levels of instructional quality, then schools and districts are inaccurately evaluated, and bad educational things happen in classrooms."

I keep returning to this theme. While I make a living building assessments of all types, recently most of my efforts and those of my colleagues have been with assessments supporting NCLB, which are "instructionally insensitive" according to Dr. Popham. It is hard to believe that any assessment that asks three or four questions regarding a specific aspect of the content standards or benchmarks (and by the way does so only once a year) can be very sensitive to changes in student behavior due to instruction on that content. At the same time, having some experience teaching, testing, and improving student learning, I have seen the power that measures just like these have for teachers who know what to do with the data and have a plan to improve instruction.

Hence my dilemma: why do I keep returning to Dr. Popham's argument? While I am not ready to admit I might have been wrong to dismiss Jim as a "reformed test builder" and to ignore his rants, I do admit he has a valid point to some extent regarding instructional sensitivity. I suppose I would have called his argument "the instructional insensitivity of large-scale assessments," but who am I to quibble with vocabulary.

Dr. W. James Popham, Professor Emeritus from UCLA welcomes all "suggestions, observations, or castigations regarding this topic...." Contact him at wpopham@ucla.edu. Or send an email to TrueScores@Pearson.com, and I will forward it to him.

Friday, January 18, 2008

IQ and the Flynn Effect

Back in the 1980s when I worked on the development of the Wechsler Intelligence Scale for Children, Third Edition (WISC-III), I was fascinated with a process commonly referred to at the time as "continuous norming." Applied by Dr. Gale Roid as developed by Professor Richard Gorsuch, continuous norming was a slick way to improve the precision of empirical norms. While things seemed to get in the way of any in-depth analysis of the procedure, and while I did stay in contact with Professor Gorsuch occasionally, I did nothing to understand or apply the process anew and simply moved on.

Over the winter holidays, I was reading The New Yorker (yes, even people who live in Iowa read The New Yorker) and discovered, much to my surprise, a story about IQ written by Malcolm Gladwell, titled "None of the Above: What IQ doesn’t tell you about race" (December 17, 2007, pp. 92-96). As you may recall, Malcolm Gladwell is the author of both The Tipping Point and Blink. Both books interested me, so I read what he had to say about IQ.

Gladwell references something he (and apparently others) call “The Flynn Effect.” The Flynn Effect comes from James Flynn, author of What is Intelligence?, and is essentially the term used to describe what Flynn claims to have discovered—that all humans are getting smarter. As Gladwell points out, Flynn looked at years of IQ assessment data from all over the world and concluded that humans gain three IQ points per decade. Gladwell then tries to put this in context. For example, if Americans' average IQ in 2000 was 100, then in 1990 it was 97, in 1980 it was 94, in 1970 it was 91, and so on. If true, this implies that my grandfather (and yours) were “dull normals” at best, but were most likely mentally retarded. Flynn claims that this is due more to the way we measure intelligence than anything else. He states, as Gladwell points out:

“An IQ, in other words, measures not so much how smart we are as how modern we are.”
For example, when members of the Kepelle tribe in Liberia were asked to associate objects such as a potato and a knife, they linked them together according to function. As Gladwell points out, after all, you use a knife to cut a potato. Most IQ assessments would expect the potato to be linked to other legumes and the knife to be linked to other tools. Flynn claims modern culture has “taught” us to think in the way the IQ assessment measures and, while this is different than how the Kepelle thought, there is no reason to believe that their thinking represents anything less intelligent.

Gladwell then tries to articulate the issue that Flynn makes regarding intelligence test norms. He observes that if the center of each new edition of the WISC is 100, and everyone is getting smarter by three IQ points per decade, than each subsequent form of the WISC (the first WISC was standardized in the 1940s) must be getting harder. Very interesting—I need to dig up references on the “continuing norming” process used for the WISC and see what impact, if any, such a process might have on "The Flynn Effect."

You can comment on this posting by emailing TrueScores@Pearson.com.

Tuesday, October 23, 2007

The 2007 Coffman Lecture Series: Lindquist was Right!

Living in Iowa City, arguably the mecca of achievement testing, I am privileged and fortunate to be able to attend professional development activities that many of my colleagues located elsewhere can not. Take for instance the 2007 William E. Coffman Lecture Series sponsored by the University of Iowa and the Iowa Testing Program. This year the guest lecture was Dr. Daniel Koretz, Harvard Professor and noted measurement scholar. I have disagreed with Professor Koretz in the past, but I disagree with everyone on virtually everything so this should be of no surprise. However, the good doctor's lecture this year, while nothing new, reminded me and my staff of why we got into measurement in the first place. The title of Dr. Koretz's presentation was:

Test-based Monitoring and Accountability: Time to Take Lindquist's Warning Seriously
Dr. Koretz started his lecture by recalling the words E.F. Lindquist composed as the introduction to the 1951 edition of Educational Measurement, of which Dr. Lindquist was the author:

"The widespread and continued use of a test will, in itself, tend to reduce the correlation between the test series and the criterion series…. Because of the…potency of the rewards and penalties associated…with high and low achievement test scores of students, the behavior measured by a widely used test tends in itself to become the real objective of instruction, to the neglect of the (different) behavior with which the ultimate objective is concerned." (Lindquist, 1951, 152–153)
What this states, simply, is that even in 1951 and prior to the NCLB rage for "education reform," "high-stakes testing" and "accountability," Dr. Lindquist anticipated the consequences warned by many when we moved to high-stakes testing. This consequence is the focus on improving test scores without the improvement in learning required.

As Dr. Koretz points out in his lecture, this does not have to be as blatant as an increased motivation to cheat or otherwise subvert the system, or even the result of teaching to the test. It could be something much more subtle; one example of which he calls "Reallocation." Reallocation means, in its most simplistic, the shifting of educational resources based on assessment results or lack there of. While this might make sense—that is, if the measures say we need help in area X, we should look at improving instruction and learning about area X. But the complications lie in how we pay attention to this area of needed improvement. Or, to quote Dr. Kortez:


"Individual elements of the domain may be measured well, but representation of the domain is undermined."
An example of this—for those of us who have to make sure our corners are square—is the use of the "3:4:5" triangle as a substitution for a clear understanding and application of the Pythagorean theorem. Students are likely to get the test questions correct without knowing the Pythagorean theorem if test builders are not careful about how they ask the questions and students apply this simple rule. In this case, incorrect inferences about the Pythagorean theorem and possibly about the more general domain of geometric shapes will be made when in reality rote learning of a "trick" led to the correct answer.

Professor Koretz ended his lecture showing that efforts to reduce the "reallocation" effect NCLB has brought is failing mainly because educators are not heeding the words Dr. Lindquist posed years ago: namely, instruction should not focus on the sample of skills tested, but rather on the domain being measured. My staff and I argue it is the instruction that the dialog should be about, but clearly we need to ensure that the efficacy of the measures are of value first.

Monday, October 08, 2007

CASMA-ACT Conference Looks Great, Despite It's Name!

The CASMA-ACT Invitational Conference on Current Challenges in Educational Testing seems to offer a very interesting slate of speakers and topics, despite it's rather academic name.
Dr. Dan Koretz, from Harvard and a national assessment expert, will speak about higher education accountability.

Joe Crick, from the National Board of Medical Examiners, has been solving assessment problems almost as long as there have been assessments and will give his perspectives on performance testing simulations.

Dr. Doug Christensen, the current Commissioner of Education in Nebraska, is an outspoken critic of NCLB and to his credit has managed to implement education the way he sees fit for his native Nebraska. All of this despite NCLB being the "law of the land."
These speakers alone would make it worth the trip to Iowa City on a warm and sunny November afternoon (Saturday, November 3rd to be exact). I promise that the fall colors will still be...well, that the leaves will have color...well, some leaves will still be on the trees!

Another interesting part of the conference will be the participation of "national media" as they provide their perspectives. I hope to leave before this part. Oh, yeah, then there is Popham.

Regardless, it looks to be a great conference, and I hope you can find the time to attend. The Hawkeyes will be out of town that weekend, so hotel rooms will be easy to find and inexpensive.