"A child's learning is the function more of the characteristics of his classmates than those of the teacher." James Coleman, 1972
Showing posts with label VAM. Show all posts
Showing posts with label VAM. Show all posts

Monday, January 24, 2022

How to Learn Nothing from the Failure of VAM-Based Teacher Evaluation

The Annenberg Institute for School Reform is a most exclusive academic club lavishly funded and outfitted at Brown U. for the advancement of corporate education in America. 

The Institute is headed by Susanna Loeb, who has a whole slew of degrees from prestigious universities, none of which has anything to do with the science and art of schooling, teaching, or learning.  

Researchers at the Institute are circulating a working paper that, at first glance, would suggest that school reformers might have learned something about the failure of teacher evaluation based on value-added models applied to student test scores. The abstract:

Starting in 2009, the U.S. public education system undertook a massive effort to institute new high-stakes teacher evaluation systems. We examine the effects of these reforms on student achievement and attainment at a national scale by exploiting the staggered timing of implementation across states. We find precisely estimated null effects, on average, that rule out impacts as small as 1.5 percent of a standard deviation for achievement and 1 percentage point for high school graduation and college enrollment. We also find little evidence of heterogeneous effects across an index measuring system design rigor, specific design features, and district characteristics.

So could this mean that the national failure of VAM applied to teacher evaluation might translate to decreasing the brutalization of teachers and the waste of student learning time that resulted from the implementation of VAM beginning in 2009? No such luck.  
 
The conclusion of the paper, in fact, clearly shows that the Annenbergers have concluded that the failure to raise test scores by corporate accountability means (VAM) resulted from laggard states and districts that did not adhere strictly to the VAM's mad methods.  In short, the corporate-led failure of VAM in education happened as a result of schools not being corporate enough:

Firms in the private sector often fail to implement best management practices and performance evaluation systems because of imperfectly competitive markets and the costs of implementing such policies and practices (Bloom and Van Reenen 2007). These same factors are likely to have influenced the design and implementation of teacher evaluation reforms. Unlike firms in a perfectly competitive market with incentives to implement management and evaluation systems that increase productivity, school districts and states face less competitive pressure to innovate. Similarly, adopting evaluation systems like the one implemented in Washington D.C. requires a significant investment of time, money, and political capital. Many states may have believed that the costs of these investments outweighed the benefits. Consequently, the evaluation systems adopted by many states were not meaningfully different from the status quo and subsequently failed to improve student outcomes.

So the Gates-Duncan RTTT corporate plan for teacher evaluation failed not because it was a corporate model but because it was not corporate enough!  In short, there were way too many small carrots and not enough big sticks.
 

Friday, February 19, 2016

Judge Refuses to Stop Unfair and Unreliable Value Added System

Not in the bolded section of this news story, the federal judge notes that the teachers have a strong case against the unreliable nature of the the Tennessee VAM system known as TVAAS.  He just refuses to do anything about it.
....TVAAS is a complex algorithm that aims to isolate the impact of individual teachers on their students’ learning, as measured by state tests. One of the nation’s first “value-added” formulas, it has inspired similar efforts in other states.

TVAAS scores have been calculated since the 1990s but started being used to help determine ratings, bonuses and tenure status only since 2011, when Tennessee overhauled its teacher evaluation law.

Under state law, TVAAS scores make up 35 percent of teachers’ ratings, with the rest based on in-person observations and “achievement measures,” which can include graduation rates, students’ AP or IB exam scores, or school-wide TVAAS scores.

The two teachers who filed the lawsuit, Lisa Trout and Mark Taylor, had strong ratings from classroom observations but TVAAS scores that were too low to make them eligible for bonuses from Knox County Schools. The district gives bonuses of up to $2,000 a year to teachers with strong ratings. Trout and Taylor charged that those scores should be discounted because only some of their students took the end-of-course exams used to generate the TVAAS scores.

Taylor’s rating was based on scores of just 22 of his 142 students, he said, rendering his TVAAS score meaningless.

Court documents reflect an exchange between William Sanders, the statistician who designed TVAAS, and Taylor’s parents, with whom he is acquainted.

Sanders was asked if a TVAAS score based on test scores of only a small fraction of a teacher’s students reflect a proper use of TVAAS. His answer: “For an overall evaluation of the effectiveness of the teacher to facilitate student academic progress, of course not.”

Mattice said in his ruling that he found the criticism compelling. But ultimately, the court ruled that Knox County had the right to hold back bonuses based on Taylor’s TVAAS. And he said the court did not have authority to tell the state legislature to come up with a different way to factor student learning into teachers’ ratings.

“It bears repeating that Plaintiffs’ concerns about the statistical imprecision of TVAAS are not unfounded,” the opinion reads. “However, this Court’s role is extremely limited. The judiciary is not empowered to second-guess the wisdom of the Tennessee legislature’s approach to solving the problems facing public education.”....

Saturday, January 02, 2016

Value Added Modeling Goes to Court


With numerous court cases filed to challenge the fairness, reliability, and/or validity of the various mis-applications of value-added modeling (VAM) to reward and punish teachers, it is appropriate to consider the huge body of evidence that may be shaped to drive a stake through the heart of the monstrous chimera that the tobacco-chewing Bill Sanders conjured just over 30 years ago from his ag desk at the University of Tennessee.  The following excerpt from The Mismeasure of Education considers the reliability issue, alone.  The other elements of the VAM critique comprise the remainder of Part 3 of the book.
 
. . . . In 1983, William Sanders was a station statistician at the UT Agricultural Campus and an adjunct professor in UT’s College of Business.  Based on his own experimentation with modeling the growth of farm animals and crops, Sanders proposed a hypothetical and reductive question:  Can student achievement growth data be used to determine teacher effectiveness?  He then built a statistical model, ran the data of 6,890 Knox County students through his model and answered his own question with an unequivocal affirmative. Proceeding from this single study, Sanders’ claims went beyond the customary correlational relationship between or among variables that statisticians find as patterns or trends in data.  He pronounced that teachers not only contributed to the rate of student growth, but that teacher effectiveness was in fact the most important variable in the rate of student growth: 
If the purpose of educational evaluation is to improve the educational process, and if such improvement is characterized by improved academic growth of students, then the inclusion of measures of the effectiveness of schools, schools systems, and teachers in facilitating such growth is essential if the purpose is to be realized. Of these three, determining the effectiveness of individual teachers holds the most promise because, again and again, findings from TVAAS research show teacher effectiveness to be the most important factor in the academic growth of students. (Sanders & Horn, 1998, p. 3).
This provocative pronouncement sent other statisticians, psychometricians, mathematicians, economists, and educational researchers into a bustle of activity to examine the Sanders’ statistical model and his various claims based on that model, which became the Tennessee Value-Added Assessment System (TVAAS), and which is now marketed by SAS as Education Value Added Assessment System SAS® EVAAS®.  In sharing the findings of that body of research, it becomes clear that using TVAAS to advance education policy is 1) undesirable when examined for high standards of reliability, validity and fairness, and 2) counterproductive in reaching high levels of academic achievement. 
At first glance, Sanders’ question seems reasonable and his answer logical.  Why entrust students to teachers for seven or more a day and expect children to grow academically, if teachers do not contribute significantly to student learning?  In the Aristotelian tradition of logical conclusions “if a=b and b=c, then c=a,” Sanders’ argument goes something like this:
a)    if student test scores are indicative of academic growth, and
b)    academic growth is indicative of teacher effectiveness, then
c)    test scores are indicative of teacher effectiveness; and therefore, they should be used in teacher evaluation.
If test scores are improving, then, academic growth has increased and teacher effectiveness is greater.  On the contrary, if test scores are not improving, academic growth has decreased and teacher effectiveness is diminished. 
         When Sanders made his causal pronouncements, he made assumptions that, in turn, made the teaching and learning context largely irrelevant.  Using complex statistical methods previously applied to business and agriculture to study inputs and outputs of systems, Sanders developed formulae that eliminated student background, educational resources, district curricula and adopted instructional practices, and the learning environment of the classroom as variables that impact student learning.   Sanders’ expressed his rationale for attempting to isolate teacher effect from the myriad of effects on student learning at an online listserv discussion with Gene Glass and others in 1994:  
The advantage of following growth over time is that the child serves as his or her own “control.”  Ability, race, and many other factors that have been impossible to partition from educational effects in the past are stable throughout the life of the child.  [http://gvglass.info/TVAAS/]

What we know, of course, is that a child’s economic, social, and familial conditions can and do change depending on the larger contexts of national recessions, mobility, divorce, crime, and changing education policies.  It stands to reason, then, that the high stakes claims made while using statistical models such as TVAAS must be held to higher standards of proof when making high-stakes declarations of causation that simultaneously render contexts irrelevant.  These standards are shaped by the following questions:  Are the value-added assessment findings reliable?  Are the findings valid? Are the findings fair?  Based on critical reviews of leading statisticians, mathematicians, psychometricians, economists, and education researchers, the Sanders Model does not meet these standards of proof when used for high-stakes purposes, with the most egregious shortcomings apparent when used for the dual purposes of diagnostics and evaluation.  Briefly, there are three reasons the Sanders Model falls short: Sanders assumes 1) that tests and test scores are a reliable measure of student learning; 2) that characteristics of students, classrooms, schools, school systems, and neighborhoods can be made irrelevant by comparing a student’s test scores from year to year; 3) that value-added modeling can capture the expertise of teachers fairly—just as fairly as Harrington Emerson thought he could apply Frederick Taylor’s (1911) Principles of Scientific Management to the inefficiency of American high schools in 1912 (Callahan, 1962).  As explained in Part II, the TVAAS is simply the latest iteration of business efficiency formulas misapplied to educational settings in hopes of producing a standard product in a cost effective fashion.  Sanders made this clear in 1994 while distinguishing TVAAS from other teacher evaluation formats,
TVAAS is product oriented.  We look at whether the child learns—not at everything s/he learns, but at a portion that is assessed along the articulated curriculum, a portion each parent is entitled to expect an adequately instructed child will learn in the course of a year ([http://gvglass.info/TVAAS/]).
This quote is even more telling for what remains implicit, rather than expressed: “the portion that is assessed” is the portion that can be reduced to the standardized test format, which is required for Dr. Sanders to be able to perform his statistical alchemy to begin with.  Not only does Dr. Sanders claim to speak here for the millions of parents who may have greater expectations for their children’s learning than those reflected in Dr. Sanders’ minimalist expectations, but he also tips his hand as to a deeper accountability and efficiency motive that is exposed by his concern for the “adequately instructed child.”  Once the other contextual factors (resources, poverty levels, parenting, social and cultural capital, leadership, etc.) have been excised from Dr. Sanders’ formulae, the only remaining contextual factor (the teacher) must absorb the full weight of any causative change.  Such contextual cleansing may make for beautiful statistical results, but it performs a devastating reduction to what is considered learning in schools, all the while acknowledging to not care a whit for either what is taught or how it is taught. All these basic shortcomings are reinforced by assuming that learning is linear, a demonstrably false assumption that will be discussed later in this section.
         While there were distinct statistical and psychometric challenges to the efficacy of TVAAS in the educational measurement and evaluation research literature from 1994 to 2012, one overarching theme ran through them all:  value-added modeling at the teacher-effect level is not stable enough to determine individual teacher contributions to student academic performance, especially as it is related to personnel decisions, i. e., evaluation, performance pay, tenure, hiring, or dismissal decisions.  As early as 1995, scholars (Baker, Xu, & Detch, 1995) offered a strong warning that the use of TVAAS for high-stakes might create unintended consequences for both teachers and students, such as teaching the test (thus narrowing the curriculum), teaching test skills instead of academic skills, over-enrolling students in special education since special education scores were not counted in TVAAS calculations, cheating to raise test scores, and using poor test performance to hurt teachers professionally.  By 2011 researchers (Corcoran, Jennings, & Beveridge, 2011) offered empirical evidence that all teachers do not teach to the test, but when they do, student learning depreciates more quickly than when teachers teach to general knowledge domains and expect students to master concepts and apply skills
In the following examination of these three standards of proof (reliability, validity, and fairness), we summarize the findings of national experts in statistical modeling, value-added assessment, education policy, and accountability practices.  Taken together, they provide irrefutable evidence that Sanders fails to meet these standards by using TVAAS for high-stakes decision-making such as reducing resources to schools, closing low-scoring schools, or sanctioning and/or rewarding teachers.  Most importantly, however, is the Sanders-sanctioned myth that if students are making some yearly growth on tests that were constructed for diagnostic rather than evaluative purposes, then those students will have received an education sufficient for a successful life, economically, socially, personally.

The Tennessee Value-Added Assessment Model—Reliability Issues
         The Tennessee Comprehensive Assessment Program (TCAP) achievement test is a standardized, multiple-choice test composed of criterion-referenced items administered to 3-8th grades.  It is purported to measure student mastery of the general academic concepts and skills as well as specific Tennessee learning standard objectives.  Controversy surrounding the use of achievement tests stems from the degree of test reliability needed for the high stakes purposes for which they are used.  To achieve reliability, achievement test scores must be consistent over repeated test measurements and free of errors of measurement (RAND, 2010).   The degree of reliability is biased by test construction such as vertical scales or test equating and test use such as diagnostic versus evaluative. 
         As early as 1995, the Tennessee Office of Education Accountability (OEA) reported “unexplained variability” in the value-added scores and called for an outside evaluation of all components of the TVAAS that included the tests used in calculating the value-added scores (p. iv). A three-person outside evaluation team included R. Darrell Bock, a distinguished professor in design and analysis of educational assessment and professor at the University of Chicago; Richard Wolfe, head of computing for the Ontario Institute for Studies in Education; and Thomas H. Fisher, Director of Student Assessment Services for the Florida Department of Education.  The outside evaluation team investigated the 1995 OEA concern over the achievement tests used by TVAAS. Of particular interest to TVAAS evaluators were the test constructions of equal interval and vertical scales and the process of test equating [needs layman’s terms or explanation].
         Bock and Wolfe found the scaling properties acceptable for the purpose of determining student academic gain scores from year to year, but unacceptable for determining district, school, and teacher effect scores (p. 32). All three evaluators had concerns about test equating. Fisher’s (1996) concerns focused on the testing contractor, CTB/McGraw-Hill, having the sole responsibility for developing multiple test forms of equal difficulty at each grade level, stating that “[t]est equating is a procedure in which there are many decisions not only about initial test content but also about the statistical procedure used.  If care is not exercised, the content design will change over time and the equating linkages will drift” (p. 23).  And indeed, Bock and Wolfe found in their examination of Tennessee’s equated tests forms that test form difficulty (due to item selection) created unexpected variation in gain scores at some grade levels.  Bock and Wolfe also emphasized the importance of how the scale scores, used in calculating value-added gain scores, are derived (pp. 12-13). Why are the scaling properties of tests important?
         Achievement tests are measurement tools designed to determine where on a continuum of learning a student’s performance falls.  The equal interval scale of a test is the continuum of knowledge and skills divided into equal units of “learning” value.  If one thinks of measuring learning along a number line ranging from 1 to 100, the assumption of equal intervals of learning would be that the same “amount” of learning occurs whether the student’s scores move from 1 to 2 or from 50 to 51 or from 98 to 99 on the number line or measurement scale.  The leap of faith here is that the student who scores 1 to 10 at the less difficult end of the testing continuum has learned a greater amount than the student who increases his or her score by fewer intervals, say 95 to 99, at the most difficult end of the continuum.  Dale Ballou (2002), an economics professor at Vanderbilt University who collaborated with Sanders during the 1990s, has maintained that the equal intervals used to measure student ability are really measuring the ordered difficulty of test items, with the ordering of difficulty determined by the test constructor (p. 15).  If, for example, a statistics and probably question requiring students make a prediction based on various representations is placed on a third grade test, it is considered more difficult than an item requiring the student to simply add or substract.  It is difficult to say who has learned more, the third grade gifted student who answers the statistical question correctly but makes less progress than the student who answers all the calculation questions correctly and appears from the test score to make more progress.  Therefore, student ability is inferred from test scores and not truly observed, thus making value-added teacher effect estimates better or worse depending on how these equal interval scales are designed for consistently measuring units of “learning” at every grade level over time.  Differences in units of measurement from scale to scale yield differences in teacher effect scores, even though the selection and use of the scales are beyond the control of any teacher.  Ballou has concluded that an built-in imprecision in scales leads to quite arbitrary results, and that  “our efforts to determine which students gain more than others—and thus which teachers and schools are more effective—turn out to depend on conventions (arbitrary choices) that make some educators look better than others” (p. 15).
The vertical scaling of an achievement test is based on the measurement of increasingly difficult test items from year to year on the same academic content and skills. Vertical scaling is important to Sanders’ TVAAS model because he uses student test scores over multiple years in estimating teacher effectiveness. Therefore, the content and skills at one grade level must be linked to the content and skills at the next grade level in order to measure changes in student performance on increasingly difficult or more complex concepts and skills from third grade through eighth grade.  Problems arise, however, when there is a shift in the learning progression of content and skills.  For example, third grade reading may focus on types and characteristics of words and the retelling of narratives, fifth grade on types and characteristics of literary genres and interpretation of non-fiction texts, and eighth grade on evaluation of texts for symbolic meaning, bias, and connections to other academic subjects such as history or science.  While the content and skills are related from grade to grade, there may be not be sufficient linkage between content and skills or consistency in the degree of difficulty across grades and subjects to render accurate performance portraits of student and the resulting teacher effect estimates, even if one has faith that test results can mirror teacher efforts in the best of all possible worlds:  Shifts in the mix of constructs across grades can distort test score gains, invalidate assumptions of perfect persistence of teacher effects and the use of gain scores to measure growth, and bias VAM [value-added model] estimates” (McCaffrey & Lockwood, 2008, p. 9). 
The same kinds of inconsistencies can occur when using the same tests for other high stakes measure such as school effectiveness.  Using the Sanders Model and eight different vertical scales for the same CTB/McGraw Hill tests at consecutive grade levels, Briggs, Weeks, and Wiley (2008) found that
the numbers of schools that could be reliably classified as effective, average or ineffective was somewhat sensitive to the choice of the underlying vertical scale. When VAMs are being used for the purposes of high-stakes accountability decisions, this sensitivity is most likely to be problematic (p. 26). 
Lockwood (2006) and his colleagues at RAND found that variation within teachers’ effect scores persisted, even when the internal consistency reliability between the procedures subtest and the problem-solving subtest from the same mathematics test was high.  In fact, there was greater variation from one subtest to the next than there was in the overall variation among teachers (p. 14).  The authors cited the source of this variation as “the content mix of the test” (p. 17), which simply means that test construction and scaling are imperfect enough to warrant great care and prudence when applying even the most perfect statistical treatments under the most controlled conditions.
For students to show progress in a specific academic subject (low-stakes) and for Sanders to isolate the teacher effect based on student progress (high-stakes), tests require higher degrees of reliability in equal interval, vertical scaling and test equating.  Tests are designed and constructed to do a number of things, from linking concepts and skills for annual diagnostic purposes to determining student mastery of assigned standards of learning.  They are not, however, designed or constructed to reliably fulfill the value-added modeling demands placed on them.  Though teachers cannot control the reliability of test scaling or the test item selection that represents what they teach, they can control their teaching to the learning objectives of the standards most likely to be on tests at their grade levels—those learning objectives that lend themselves easily to multiple-choice tests.  For example, an eighth teacher might have students identify bias in different reading selections, easily tested in a multiple choice format, instead of studying the effect of reporting bias in the news and research literature on current political issues and policy decisions in the students’ community.   Or if she does teach the later lesson, she and her students get no credit on a multiple-choice test for their true level of expertise in teaching and understanding the concept of bias.  In fact, if she spends the time to examine reporting bias as part of the student’s social and political environment at the expense of another test objective, student scores and her resulting value-added effect designation may suffer.  This is an unintended consequence of high-stakes testing and a survival strategy for teachers whose position and salary are bound to policies and practices that focus on high test scores.  As an invited speaker to the National Research Council workshop on value-added methodology and accountability, Ballou pointedly went to the heart of the matter when he acknowledged the “most neglected” question among economists concerned with accountability measures:
The question of what achievement tests measure and how they measure it is probably the [issue] most neglected by economists…. If tests do not cover enough of what teachers actually teach (a common complaint), the most sophisticated statistical analysis in the world still will not yield good estimates of value-added unless it is appropriate to attach zero weight to learning that is not covered by the test. (Braun, Chudowsky, & Koneig, 2010, p. 27).
In addition to these scaling issues, the reliability of the teacher effect estimates is a problem in high-stakes applications when compromised by the timing of the test administration, summer learning loss, missing student data, and inadequate sample size of students due to classroom arrangements or other school logistical and demographic issues.
Achievement tests used for value-added modeling are generally administered once a year.  Scores from these tests are then compared in one of three ways:  (1) from spring to spring, (2) from spring to fall, or (3) from fall to fall.   Spring to spring and fall to fall schedules introduce what has become known as summer learning loss—what students forget during summer vacation.  This loss is different for different students depending on what learning opportunities they have or do not have during the summer, e. g., summer tutoring programs, camps, family vacations, access to books and computers.  What John Papay (2011) found in comparing different test administration schedules was that “summer learning loss (or gain) may produce important differences in teacher effect” and that even “using the same test but varying the timing of the baseline and outcome measure introduces a great deal of instability to teacher rankings” (p. 187).  Papay’s warned policymakers and practitioners wishing to use value-added estimates for high-stakes decision making that they “must think carefully about the consequences of these differences, recognizing that even decisions seemingly as arbitrary as when to schedule the test within the school year will likely produce variation in teacher effectiveness estimates” (p. 188). 
In addition to the test schedule problem for pre/post test administration, achievement tests are usually administered before an entire school year is completed, meaning the students’ achievement test scores impact two teachers’ effect scores each year instead of just one.  By using multiple years of student data to estimate teacher effect scores, Sanders has remained unconcerned with this issue by assuming the persistence of teacher effect on student performance is an assumption of his model.  Ballou (2005) described Sanders’ assumption in the following way “. . . teacher effects are layered over time (the effect of the fourth grade teacher persists into fifth grade, the effects of the fourth and fifth grade teachers persist into sixth grade, etc.)” (p. 6).  However, the possible “contamination” of other teachers’ influence on an individual teacher’s effect estimate was noted in the first outside evaluation of TVAAS by Bock and Wolfe (1996), who questioned the three years of data that Sanders used in his model.  Bock and Wolfe agreed that three years of data would help stabilize the estimated gain scores, but they were concerned, nonetheless, that “the sensitivity of the estimate as an indicator of a specific teacher’s performance would be blunted” (p. 21).  Fourteen years after Bock and Wolfe’s neglected warning, the empirical research presented in a study completed for the U.S. Department of Education’s Institute of Education Sciences (Schochet & Chiang, 2010) found that the sensitivity of the estimate of a specific teacher’s effect was, indeed, blunted.  They found, in fact, that the error rates for distinguishing teachers from the average teaching performance using three years of data was about 26 percent.  They concluded
more than 1 in 4 teachers who are truly average in performance will be erroneously identified for special treatment, and more than 1 in 4 teachers who differ from average performance by 3 months of student learning in math or 4 months in reading will be overlooked (p. 35).  
Schochet and Chiang (2010) also found that to reduce the effect of test measurement errors to 12 percent of the variance in teachers’ effect scores would take 10 years of data for each teacher (p. 35), an utter impracticality when using value-added modeling for high-stakes decisions that alter school communities and students’ and teachers’ lives. 
McCaffrey, Lockwood, Koretz, Louis, & Hamilton (2004) challenged Sanders’ assumption of the persistence of a teacher’s effect on future student performance.  In noting the “decaying effects” that are common in social science research, they concluded the Sanders claim of teacher effect immutability over time “is not empirically or theoretically justified and seems on its face not to be entirely plausible (p. 94).  In fact, in earlier research, McCaffrey and his colleagues at RAND (2003) developed a value-added model that allowed for the “estimation of the strength of the persistence of teacher effects in later years” (p. 59) and found that “teacher effects dampen [decay] very quickly” (p. 81).  As a result, they called for more research concerning the assumption of persistence.  Mariano, McCaffrey, and Lockwood’s (2010) research concerning the persistence of teacher effect showed that “complete persistence of teacher effects across future years is not supported by data” (in Lipscomb et al, 2010, p. A14).   Using statistical methods to measure teacher persistence effect on student performance across multiple years in math and reading, Jacob, Lefgrens and Sim (2010) determined that “only about one-fifth of the test score gain from a high value-added teacher remains after a single year…. After two years, about one-eighth of the original gain persists” (p. 33).  They went on to say that “if value-added test score gains do not persist over time, adding up consecutive gains [over multiple years] does not correctly account for the benefits of higher value-added teachers” (p. 33).  In light of these more recent research studies, Sanders’ unwavering claims have proven more persistent than the teacher effect persistence that he claims.  In light of the mounting body of research that, at a minimum, acknowledges deep uncertainty regarding the persistence of teacher effect, the claim by Sanders and Rivers (1996) that the “residual effects of both very effective and ineffective teachers were measurable two years later, regardless of the effectiveness of teachers in later grades,” (p. 6) clearly needs to be reexamined and explicated further.
By claiming the persistence of a teacher’s influence on student performance, Sanders is able to assume that access to three years of data lessens the statistical noise created by missing student test scores, socioeconomic status, or other factors that affect teacher effect scores (Sanders, Wright, Rivers, & Leandro, 2009).  Missing data was a primary issue in the 1996 evaluation of TVAAS by Bock, Wolfe and Fisher. Their examination of the data quality showed that missing data could cause distortion to the TVAAS results and the “linkage from students to teachers is never higher than about 85 percent, and worse in grades 7-8, especially in reading” (p.18). It is important, of course, that test scores of every student in every classroom in every school are accounted for and attached or linked to the correct teacher when computing teacher effect scores.  Poor linking commonly occurs in schools, however, due to student absences, students being pulled out of class for special education, student mobility, and team teaching arrangements, just to name a few (Baker et al, 2010).  Missing or faulty data contributes to teachers having incomplete sets of data points (student test scores) and “can have a negative impact on the precision and stability of value-added estimates and can contribute to bias” (Braun, Chudowsky & Koneig, 2010, p. 46).  A small set of student test scores, for example, can be impacted by an overrepresentation of a subgroup of students (i.e., socioeconomic status or disabilities), or that same set of scores may be significantly impacted by a single student with very different scores, high or low, from the other students in a class.  
In addition to assuming multiple years of data will increase the number of data points per teacher sufficiently, Sanders assumes that missing data are random, but this is doubtful as students whose test data are missing are most often low-scoring students (McCaffrey et al, 2003, p. 83) who missed school or moved, entered personal data incorrectly, or took the test under irregular circumstances (i.e. special education modifications, make-up exams) and may have been improperly matched to teachers (Dunn, Kadane, & Garrow, 2003).  Sanders attempts to capture all available data for teacher effect estimates by recalculating those estimates to include missing data from previous years that is eventually matched properly to the correct teacher (Eckert & Dabrowski, 2010).[i]  This causes variability in past effect scores for the same teacher and increases skepticism about TVAAS accuracy, especially when such retrospective conversions come too late to alter teacher evaluation decisions based on the earlier version of scores.  In his primer on value-added modeling, Braun (2005) pointed out that the Sanders claim that multiple years of data can resolve the impact of missing data, “required empirical validation” (p. 13).  No such validation has been forthcoming from the Sanders Team.
Ballou (2005) explained that imprecision arises when teacher effect scores are based on too few data points (student test scores linked to a particular teacher).  The number of data points can be too few based on the number of years the data are collected, whether one, two, or three years.  The number of data points can be reduced, too, by small class size, missing student data, classes with a large percentage of special education students whose scores do not count in teacher effect data, or students who have not attended a teacher’s class for at least 150 days of instruction, and by shifting teaching assignments for either grade level and subject area.  Data points may be reduced, too, by team teaching situations whereby only one teacher on the team is linked to student data. Ballou’s (2005) research indicated further that the “imprecision in estimated effectiveness due to a changing mix of students would still produce considerable instability in the rank-ordering of teachers [from least effective to most effective]” (p. 18).  Ballou recommended adjusting TVAAS to account for all of a teacher’s grade level and subject area data, as too few teachers teach the same grade level and same subject over a three-year period (p. 23).  Ballou concluded, too, that only one year of student test data makes teacher effect data too imprecise to be meaningful or fair for its use in teacher evaluation.
In 2009, McCaffrey, Sass, Lockwood and Mihaly published their research concerning the year-to-year variability in value-added measures applied to teachers assigned small numbers of students such as special education teachers since special education students are often exempted from taking the tests.  The small number of student scores are impacted by extremely high or extremely low scores resulting is extremes in teacher value-added scores, “so rewarding or penalizing the top or bottom performers would emphasize these teachers and limit the efficacy of polices designed to identify teachers whose performance is truly exceptional” (p. 601).  Even though using multiple years of data helps reduce the variability in teacher effect scores, “one must recognize that even when multiyear estimates of teacher effectiveness are derived from samples of teachers with large numbers of students per year, there will still be considerable variability over time” (p. 601).
With these unresolved issues and deep skepticism related to test reliability, Sanders’ logic in justifying the use of value-added modeling for teacher accountability weakens, as in our slightly modified syllogism:
a)     if student test scores are unreliable measures of student growth and
b)    unreliable measures of student growth are the basis for calculating teacher effectiveness, then
c)   test scores are unreliable measures for calculating teacher effectiveness, at least for high-stakes decisions concerning teachers’ livelihoods and schools’ existence. 


[i] “[The current] year’s estimates of previous years’ gains may have changed as a result of incorporating the most recent student data. Re-estimating all years in the current year with the newest data available provides the most precise and reliable information for any year and subject/grade combination. Find district and school information at the following: TVAAS Public https://tvaas.sas.com/evaas/public_welcome.jsp, TVAAS Restricted: https://tvaas.sas.com/evaas/login.jsp” (Eckert & Dabrowski, 2010, p. 90).

Saturday, November 14, 2015

AERA Concludes the Facts Are Factual about VAM

The American Education Research Association (AERA) can be counted on to remain irrelevant to research discussions of great social significance.  Not surprisingly, AERA's shrinking membership numbers have coincided with the org's steady drift into the arms of education reform schoolers and the corrupt CorpEd foundations that are laser focused on redirecting education at all levels into corporate revenue streams.  

AERA's complacency and complicity have been sickening to watch, with their kowtowing to Bill Gates and his various bad ideas culminating last year when AERA announced a fellowship program for doctoral students interested channeling and then relinquishing their doctoral research to the Gates's MET database.

Now almost 20 years after legitimate researchers starting ringing the alarm bell on the value-added muddle that was thrust upon the education world by the tobacco-chewing ag statistitian, Bill Sanders, and six years after the National Academy of Sciences sent their hair-on-fire letter to Arne Duncan (which was ignored), warning him about including not-ready-for-prime-time VAM in Race to the Top requirements , AERA has finally concluded that the truth must be true:  VAM is not a legitimate tool for ANY high stakes education decisions.  

From AERA's announcement:
. . . . In recent years, many states and districts have attempted to use VAM to determine the contributions of educators, or the programs in which they were trained, to student learning outcomes, as captured by standardized student tests. The AERA statement speaks to the formidable statistical and methodological issues involved in isolating either the effects of educators or teacher preparation programs from a complex set of factors that shape student performance.

“This statement draws on the leading testing, statistical, and methodological expertise in the field of education research and related sciences, and on the highest standards that guide education research and its applications in policy and practice,” said AERA Executive Director Felice J. Levine.

The statement addresses the challenges facing the validity of inferences from VAM, as well as specifies eight technical requirements that must be met for the use of VAM to be accurate, reliable, and valid. It cautions that these requirements cannot be met in most evaluative contexts.

The statement notes that, while VAM may be superior to some other models of measuring teacher impacts on student learning outcomes, “it does not mean that they are ready for use in educator or program evaluation. There are potentially serious negative consequences in the context of evaluation that can result from the use of VAM based on incomplete or flawed data, as well as from the misinterpretation or misuse of the VAM results.”

The statement also notes that there are promising alternatives to VAM currently in use in the United States that merit attention, including the use of teacher observation data and peer assistance and review models that provide formative and summative assessments of teaching and honor teachers’ due process rights.

The statement concludes: “The value of high-quality, research-based evidence cannot be over-emphasized. Ultimately, only rigorously supported inferences about the quality and effectiveness of teachers, educational leaders, and preparation programs can contribute to improved student learning.” Thus, the statement also calls for substantial investment in research on VAM and on alternative methods and models of educator and educator preparation program evaluation. 
For a full research review and history of VAM's origin and growth in Tennessee, see The Mismeasure of Education (Horn & Wilburn, 2013).
 

Tuesday, August 18, 2015

Petrilli and Ravitch Blame Team Obama for What Bush's Hacks Cooked Up

Even though much has changed about Diane Ravitch's rhetoric over the past few years, some things have not.  For instance, she has never explained the role of her conservative policy pals in the miseducating testing accountability policies that began with Richard Nixon, gained steam with Reagan, held firm through Clinton, and soared to new heights with Bush. 

Instead, she prefers to focus on Team Obama's role, as if Obama's administration represents something more than the final delivery of rotted education policies that Ravitch, herself, helped to craft over the three decades before Obama came to Washington.  

Ravitch found companionship for her revisionist and selective policy account this week when fellow neolib, Michael Petrilli, confessed that VAM-based teacher evaluation was a policy mistake that, hold on to your Scantron,  Arne Duncan was responsible for.  

Somehow, both Ravitch and Petrilli have chosen to forget that using VAM to hold teachers and schools "accountable" for differences in scores that poverty levels create had been been around for two decades when Duncan was named chief water carrier for plans that had been formulated in the early days of Bush's NCLB.  

Besides, the war on teachers was not simply a matter of bribing states to adopt VAM teacher evaluations.  Under Bush's Rod Paige and, later, Margaret Spellings, policies were put in place to make it harder for experienced teachers to be "highly qualified," and easier for unprepared beginners to be named as "highly qualified." 

Many other teachers were run out of teaching by unethical and abusive testing practices, the proliferation of scripted teaching and curriculum, and a ramped-up policy elite rhetoric that blamed teachers for low test performance.  All of these were put into place back when Ravitch was still openly riding the corporate bandwagon.

Sure, sure, Duncan had his role.  He used the ill-fated Race to the Top grants to bribe states into adopting VAM teacher eval, Common Core, unlimited charters, and Big Data, but VAM was viewed as the new tool of education industry and conservative ed ideologues back when Arne Duncan was learning his corporate trade as CEO of Chicago schools.  

To be sure, the VAM bandwagon had been christened and launched by Bush's Margaret Spellings in 2005, as explained in this short excerpt from The Mismeasure of Education.  The Growth Model Pilot Project had Bill Sanders' version of VAM at its center.

To pretend that Arne Duncan and Team Obama were largely responsible for the coming of VAM is equivalent to blaming the pizza delivery man for the awfulness of the pie.  In fact, Arne was just delivering what had already been cooked up before his shift ever started.


The Growth Model Pilot Project (GMPP) and Sanders’ Testimony in Washington
Following a 2004 letter pleading for flexibility in NCLB accountability requirements from sixteen “state school chiefs” (Olsen, 2004), Secretary of Education, Margaret Spellings, announced the Growth Model Pilot Project in November 2005, as predicted by Dr. Sanders in his testimony in 2004 to the Tennessee House Education Committee.  Under pressure from states and municipalities faced with increasingly-impossible proficiency targets that NCLB required students from poorer districts where students were farther behind, the U.S. Department of Education developed a Peer Review Committee to evaluate state growth model proposals.
As indicated in Part I, the potential effects of NCLB’s unachievable proficiency targets were not a secret, even prior to passage. In her policy history of NCLB, Debray (2006) cites Dr. Joseph Johnson’s comments in a public address prior to NCLB passage: “[p]eople are looking at the data and saying, ‘This is going to be catastrophic because there are going to be so many low-performing schools and this isn’t going to work’” (p. 138).  Debray notes, however, that the Bush Administration, which had included a school voucher provision that was eventually struck from the final version of the Act, “had a political reason to want to see nonimproving schools identified so that NCLB’s options to exit such schools for better ones or receive private supplemental instruction would produce visible results of Bush’s educational innovations in the first term.  There was political interest in identifying lots of failing schools” (p. 115).  This would also be a boon for tutoring companies, canned remedial intervention programs, and other “learning corporations” to hawk their wares, including value added testing models for assessing test score improvement over time, in a mass market of desperate educators trying to achieve unrealistic testing targets.
Originally, NCLB disallowed states the use of value-added models and nationally norm-referenced tests to measure the effects of teachers, schools, and schools on student test performance.  Instead, the USDOE directed states to use proficiency benchmarks based on criterion-referenced tests aligned with their own state standards.  In a nod to growing criticism, however, the Spellings Growth Model Pilot Project allowed states to use “projection models” that could predict student performance on future assessments, thus answering the question: “Is the student on an academic trajectory to be proficient or advanced by 2014?” On May 17, 2006, Tennessee and North Carolina, the two states where the Sanders Model was in use, were approved to use their value-added projection models to track individual student progress in meeting NCLB academic goals.  The GMPP allowed projection models as an acceptable “safe harbor” option in providing evidence that a state was making significant progress toward AYP proficiency targets.
During the implementation of the GMPP, Sanders testified to the U.S. House Committee on Education and Workforce on “No Child Left Behind: Can Growth Models Ensure Improved Education for All Students” (July 27, 2006). On March 6, 2007, Sanders presented at the U.S. Senate Committee on Health, Education, Labor, and Pensions as part of a roundtable discussion entitled “NCLB Reauthorization: Strategies for Attracting, Supporting, and Retaining High Quality Educators” (March 6, 2007).  He called on Congress to replace the existing “safe harbor” options of NCLB with projection value-added models like his, promising that “effective schooling will trump socio-economic influences.”  Even though the focus of each hearing was different,  Sanders used both opportunities to promote his own brand of value-added modeling that could separate educational influences from “exogenous factors (if not completely, then nearly so) allowing an objective measure of the influence of the district, school, and teacher on the rate of academic progress.” Without naming any of them, Sanders described his growth model competitors as having “been shown to produce simplistic, potentially biased, and unreliable estimates.” However, later in his testimony, he admitted that he “had to engineer the flexibility to accommodate other ‘real world’ situations encountered when providing effectiveness measures at the classroom level: the capability to accommodate different modes of instruction (i.e. self-contained classrooms, team teaching, etc.), ‘fractured’ student records, and data from a diversity of non-vertically scaled tests.” Sanders expressed no doubt that his engineered flexibility was up to the task, even if other growth models on the market could not “and should be rejected because of serious biases.” 
Sanders provided the members of the U.S. Senate Committee on Health, Education, Labor, and Pensions a summary of the research finding that he attributed entirely to his “millions of longitudinal student records.”  According to Dr. Sanders, by “addressing research questions that heretofore were not easily addressed,” queries of his databases had yielded that beginning teachers are less effective than veteran teachers, that inner city schools have a “disproportionate number of beginning teachers,” that turnover rates were higher in inner city schools, that high poverty schools have a lower percentage of highly effective teachers as measured by test scores, that math teachers in inner city middle schools were less likely to have high school math certification, that high poverty students assigned to effective teachers “make comparable academic progress” with low poverty students.  From talking to highly effective teachers across the country, Sanders said that he learned that these teachers knew how to “differentiate” instruction, how to use feedback from formative assessment to make instructional decisions, and how to maximize their instructional time.  Never in his testimony did Dr. Sanders indicate that most of research questions had been asked and answered before his developing and marketing of value-added assessment took place. 
Also known were the obstacles teachers dealt with daily in applying teaching skills consistently across classrooms and schools, especially in high poverty classrooms and schools. This fact, however, is minimized by Dr. Sanders’ misleading and obfuscating claims that “differences in teaching effectiveness is the dominant factor affecting student academic progress” and “the evidence is overwhelming that students within high poverty schools respond to highly effective teaching.” Clearly, there are a couple of important qualifiers that Sanders fails to mention.  First, the more obvious one: Sanders does not make explicit that his claim regarding the importance of teaching effectiveness is based solely on gauging academic progress of individuals on tests as measured by the Sanders algorithm. By omitting this most important point in his presentation to the senators and their staffs, Sanders allows the false impression to be drawn and/or perpetuated that teaching effectiveness is more important than all the other factors, whether in school or outside school, that determine the variability in student achievement across income levels, family education levels, diversity levels, social capital levels, or any of the other variables that researchers have demonstrated are more important than teacher quality in determining variability in levels of achievement among students.  As noted earlier in Part II, an impressive group of researchers (Nye, Konstantopoulos, & Hedges, 2004) just two years before Dr. Sanders’ testimony noted in a most reputable peer-reviewed journal that, among seventeen studies the researcher examined, “7% to 21% of the variance in achievement gains is associated with variation in teacher effectiveness” (p. 240). 
The second point is not so easy to tease out or to discount, for there is commonsense and empirical veracity to claiming “the evidence is overwhelming that students within high poverty schools respond to highly effective teaching.”  When set atop the previous claim, however, the weight of potential misconception becomes too heavy to ignore.  All sentient beings, we suggest, are more responsive to effective teaching than to ineffective teaching.  Since Dr. Sanders never tells us what effective teaching is; we can only assume it is the kind that produces greater test score gains than would less effective teaching. Thus, higher test score gains are produced by more effective teachers, and we know they are more effective teachers because they have higher test score gains.  We are not the first to point out the obvious circularity of this definition, but the resulting unquestioned tautology is worth keeping in mind when the term “effective teaching” is bandied about.  The more serious difficulty with the Sanders claims about highly effective teaching in high poverty schools comes from the unstated conclusion to this unfinished syllogism that high-ranking politicians rush to, when given one thoroughly misleading premise and another that is full of emotional appeal and that can’t be argued with: If teacher quality largely determines student achievement, and if poor and hungry children respond with higher achievement to good teaching, then teaching holds the key to closing that achievement gap left gaping from the last round of less than effective reforms. 
This conclusion has proved appealing to both liberal political elites and conservative political elites: to the former because of a long-held suspicion that teachers are lazy and are just not trying hard enough, and to the latter because of the long-held suspicion that teachers are self-serving louts protected by their big unions.  In either case, the political solution must be better teachers, and any policy to help advance that priority, then it must be a good policy.  And if Dr. Sanders has a tool that can help tell us know who is doing a good job and who is not, then we have an intervention worth investing in that is much less expensive and with a wider appeal, by the way, than trying to do something substantive about poverty, which has for a hundred years remained the inseparable shadow of the testing achievement gap, from whatever direction it is viewed.
Just days before Dr. Sanders offered his testimonial to the Senate Committee in March 2007, the U. S. Chamber of Commerce (2007) published a state-by-state assessment that compared state test results to achievement levels of NAEP.  Tennessee did not fare well, earning an F for “truth in advertising about student proficiency” (p. 52). By Tennessee’s own standards, however, and by Dr. Sanders’ value-added calculations, the state seemed well on its way to meeting its NCLB benchmarks. The 2005 Tennessee Report Card, which provided one score for achievement proficiency and another score for value-added gains, showed 48 percent of 3-8th grade students proficient in Math and 40 percent advanced. In reading, Tennessee 3-8th grade students were 53 percent proficient and 38 percent advanced.  Based on Tennessee’s own scale and timeline for achieving NCLB targets[1], the state gave itself a B for achievement and a B for value-added in both subjects.
For a state with 52 percent of its students economically disadvantaged, 25 percent African American, and 16 percent with special needs, Tennessee’s self-generated report card results looked respectable until set alongside results from the National Assessment of Educational Progress (NAEP), which showed Tennessee’s proficiency rates dropping in 2005, rather than moving up as measured by state proficiency scores and value-added scores.  In 2005, Tennessee claimed 87 percent of its 4th and 8th grade students were proficient in math, while NAEP proficiency levels for 4th and 8th graders were 27.7 and 20.6 percent, respectively. In reading the discrepancy was no less startling; state proficiency scores for fourth and eighth grade were 87 percent, and the NAEP proficiency scores were 26.7 and 26.2 percent, respectively.
       In April 2010, The U.S. Department of Education issued an interim Growth Model Pilot Project (GMPP) Report that reviewed the approved growth models, including Tennessee’s.  In comparing the use of the NCLB’s status model that measured the percentage of proficient students each year and the pilot states’ growth model projections, the U.S. Department of Education came to the following conclusions:
1) Simply stated, “making AYP under growth,” as defined by the GMPP, does not mean that all students are on-track to proficiency (p. 54).
2) There was little evidence that the type of model selected had an impact on the extent to which schools were identified as making AYP by growth within the GMPP framework (p. 54).
3) Schools enrolling higher proportions of low-income and minority students were more likely to make AYP under growth in the status-plus-growth framework than were schools enrolling higher proportions of more affluent and nonminority students. However, if growth were the sole criterion for determining AYP, schools enrolling higher proportions of low-income and minority students would be more likely to move from making AYP to not making AYP (p. 55).
      
       In January 2011, the U.S. Department of Education published the final GMPP Report that analyzed two years of growth data from the participating states.  The findings from the second year of the study were similar to those of the first.  For Tennessee, that meant very few additional schools made AYP (22 in 2007-2008) using the projection growth model developed by Sanders. The 2011 Report also found that the type of model does make some difference in the number of students identified to be “on-track” to reach the 2014 NCLB target of 100 percent proficiency in math and reading, and with the Tennessee projection growth model, “relatively few students with records of low achievement but evidence of improvement are predicted to meet or exceed future proficiency standards, while students with records of high achievement but evidence of slipping are very likely to be predicted to meet or exceed future proficiency standards” (p. xix).  In short, the Sanders projection model did little to address the essential unfairness perpetuated by NCLB proficiency requirements that required those further behind and with fewer resources to achieve more than privileged schools whose gains are required to be much smaller to reach the same proficiency point.  
       An interesting finding tucked away in Appendix A of the 2011 Report suggested how the parameters for using growth models could be adjusted to help identify more schools making AYP annual targets:  “…results from growth measures used for state accountability purposes suggest that many more schools would make AYP if the first of the seven core principles of the ESEA project was relaxed” (p. 112). The first of the seven core principles “requires that the growth model, like the status model, be applied to each targeted subgroup as well as all students in the school.”  This means that growth outcomes are to be monitored separately, or “disaggregated,” for major racial and ethnic groups, limited English proficient (LEP) students, special education students, and low-income students. To sacrifice or to “relax” the core principle of disaggregation so that the value-added estimates would identify more schools as making AYP would appear to neutralize NCLB’s purported goal of bringing attention to those subpopulations of students who traditionally have not made adequate progress by any measure. It seems, too, that elimination of the core principle requiring disaggregation could serve to mask the amount of growth that low SES and minority students need to make in order to become proficient by means other than the relaxing of first principles. 

       Even though the U.S. Department of Education had approved the use of value-added projection models to demonstrate which schools and districts were making AYP, skepticism remains deep among highly respected statisticians, psychometricians, and economists when VAM is used for this and other high-stakes purposes. Some of those critical reviews will be examined in Part III.


[1] The Tennessee NCLB target (AYP) for math in 2005 was 79 percent proficient/advanced and 83 percent for reading.