ברוכים הבאים! בלוג זה נועד לספק משאבים לפסיכולוגים חינוכיים ואחרים בנושאים הקשורים לדיאגנוסטיקה באורייטנצית CHC אבל לא רק.

בבלוג יוצגו מאמרים נבחרים וכן מצגות שלי וחומרים נוספים.

אם אתם חדשים כאן, אני ממליצה לכם לעיין בסדרת המצגות המופיעה בטור הימני, שכותרתה "משכל ויכולות קוגניטיביות".

Welcome! This blog is intended to provide assessment resources for Educational and other psychologists.

The material is CHC - oriented , but not entirely so.

The blog features selected papers, presentations made by me and other materials.

If you're new here, I suggest reading the presentation series in the right hand column – "intelligence and cognitive abilities".

נהנית מהבלוג? למה שלא תעקוב/תעקבי אחרי?

Enjoy this blog? Become a follower!

Followers

Search This Blog

Friday, June 17, 2016

מעגל הסיווגים: DSM1 עד DSM5 – חלק ראשון

מעגל הסיווגים:   DSM1 עד DSM5 – חלק ראשון   

Blashfield, R. K., Keeley, J. W., Flanagan, E. H., & Miles, S. R. (2014).
The Cycle of Classification: DSM-I Through DSM-5. Annu. Rev. Clin. Psychol, 10, 25-51.

המאבקים על הגדרות מחלות הנפש ב – DSM ים השונים מזכירים בעוצמתם ובמאפייניהם את המאבקים על הגדרת לקות למידה.  זו אחת הסיבות שכדאי להכיר אותם. 

DSM = Diagnostic and Statistical Manual of Mental Disorders

מאמר זה סוקר את התפתחות ה – DSM.  פוסט זה (הראשון מבין שניים) אינו סיכום של המאמר אלא מביא מספר דברים שעניינו אותי במיוחד.       

המאמר משתמש במונח "פציינטים" כדי לתאר את האנשים הנזקקים להתערבות פסיכיאטרית ולכן לעתים גם להגדרות DSM.  בהתאם לכך אשתמש במונח "פציינטים" בפוסט הזה.  הדרך בה אנו מכנים את הלקוחות שלנו משקפת לדעתי עמדה ערכית (אני מכנה אותם "קליינטים").  את המונח mental disorder המופיע במאמר תרגמתי כ"הפרעה נפשית/הפרעת נפש".  גם הדרך בה אנו מכנים את הבעיות עמם מתמודדים הלקוחות שלנו משקפת, כמובן, עמדה ערכית.  

נתחיל בטבלה מרתקת זו:

הכנסה במליוני דולרים לאגודה הפסיכיאטרית האמריקנית
מחיר  בדולרים  
מספר  אבחנות
מספר עמודים
שנה
שם
לא ידוע
3$
128
132
1952
DSM1

1.27
 3.5$
193
119
1968
DSM2

9.33
31.75$
228
494
1980
DSM3

16.65
לא כתוב במאמר
253
567
1987
DSM3R

120
$48.95
383
886
1994
DSM4

לא ידוע
$74.95
383
943
2000
DSM4TR

לא ידוע
$199
541
947
2013
DSM5



DSM1

לאחר מלחמת העולם השניה, היו בארה"ב ארבע מערכות של קלאסיפיקציות של מחלות נפש.  האיגוד הפסיכיאטרי האמריקני APA- AMERICAL PSYCHIATRIC ASSOCIATION החלוט להתגבר על מגדל בבל זה על ידי יצירת קלאסיפיקציה שתהיה מקובלת על כל חברי האירגון ושתאחד את המונחים הדיאגנוסטים של הפסיכיאטרים.  התוצאה היתה DSM1.

מערכת הסיווג של ה - DSM1         היתה היררכית.  בצומת הראשונה הקלינאי נדרש להבחין בין תסמונות/סינדרומים מוחיים אורגנים לבין הפרעות בתפקוד (כנראה ללא גורם אורגני).  ההפרעות בתפקוד חולקו לשלושה סוגים:  פסיכוטי, נוירוטי או הפרעות אופי CHARACTER.  ארגון זה של ה – DSM היה דומה לתהליך בו אנשי המקצוע קיבלו החלטות דיאגנוסטיות.  תיאורי ההפרעות נכתבו בפרוזה שכללה מאפיינים התנהגותיים ותכונות TRAITS.  המונחים התיאורטיים היו יחסיים, כל קלינאי פירש אותם אחרת, וזה הוביל לבעיות של מהימנות בין אנשי המקצוע. 

DSM2

DSM2 שפורסם בשנת 1968 היה תוצר של מאמץ לאחד את מערכות הקלסיפיקציה בעולם.   הוא היה מאורגן באופן דומה ל – DSM1 ונוספו לו קטגוריות שהיו קשורות לטיפול בפציינטים שאינם מאושפזים.
בשנת 1971, הוכן סט של שמונה סרטוני וידאו של פציינטים מארה"ב ומבריטניה.  סרטונים אלה שהכינו KENDELL וחבריו הוצגו בפני פסיכיאטרים אמריקנים ובריטים.  האמריקנים אבחנו את כל הפציינטים שהופיעו בסרטונים כסכיזופרנים.  הבריטים אבחנו את חלקם כמאנים דפרסיבים, חלקם כסכיזופרנים וחלקים כסובלים מהפרעות אישיות.  מחקר זה סיפק עדות לכך שעדיין קיימים הבדלים בגישה בין אנשים מקצוע ממדינות שונות.  יש שראו בו עדות לכך שהאמריקנים נוטים לאבחנת יתר של סכיזופרניה.

בשנת 1973, ROSENHAN פירסם מחקר פרובוקטיבי בו הוא תיאר כיצד קבוצה של אנשי מקצוע פנתה למספר מרכזי אישפוז בארה"ב בבקשה להתאשפז.  במהלך האינטייק הם ענו בכנות על עצמם למעט שני דברים:  א.  הם נתנו שמות בדויים.  ב.  הם דיווחו ששמעו קול שאומר "EMPTY" או "THUD".  כולם אושפזו עם דיאגנוזה של סכיזופרניה.  זמן האישפוז הממוצע שלהם היה 19 יום.  כאשר הם שוחררו, רובם קיבלו אבחנה של "סכיזופרניה ברמיסיה" (סכיזופרניה בהפוגה).  רוזנהן וחבריו ציינו שמרבית הפציינטים המאושפזים הבחינו שאנשי המקצוע שאושפזו אינם חולים אמיתיים, אבל איש מאנשי הצוות לא הבחין בכך.  רוזנהן הסיק שהמוסדות האישפוזיים לא הצליחו להבחין בין חולי נפש לאנשים בריאים.  מחקרו עורר תגובות סוערות. 

מחקרים אלה ודומים להם השפיעו על יצירת ה – DSM3

DSM3

פרסום ה - DSM3 ב – 1980 היה חלק משינוי פרדיגמטי בפסיכיאטריה ובתחום בריאות הנפש בכללותו.  לפני ה – DSM3, הפסיכיאטריה נשלטה על ידי פסיכיאטרים שהושפעו מהפסיכואנליזה.  קבלת אבחנה פסיכיאטרית לא נחשבה בעיניהם כחשובה במיוחד או כמשפיעה על הפסיכותרפיה שהפציינטים שלהם קיבלו.   לעומתם, המחברים הראשיים של ה- DSM3  היו קבוצת פסיכיאטרים שניסו להחזיר את הפסיכיאטריה לשורשים הרפואיים שלה.  רעיונותיהם התאימו למעבר מהתמקדות בפסיכותרפיה להתמקדות בשימוש בתרופות. 

הרצון של מחברי ה – DSM3 להשמיט את מושג הנוירוזה עורר מחלוקת רבה בין אנשי המקצוע.  בסופו של דבר הגיעו לפשרה בה המלה "נוירוזה" הופיעה בסוגריים לאחר המפה "הפרעה" במקרים מסויימים.  

SPITZER , שהוביל את עריכת ה –  DSM3 והועדה המארגנת ל – DSM3 עשו צעד אמיץ והציעו הגדרה טנטטיבית של המושג של הפרעה נפשית.  הם נזקקו להגדרה זו מכיוון שמטרה מפורשת של יוצרי ה – DSM3 היתה להימנע מספקולציות על הגורמים לפסיכופתולוגיה (ובמיוחד ממושגים פסיכואנליטים).  הגדרה זו היתה גם בניגוד ישיר לתנועת האנטי פסיכיאטריה שניסתה להגדיר מחלת נפש כדרך של החברה להתמודד עם אנשים בלתי רצויים – על ידי כך שהיא שמה עליהם תוית של מחלת נפש כדי לשמור אותם שקטים ומופרדים מהחברה ה"נורמטיבית". 

ההגדרה של ה – DSM3 היתה כזו:  הפרעה נפשית היא תסמונת או דפוס התנהגותי או פסיכולוגי בעל חשיבות קלינית המופיע באנשים והקשור בדרך כלל לסימפטום כואב (מצוקה) או לפגיעה בתחום חשוב אחד או יותר של התפקוד (לקות).  בנוסף לכך, ניתן להסיק שקיימת דיספונקציה התנהגותית, פסיכולוגית או ביולוגית, ושההפרעה היא לא רק ביחסים בין האדם לחברה.

ובאנגלית:

Each of the mental disorders is conceptualized as a clinically significant behavioral or psychological syndrome or pattern that occurs in an individual and that is typically associated with either a painful symptom (distress) or impairment in one or more important areas of functioning (disability). In addition, there is an inference that there is a behavioral, psychological, or biological dysfunction, and that the disturbance is not only in the relationship between the individual and the society.

הגדרה זו של הפרעה נפשית הובילה לדיון מעניין ורחב של פילוסופים, פסיכולוגים קוגניטיבים, אנתרופולוגים חברתיים והיסטוריונים על אודות סיווגים פסיכיאטרים.

ה – DSM3  הכיל קריטריונים דיאגנוסטים בנוסף לתיאורי פרוזה.  כל קטגוריה של הפרעת נפש כללה תיאור של פרופיל דמוגרפי טיפוסי של פציינטים שסבלו מהפרעה זו, הסבר על משמעות הקטגוריה, תיאור כיצד לבצע אבחנה מבדלת, ודיון קצר במה שהיה ידוע על מהלך ההפרעה ועל הופעתה.  בנוסף לכך היו חמישה "צירים", שכל פציינט מוקם בכל אחד מהם:  א.  הקטגוריה של ההפרעה הנפשית.  ב.  הפרעת אישיות ו/או הפרעה אינטלקטואלית  ג.  בעיות רפואיות רלוונטיות לסימפטומים הפסיכיאטרים.  ד.  גורמי לחץ פסיכוסוציאלים בסביבה של הפציינט.  ה.  רמת התפקוד המסתגל הגבוהה ביותר של הפציינט בשנה האחרונה.  

לאחר פרסום ה – DSM3 חל חידוש נוסף בפסיכיאטריה – הראיון המובנה, שפותח לראשונה על ידי SPITZER וחבריו.  עד לשנת 2000 היו מעל 240 סוגים של ראיונות מובנים שבדקו היבטים שונים של פסיכופתולוגיה והפרעות נפשיות.  המהימנות של הערכה דיאגנוסטית באמצעות כלים אלה היתה טובה הרבה יותר מכפי שהיתה קודם.  בעוד שבאמצעות שימוש ב – DSM3 בלבד, מידת ההסכמה בין קלינאים היתה בין 0.4 ל – 0.6,  באמצעות שימוש בראיון מובנה, המהימנות עלתה לטווחים של 0.75-0.90.  הראיונות המובנים הפכו לדומיננטים במחקר למרות שהשימוש הקליני בהם הוא מועט גם כיום.

עד כמה מבנה ה – DSM3  תואם את הדרך ה"טבעית" בה קלינאים מאבחנים?

בשנת 1980 החוקר קנטור וחבריו ביקשו מ – 13 קלינאים לרשום מאפיינים שהם מקשרים עם תשע קטגוריות דיאגנוסטיות של פסיכוזה מה – DSM2.  כל מאפיין שנבחר על ידי שלושה קלינאים לפחות נשמר לרשימה הסופית.  לאחר מכן החוקרים לקחו 12 תיאורי מקרים של פציינטים שקיבלו אחת מארבע אבחנות פסיכוטיות (סכיזופרניה מאנית, סכיזופרניה דכאונית, סכיזופרניה פרנואידית וסכיזופרניה בלתי מובחנת).  ארבעה מתיאורי המקרים נחשבו לדי טיפוסיים לארבע האבחנות (הם הכילו כמעט את כל המאפיינים ש – 13 הקלינאים כתבו), ארבעה היו פרוטוטיפים באופן בינוני, וארבעה היו לא טיפוסיים (כלומר, היו להם ארבעה או פחות מהמאפיינים שכתבו 13 הקלינאים).  נתנו את תיאורי המקרים הללו לקלינאים וביקשו מהם לתת להם אבחנות.  מהימנות האבחנה ביתה תלויה במידת הפרוטוטיפיות של המקרה:  ככל שהמקרה היה יותר טיפוסי, פרוטוטיפי, הוא אובחן באופן מהימן יותר.  

ממחקר זה וממחקרים אחרים הסיקו שקלינאים לא משתמשים בקריטריונים דיאגנוסטים כדי לאבחן.  הם נוטים לאבחן באמצעות השוואת מאפייני הפציינט לפרוטוטיפ של אותה תסמונת או הפרעה, ולא באמצעות בדיקה אם הפציינט עונה לקריטריונים של אותה הפרעה.  זה דומה לדרך בה ילדים לומדים מושגים קטגוריאלים – באמצעות השוואתם לפרוטוטיפ.  וזה אומר שארגון ה – DSM לפי תיאורים פרוטוטיפים ולא לפי קריטריונים מתאים יותר לדרך הטבעית בה קלינאים עובדים. 
 .
DSM3R

DSM-III-R    שפורסם ב – 1987 לא היה שונה במבנהו מה – DSM3 (גם הוא הכיל כמה צירים, שימוש בקריטריונים דיאגנוסטים וההפרעות הנפשיות אורגנו בו באופן דומה).  אך הוא כלל קטגוריות חדשות.

בעוד שלאחר פרסום ה – DSM3 המאבק הפוליטי התמקד בין הקבוצה הפסיכואנליטית באגודה הפסיכיאטרית האמריקנית לבין הקבוצה בעלת האוריינטציה הביולוגית, כאשר נוצר ה – DSM3R המאבק הפוליטי לבש פנים אחרות.  פמיניסטיות הביעו דאגה על כך שעמדו לכלול ב- DSM קטגוריות כמו סינדרום טרום-וסתי והפרעת אישיות מזוכיסטית.  כתוצאה מהמחלוקת ה – DSM3R הוסיף נספח שנקרא "קטגוריות דיאגנוסטיות מוצעות שדורשות מחקר נוסף".  בנספח זה נכללו שלוש קטגוריות: LATE LUTEAL PHASE DYSPHORIC DISORDER (שם חדש לסינדרום טרום-וסתי),  הפרעת אישיות סדיסטית (כדי לאזן את הפרעת האישיות המזוכיסטית), והפרעת אישיות המביסה את עצמה    SELF DEFEATING    (שם חדש להפרעת אישיות מזוכיסטית). 

בשנת 1992, WAKEFIELD הסב את תשומת הלב לכך שההגדרה של הפרעה נפשית מבוססת על שיפוט ערכי.  אותם סימפטומים עשויים להישפט כהפרעה נפשית בהקשר אחד ולא בהקשר אחר.  כאשר ערכים חברתיים משתנים במשך הזמן, מצבים מסויימים שנחשבו קודם להפרעות נפשיות כבר לא ייחשבו לכאלה (למשל, הומוסקסואליות).  מצבים אחרים שלא נחשבו להפרעות נפשיות עשויים להפוך לכאלה (למשל, שימוש באינטרנט).  כך, לעולם לא תהיה גירסה סופית של ה – DSM
  

מהדורות ההמשך של ה – DSM יסקרו בקצרה בפוסט הבא.


Tuesday, June 14, 2016

תשובתו של פרופ' מקגיל לפוסט מתאריך 11 ביוני : "האם הגדרת לקות למידה על פי CHC עובדת?"

   

יש לי הכבוד לפרסם את תשובתו של פרופ' מקגיל לפוסט שכתבתי על מאמר שלו.  הפוסט נכתב בתאריך 11 ביוני ואפשר לקרוא אותו למטה או בלחיצה כאן.
  

Prof. McGill's response to my June 11th post: "When Theory Trumps Science: a Critique of the PSW Model for SLD Identification"

I have the honor to present Prof. McGill's comments on this post:


Smadar,

Thank you for your interest and overview of our paper. I have been following your blog for several years now and always enjoy your thoughtful posts on intelligence and cognition.

Our commentary was somewhat brief and is conceptually similar to a more substantive paper that was recently published in Learning Disability Quarterly that was more focused on potential measurement issues with the PSW model:

McGill, R. J., Styck, K. S., Palomares, R. S., & Hass, M. R. (2015). Critical issues in specific learning disability identification: What we need to know about the PSW model. Learning Disability Quarterly. Advance online publication. doi: 10.1177/0731948715618504

The point you raise about diagnostic validity studies and LD is a good one and an issue we raised in the LDQ paper as there is no “gold standard” for SLD diagnoses, as a consequence, we have no way of knowing who truly has SLD and thus the results from SLD DV studies will always have this limitation. While I still think they are of some value as they give us some estimate of the potential DV of identification models, this limitation must always be considered when interpreting those results.

I also concur that the simulation studies conducted by Steubing et al. and the Kranzler study have limitations (as all studies do). Most germane, the fidelity in which the authors attempted to model various PSW implementations. With all the potential permutations, these models are incredibly complex which renders them difficult to conceptualize without access to significant sources of clinical assessment data (really hard to obtain). Even if you have the data, simulating the multi-step decision-making that these models require of clinicians is the biggest hurdle that a researcher faces and one that so far has not been able to be overcome. I am hopeful that with advent of machine learning algorithms in advanced statistical software programs such as R, one day we may be able to model this stuff better.

What I always want to stress when discussing these things is that I am not anti-PSW, I actually think it makes a lot of conceptual sense. I don’t really have a dog in this fight. My major concern with the model has to do with the tools that we are using to make decisions within these models (i.e., IQ tests). While I love IQ tests and think they are all very good at estimating overall cognitive ability, I do think that they have significant limitations when we try to get more than that out of them. As an example, the results from numerous independent factor analytic studies raise questions about the viability of publisher suggested measurement models. Most pertinent, do these instruments measure lower-order abilities well if at all? What we consistently find is that g dominates all levels of IQ tests and the scores provided by those instruments and when this source of variance is accounted for, there is often only a small proportion of reliable variance attributable to the lower-order abilities (e.g., auditory processing, visual processing, etc.) that are of most interest to clinicians and the focus of clinical interpretation in PSW models. These are significant confounds as we use these models as the basis for clinical interpretation of scores. In my opinion, there has been insufficient discussion of these issues and their potential impact on clinical decision-making…especially for those using and advocating the PSW model.

My opinion is that we are not measuring these constructs very well (not that they don’t exist) and that is why we have the issues with long-term stability and incremental prediction of achievement. Of course, the issues with cognitive profile analysis (regardless of the level of the scores) have long been known (see Canivez, 2013; Glutting, Watkins & Youngstrom, 2003; Watkins, 2000) and that is all PSW really is….profile analysis at the factor score level rather than subtest level. Scatter and variability are endemic in the population. As an example, my analyses indicate that over 30% of the KABC-II normative sample have at least a 23 points difference between their highest and lowest factor scores. That’s a lot of noise that PSW models will have to sift through in order to find the “signal.” To be fair, this is something that Flanagan and colleagues have repeatedly discussed in their writings.

In sum, these are complex issues that we think clinicians need to be aware as it relates to the PSW model and SLD identification in general. Perhaps these limitations will be overcome in the future however presently my opinion is that we need to know more before we utilize these models to make important diagnostic and treatment decisions in practice.

As you rightly note, in general I advocate more circumspect interpretation of IQ tests. Whereas, the corpus of the empirical literature indicates that one can interpret FSIQ with confidence, significant questions remain as one moves to lower levels of dimensionality. As we indicate in our paper, if one uses these scores within a diagnostic decision-making model (i.e., PSW), their shortcomings will be encapsulated in those models and render consistent and defensible decision-making very difficult. 

In spite of this, I do not advocate a return to the flawed discrepancy model. I advocate using FSIQ as a rule out element within the broader conceptual definition of LD (i.e., unexpected underachievement). If a kiddo has low average or higher ability than I can deduce that is probably not the reason for their underachievement and thus LD or some other condition is a more viable explanation. This requires thinking about performance on IQ tests more from a criterion-based perspective which when you think about it is the level of precision with which we measure functioning on these instruments (think confidence intervals), that’s really all we can get out of them anyway. As a colleague of mine says, we are trying to measure really complex aspects of cognition with what are virtually stone tools. When it comes to additional assessment, I stipulate that we need more than CBM and other related achievement data but it remains to be seen whether the use of multi-factored cognitive batteries and hours and hours of additional assessment is indeed the answer. Nevertheless, I think certain dimensions are more important than others (e.g., working memory and processing speed).

We have been trying to figure out how to validly diagnose LD for a long time now. I don’t profess to have the answer. As previously mentioned, this is a complex issue that we have been attempting to adjudicate for a long time now. Unfortunately, proposed remedies have consistently been found wanting once they have been implemented. Perhaps we would be better served if we quit this quixotic quest to diagnose an illusive construct and just figured out which kids need help and get to helping. 

      

Sunday, June 12, 2016

הגדרות היכולות הרחבות והצרות על שני דפים





McGrew חיבר מסמך חדש, קצר וצפוף ובו הגדרות יכולות ה – CHC הרחבות והצרות על שני דפים.

When Theory Trumps Science: a Critique of the PSW Model for SLD Identification

McGill, R. J., & Busse, R. T. (2016). When Theory Trumps Science: a Critique of the PSW Model for SLD Identification. Contemporary School Psychology, 1-9.

In this paper, McGill & Busse criticize the use of PSW operational definitions of learning disabilities (such as the Flanagan/CHC operational definition).

An operational definition does not define the concept itself (an operational definition of learning disability (LD) does not define the essence of LD).  Rather, it portrays the way in which the concept is measured (what we should actually do in order to determine whether a child is learning disabled or not).

Each operational definition of learning disability has advantages and disadvantages, and we can find papers criticizing each definition.  It's important to be familiar with the critique.  The mere existence of critique does not make the use of an operational definition wrong.  It's important to weigh each definition's critique and to choose the definition with the least heavy- weight critique.

Now we turn to the paper, with remarks and explanations by me in green.  

Within the professional literature, there is growing support for educational agencies to adopt an approach to SLD identification that emphasizes the
importance of an individual’s pattern of cognitive and achievement strengths and weaknesses (PSW).  The Flanagan/CHC definition of learning disability is one of these operational definition.  Cognitive strengths and weaknesses can manifest in CHC abilities such as Fluid Ability, Short Term Memory, Long Term Storage and Retrieval, Visual Processing, Auditory Processing, Processing Speed, Comprehension Knowledge.  Achievement strengths and weaknesses can manifest in tests measuring Reading Decoding/Reading Comprehension/Writing/Written Expression/Math Calculations/Math Reasoning.

The Flanagan definition is based on five major criteria:  A. Significantly poor performance on achievement tests. B.  One or more of the CHC cognitive abilities is significantly below average.  C.  Concordance/linkage between the poor area of achievement and the low cognitive ability.  The low cognitive ability can explain/be the cause of the poor achievement area.  D. Other cognitive abilities are intact.  E. Exclusionary factors are not the primary reason for the poor performance on the achievement tests.

In 2014, the California Association of School Psychologists released a  position paper endorsing this approach.  As a vehicle for examining the PSW model, the authors respond critically to three fundamental positions taken in the position paper: (a) diagnostic validity for the model has been established; (b) cognitive profile analysis is valid and reliable; and (c) PSW data have adequate treatment utility. The authors conclude that at the present time there is insufficient support within the empirical literature to support adoption of the PSW method for SLD identification.

Prior to IDEA (INDIVIDUALS WITH DISABILITIES EDUCATION ACT) 2004, federal regulations emphasized the primacy of the discrepancy model, wherein Specific Learning Disability was operationalized as a significant discrepancy between an individual’s achievement and their cognitive ability (Full Scale IQ score). This model was heavily criticized.  First, it causes "waiting for failure", because such a discrepancy can be proven only in third grade, when the child "achieves" a two-year discrepancy between his IQ score and his performance in reading/writing/math.  Second, it causes under-identification of learning disabilities in adolescence.  LD affects full scale IQ (for instance, a learning disabled child reads less, thus his comprehension knowledge is less developed, which lowers his IQ).  Thus in adolescence there is a lowered chance for a learning disabled child to have a discrepancy between his IQ score and his achievement scores.  This renders it impossible to diagnose him as learning disabled despite him being so.

In contrast to previous legislation, IDEA 2004 permitted local educational agencies the option of selecting between the discrepancy method and alternatives such as response-to intervention (RTI). The RTI model defined learning disability as persistent low achievement (in reading/writing/math) despite adequate intervention. This model permits but does not require to conduct a psychological assessment  for a differential diagnosis with intellectual disability, language disorder and emotional /behavioral disorders.  This point (the lack of differential diagnosis) came under a lot of criticism.  Over the last decade, RTI has been widely embraced within the technical literature and adopted as an SLD classification model by many educational agencies across the country, resulting in renewed concern regarding the validity of identification approaches that deemphasize the role of
cognitive testing.

PSW models were developed in response to these problems.  There are several such models which are quite similar to each other:  (a) the concordance/discordance model (C/DM; Hale and Fiorello 2004), (b) the
Cattell-Horn-Carroll operational model (CHC; Flanagan et al. 2011), and (c) the discrepancy/consistency model (D/CM; Naglieri 2011). It is noteworthy that, although the models differ with respect to their theoretical orientations
and the statistical formulae used to identify patterns of strengths and weakness, all three PSW models share at least three core assumptions as related to the diagnosis of Specific LD: (a) evidence of cognitive weaknesses must be present, (b) an academic weakness must also be established, and (c) there must be evidence of  spared (i.e., not indicative of a weakness)
cognitive-achievement abilities.  The authors go on to briefly discuss each model.  This will not be done here.

Now the authors criticize PSW models on three points:

Critical Assumption One: Diagnostic Validity for the Model Has Been Established

Steubing et al. (2012) investigated the diagnostic accuracy of several PSW models and reported high diagnostic specificity (a high percentage of children who do not have LD and are correctly identified as not having LD) across all
models. However, the models had low to moderate sensitivity (a low to moderate percentage of children who have LD and are identified as having LD).  Only a very small percentage of the population (1%-2%) met criteria for specific learning disabilities using these models.  Clinically, it feels like the CHC definition does lower the percentage of children identified as LD, but other definitions, like DSM5, inflate the percentage of children identified as LD.  We have no way of knowing what is the "real" percentage of LD in the population, because every study that assesses this percentage is conducted in light of some operational definition, usually a relatively "inflating" one.

Since there are no objective criteria with which we can know who is really learning disabled, Steubing et al cannot argue that "these models had low to moderate sensitivity".  The only thing that Steubing et al can say is that PSW models identify less children as learning disabled than other models.  But we cannot know whether this fact makes PSW models better or worse at identifying LD.   

Kranzler et al. (2016) examined the broad cognitive abilities of the Cattell-Horn-Carroll theory held to be meaningfully related to basic reading, reading comprehension, mathematics calculation, and mathematics reasoning across age groups. Results of analyses of 300 participants in three age groups (6–8, 9–13, and 14–19 years) indicated that the XBA method (a method for implementing the PSW approach) is very reliable and accurate in detecting true negatives (the percentage of children who are not LD and are correctly identified as not having LD).  The model identified 92% of children who were not LD as not having LD.  However Kranzler et al found the model to have quite low sensitivity, indicating that this method is very poor at detecting the percentage of LD children who are correctly identified as having LD.  Only 21% of children who were LD were identified as having LD according to this model.

 A brief peek at Kranzler et al's study makes me wonder at the way Kranzler et al determined which abilities were related to which areas of achievement.  For example, at age 6-8, Kranzler et al considered the broad abilities Comprehension Knowledge, Long Term Storage and Retrieval, Processing Speed and Short Term Memory as related to basic reading skills.  Kranzler et al say they took these links out of McGrew and Wendling's 2010 study.  But that study (presented on slide no. 11 in the second presentation of the Intelligence and Cognitive Abilities presentation series on the right hand column of this blog) found that the narrow ability Phonological Coding is also related to basic reading skills at age 6-8! Generally, McGrew and Wendling recommend using combinations of broad and narrow abilities to predict performance in different areas of achievement.  Kranzler et al used only broad abilities.

Thus it may be that the low sensitivity found in Kranzler et al's study results from the fact that Kranzler et al did not implement McGrew and Wendling's findings accurately (I have to say this very carefully since I did not read the entire Kranzler paper).
 
Critical Assumption Two: Cognitive Profile Analysis Is Reliable and Valid

In order to identify LD according to the Flanagan method we have to use cognitive ability/index scores.  If they are not reliable – the identification of LD will not be reliable as well.   

Significant questions have been raised about the long term stability and structural and incremental validity of factor level measures from intelligence.  Structural validity investigations using exploratory factor analysis have revealed conflicting factor structures from those reported in the technical manuals of contemporary cognitive measures  which indicates that these instruments may be overfactored (Frazier and Youngstrom 2007).  Additionally, the long-term stability and diagnostic utility of these indices has been found wanting.  The authors cite a study by Watkins and Smith (2013) who investigated the long-term stability of the  WISC-IV  with a sample of 344 students twice evaluated for special education eligibility at an average interval of 2.84 years. Test-retest reliability coefficients for the Verbal Comprehension Index (VCI), Perceptual Reasoning Index (PRI), Working Memory Index (WMI), Processing Speed Index (PSI), and the Full Scale IQ (FSIQ) were .72, .76, .66, .65, and .82, respectively.  As far as I know, good reliability is considered to be above 0.7.  Thus the WMI and the PSI were found to be not reliable in this research and the VCI and PRI had low reliability.   However, 25% of the students earned FSIQ scores that differed by 10 or more points, and 29%, 39%, 37%, and 44% of the students earned VCI, PRI, WMI, and PSI scores, respectively, that varied by 10 or more points. Given this variability, Watkins and Smith argue that it cannot be assumed that WISC-IV scores will be consistent across long test-retest intervals for individual students.

In light of this study we ought to give up using index scores and use only the FSIQ score as the lesser evil.  In the context of LD this brings us back to the discrepancy model.

However it's possible that during the 2.84 years these children received special education services their cognitive abilities were improved.  This can explain the unstableness of the indices in this study.  Does this unstableness exist in the general population?  In other intelligence tests?

Critical Assumption Three: PSW Methods Have Adequate Treatment Utility

Despite many attempts to validate group by treatment interactions, the efficacy of interventions focused on cognitive deficits remains speculative and unprovenParticularly noteworthy, are the findings obtained from a recent meta-analysis of the efficacy of academic interventions derived from neuropsychological assessment data by Burns et al. (2016). In contrast to the effects attributed to more direct measures of academic skill, it was found that the effects of interventions developed from cognitive data were consistently
small (g=0.17). As a result, Burns et al. (2016) concluded, "the current and previous data indicate that measures of cognitive abilities have little to no [emphasis added] utility in screening or planning interventions for reading and mathematics".

Do other operational definitions of learning disability lead to more efficient interventions?

To summarize, McGill & Busse criticize the CHC/PSW LD definition in three ways:  A.  the diagnostic validity of the model is weak.  B. index scores are not reliable and valid.  C. the model does not lead to efficient interventions.


Are you convinced?  

Saturday, June 11, 2016

האם הגדרת לקות למידה על פי CHC עובדת?

האם הגדרת לקות למידה על פי CHC עובדת?

McGill, R. J., & Busse, R. T. (2016). When Theory Trumps Science: a Critique of the PSW Model for SLD Identification. Contemporary School Psychology, 1-9.

McGill & Busse מותחים במאמר זה ביקורת על השימוש בהגדרה האופרציונלית ללקות למידה בשיטה של FLANAGAN

הגדרה אופרציונלית אינה מגדירה את מהות התופעה (הגדרה אופרציונלית של לקות למידה לא מגדירה מה זה לקות למידה) אלא מאפיינת את הדרך בה מודדים את התופעה (מה עושים בפועל כדי לקבוע אם הילד לקוי למידה או לא). 

ראוי לציין, שלכל הגדרה אופרציונלית של לקות למידה יש יתרונות וחסרונות, ועל כל הגדרה אופרציונלית ניתן למצוא מאמרי ביקורת.  חשוב להכיר את הביקורות השונות.  עצם קיומם של מאמרי ביקורת עדיין לא פוסל את השימוש בהגדרה מסויימת.  רצוי שנדע לשקול את הביקורות על ההגדרות השונות ולבחור בהגדרה שטיעוני הביקורת עליה פחות כבדי משקל.

מכאן רשות הדיבור למאמר, עם תוספות, הערות והסברים שלי (בצבע ירוק):  

בספרות המקצועית ניכרת תמיכה הולכת וגוברת באימוץ הגדרה אופרציונלית ללקויות למידה המדגישה את חשיבותן של חוזקות וחולשות בתחומי הקוגניציה ובתחומי ההישג של האדם.   הגדרת לקות למידה לפי /CHC הגדרת FLANAGAN             היא אחת מההגדרות האופרציונליות הללו.  תחומי הקוגניציה הם, למשל, יכולות ה – CHC:  יכולת פלואידית, זיכרון לטווח קצר, אחסון ושליפה לטווח ארוך, עיבוד חזותי, עיבוד שמיעתי, מהירות עיבוד, ידע מגובש.  תחומי ההישג הם הביצוע של הילד במבחנים הבודקים קריאה טכנית/הבנת הנקרא/כתיבה טכנית/הבעה בכתב/חישוב מתמטי/חשיבה מתמטית. 

הגדרות אלה נקראות בשם הכללי PSW – PATTERNS OF STRENGTHS AND WEAKNESSESכאמור הגדרת לקות למידה לפי FLANAGAN היא אחת הגישות הללו.  הנה תזכורת מהירה להגדרה:  ההגדרה מבוססת על חמישה שלבים עיקריים:   א.  הנמכה מובהקת באחד או יותר מתחומי ההישג.  ב.  הנמכה מובהקת באחת או יותר מהיכולות הקוגניטיביות.  ג.  התאמה בין תחום ההישג המונמך לבין היכולת הקוגניטיבית המונמכת, באופן שהיכולת הקוגניטיבית המונמכת תסביר את ההנמכה בהישג.  ד.  מרבית היכולות הקוגניטיביות תקינות.  ה.  גורמי הדרה אינם סיבה עיקרית להנמכה בתחומי ההישג.   

בשנת 2014 האגודה של הפסיכולוגים החינוכיים בקליפורניה פרסמה נייר עמדה בו היא תומכת בשימוש בהגדרות מסוג זה.  מחברי המאמר רוצים לבחון שלוש עמדות שננקטו בנייר העמדה של האגודה של הפסיכולוגים החינוכיים בקליפורניה:  א.  שיש עדות לתוקף הדיאגנוסטי של הגדרות על פי מודל PSW.  ב.  שניתוח פרופיל קוגניטיבי הוא תקף ומהימן.  ג.  שניתן לתכנן התערבויות מתאימות לפי המודל.  מחברי המאמר מסיקים, שבנקודת הזמן הנוכחית, אין מספיק תמיכה בספרות המקצועית באימוץ מודל ה – PSW להגדרת לקות למידה.

החוק האמריקני לפני  2004  הגדיר לקות למידה לפי מודל הפערים.   על פי מודל זה, לקות למידה היא פער מובהק בין ההישגים של הילד בקריאה/כתיבה/חשבון לבין רמת המשכל הכללית שלו.  מודל זה ספג ביקורת רבה והתעורר צורך לשנותו.  ראשית, הוא גרם להמתנה לכשלון, מכיוון שניתן היה להוכיח פער כזה רק בכיתה ג', בה הצטבר אצל הילד פער של שנתיים בין רמת המשכל הכללית לבין רמת ההישגים בקריאה/כתיבה/חשבון.  שנית, הוא גרם לתת-אבחון של לקות למידה בגיל ההתבגרות.  זאת מכיוון שלקות למידה משפיעה על רמת המשכל הכללית (למשל, ילד לקוי למידה ממעט לקרוא.  כתוצאה מכך הידע המגובש שלו מתפתח פחות, מה שמוריד את רמת המשכל הכללית).  כך בגיל ההתבגרות מצטמצם הסיכוי שיהיה לילד פער בין רמת המשכל הכללית לבין תחומי ההישג, ואז לא נוכל להגדירו כלקוי למידה למרות שהוא כזה.

IDEA- INDIVIDUALS WITH DISABILITIES EDUCATION ACT , החוק האמריקני ששונה בשנת 2004 , אפשר למדינות לבחור בין מודל הפערים לבין מודלים חלופיים כמו מודל תגובה להתערבות.  מודל תגובה להתערבות הגדיר לקות למידה כהנמכה מתמשכת בתחומי ההישג (קריאה/כתיבה/חשבון) למרות קבלה של הוראה מתקנת מתאימה.  מודל זה מאפשר אך אינו מחייב לבצע הערכה פסיכולוגית כדי לקבוע אבחנה מבדלת עם פיגור שכלי, הפרעת שפה וקשיים רגשיים – התנהגותיים, ומכאן הביקורת עליו (חוסר הדרישה לאבחנה מבדלת).  אימוץ מודל תגובה להתערבות באופן נרחב במהלך העשור הקודם גרם לדאגה לגבי התוקף של הגדרה אופרציונלית זו שמיתרת את השימוש במבחנים קוגניטיבים ובמציאת קישור בין ההנמכה בהישג לבין הנמכה קוגניטיבית שעומדת בבסיסו. 

כתוצאה מבעיות אלה בהגדרות לקות למידה, התפתחו מודלים של PSW.  יש כיום מספר מודלים כאלה שהם מאד דומים זה לזה:  א.  מודל ההתאמה/אי התאמה CONCORDANCE/DISCORDANCE של HALE AND FIORELLO משנת 2004, ב.  הגדרה אופרציונלית לפי CHC שפיתחה FLANAGAN בשנת 2011, ו-  ג.  מודל הפער/עקביות DISCREPANCY/CONSISTENCY  של NAGLIERI  (אחד ממפתחי תאורית ה – PASS שקשורה למבחן הקאופמן).  לכל שלושת המודלים שלוש הנחות בסיסיות משותפות לגבי הגדרת לקות למידה:  א.  חייבת להיות עדות לחולשה קוגניטיבית.  ב.  חייבת להיות עדות לחולשה באחד מתחומי ההישג.  ג.  חייבת להיות עדות ליכולות קוגניטיביות תקינות המתקשרות לתחומי הישג תקינים.  המאמר מפרט מעט על כל אחד מהמודלים אך בשל הדמיון הרב ביניהם לא ניכנס לזה כאן.

כעת המאמר מנסה למתוח ביקורת על מודל ה – PSW (שכאמור הגדרת FLANAGAN היא חלק ממנו)

הנחה ראשונה:  יש תוקף דיאגנוסטי למודל PSW

Steubing et al.    חקרו ב - 2012 את הדיוק הדיאגנוסטי במודלים השונים של PSW .  הם מצאו שלמודלים אלה יש ספציפיות מעולה (אחוז האנשים שאינם לקויים ואכן זוהו כלא לקויים היה גבוה בכל המודלים).  לעומת זאת למודלים אלה היתה רגישות נמוכה עד בינונית (אחוז האנשים שהיו לקויי למידה ואכן זוהו כלקויי למידה היה נמוך עד בינוני במודלים אלה).  רק אחוז נמוך של האוכלוסיה, 1-2%, עמד בקריטריונים ללקות למידה ספציפית על פי מודלים אלה.  בתחושה הקלינית ההגדרה על פי CHC אכן מצמצמת את אחוז הילדים המוגדרים כלקויים, אולם הגדרות אחרות, כדוגמת DSM5 מנפחות את אחוז הילדים המוגדרים כלקויים.  יש לציין כי אין לנו שום דרך לדעת מהו האחוז ה"אמיתי" של לקויי למידה באוכלוסיה, שכן כל מחקר שבדק זאת נערך על פי הגדרה אופרציונלית כלשהי, ובדרך כלל על פי הגדרות "מנפחות". 


מכיוון שלא קיימים קריטריונים אובייקטיבים בעזרתם אנחנו יכולים לדעת מי באמת לקוי למידה,  סטובינג וחבריו לא יכולים לטעון, כפי שהם טוענים, ש"למודלים אלה היתה רגישות נמוכה עד בינונית" (אחוז האנשים שהיו לקויי למידה ואכן זוהו כלקויי למידה היה נמוך עד בינוני במודלים אלה).  הדבר היחיד שסטובינג וחבריו יכולים לטעון הוא שמודלים של PSW מזהים פחות אנשים כלקויי למידה מאשר מודלים אחרים.  אבל אנחנו לא יכולים לדעת אם עובדה זו הופכת מודלים שלPSW   לטובים יותר להגדרת לקות למידה או לא.

Kranzler et al.     בחנו ב - 2016 את תוקף ההגדרה בשיטת PSW תוך שימוש מנתונים של הוודקוק ג'ונסון 3 (מבחן המשכל החדש של ישראל).  הוא בדק את היכולות הקוגניטיביות הרחבות במודל ה-  CHC שנחשבו כקשורות לקריאה בסיסית, הבנת הנקרא, חישוב מתמטי וחשיבה מתמטית ב – 300 ילדים בגילאים 6-8, 9-13 ו – 14-19.  בדומה לסטובינג, הוא מצא, שההגדרה האופרציונלית על פי CHC היא מאד מהימנה ומדויקת בזיהוי אנשים שאינם לקויים ככאלה (כלא לקויים).  המודל זיהה 92% מהאנשים שלא היו לקויים כלא לקויים.  אבל החוקרים מצאו שלמודל יש רגישות נמוכה, כלומר אחוז הילדים שהם אכן לקויי למידה וזוהו כלקויי למידה היה נמוך.  רק 21% מהילדים שהיו לקויי למידה זוהו ככאלה לפי מודל זה. 

הצצה שטחית במחקר של קרנצלר וחבריו מעוררת אצלי השגות על הדרך בה קבעו אילו יכולות קוגניטיביות קשורות לאילו תחומי הישג.  למשל, בגיל 6-8 קרנצלר וחבריו בדקו את היכולות הרחבות ידע מגובש, אחסון ושליפה לטווח ארוך, מהירות עיבוד וזיכרון לטווח קצר כקשורות לכישורי קריאה בסיסיים.  קרנצלר וחבריו טוענים שהם לקחו את הקישורים בין היכולות לתחומי ההישג מממצאי מחקר של MCGREW  ו – WENDLING מ- 2010.   אבל במחקר זה (המוצג בשקופית 11 במצגת השניה בסדרת המצגות על ה – CHC בטור הימני של הבלוג) גם היכולת הצרה קידוד פונולוגי בתוך היכולת הרחבה עיבוד שמיעתי נכללה כאחד הכישורים הקשורים לקריאה בסיסית בגילאים אלה!

לכן ייתכן שהרגישות הנמוכה כביכול שהתגלתה במחקר קרנצלר במודל הגדרת לקות למידה על פי CHC נובעת מכך שקרנצלר לא יישם את הממצאים של MCGREW ו – WENDLING  באופן מדויק (אני חייבת לסייג מסקנה זו עד שאקרא בעיון את מחקר קרנצלר).   

 הנחה שניה:  ניתוח פרופיל קוגניטיבי הוא מהימן ותקף

כדי לקבוע לקות למידה על פי CHC  אנחנו צריכים להשתמש בציוני היכולות הקוגניטיביות השונות/האינדקסים השונים במבחני המשכל.  אם הם אינם מהימנים – קביעת לקות למידה לא תהיה מהימנה.   

המחברים מעלים שאלות על היציבות לאורך זמן והתוקף של אינדקסים ממבחני משכל.    חוקרים מסוימים כמו   Frazier and
 Youngstrom 2007 ביצעו ניתוח גורמים על מבחני המשכל באופן שונה מהמקובל ומצאו שמבחני משכל מסוימים בודקים פחות גורמים מכפי שהם טוענים שהם בודקים.   בנוסף, המחברים טוענים שהיציבות של האינדקסים במבחני המשכל לא מספיק טובה.   המחברים מצטטים מחקר של Watkins and Smith 2013 שבדקו את היציבות של מבחן WISC4.  הם השתמשו במדגם של 344 תלמידים שעברו את המבחן פעמיים כחלק מאבחון זכאותם לחינוך מיוחד.  פער הזמן בין שתי ההעברות היה 2.84 שנים.  מהימנות מבחן חוזר של אינדקס הבנה מילולית היתה 0.72, של אינדקס היסק תפיסתי 0.76, אינדקס זיכרון עבודה 0.66, אינדקס מהירות עיבוד 0.65 ושל מנת המשכל הכללית 0.82.  למיטב ידיעתי, מהימנות טובה נחשבת למעל 0.7.  כך שמהימנותם של אינדקס זיכרון עבודה ואינדקס מהירות עיבוד לא מספיק טובה, וגם מהימנותם של אינדקס הבנה מילולית ואינדקס היסק תפיסתי אינה גבוהה.  המחברים כותבים ש – 25% מהתלמידים קיבלו מנת משכל כללית שונה בעשר נקודות או יותר בין שתי המדידות, 29% קיבלו אינדקס הבנה מילולית, 39% קיבלו אינדקס חשיבה תפיסתית, 37% קיבלו אינדקס זיכרון עבודה ו – 44% קיבלו אינדקס מהירות עיבוד שונים בעשר נקודות או יותר בין שתי המדידות. 

 כזכור, כדי לקבוע לקות למידה על פי CHC  אנחנו צריכים להשתמש בציוני היכולות הקוגניטיביות השונות/האינדקסים השונים במבחני המשכל.  אם הם אינם מהימנים – קביעת לקות למידה לא תהיה מהימנה.  לפי ממצאי המחקר הזה, עלינו לותר על השימוש בציוני האינדקסים ולהשתמש בציון המשכל הכללי כ"רע במיעוטו" (משום שגם המהימנות של ציון המשכל הכללי לא היתה מאד גבוהה במחקר הזה).  בהקשר ללקות למידה, זה אכן מחזיר אותנו להגדרת לקות למידה על פי מודל הפערים. 

אבל, נרצה לקוות שבמשך ה - 2.84 שנים שילדים אלה קיבלו שירותי חינוך מיוחד כישוריהם הקוגניטיבים השתפרו.  אם זה כך, זה עשוי להסביר את חוסר היציבות במדדים.  כדאי לבדוק האם חוסר היציבות הזה קיים גם במחקר עם אוכלוסיה בחינוך הרגיל, וכן האם זה כך גם במבחני משכל אחרים?

הנחה שלישית:  שיטות PSW, שהגדרת לקות למידה לפי    CHCהיא אחת מהן, הן יישומיות ושימושיות לטיפול

המחברים טוענים שהיעילות של התערבויות שמתמקדות ביכולות קוגניטיביות נמוכות אינה מוכחת.  החוקר BURNS וחבריו (2016) גילו שההשפעה והיעילות של התערבויות שנגזרו מנתונים קוגניטיבים היו קטנות.  הוא הסיק שמדדים של יכולות קוגניטיביות אינם שימושיים בסקרינינג או בתכנון התערבויות בקריאה ובמתמטיקה. 

השאלה היא האם הגדרות לקות למידה על פי גישת הפערים או RTI מובילות להתערבויות יעילות יותר?

לסיכום, McGill & Busse תוקפים את הגדרת לקות למידה על פי CHC בשלוש דרכים:  א.  התוקף הדיאגנוסטי של המודל חלש.  ב.  שימוש באינדקסים אינו מהימן ותקף.  ג.  המודל לא מוביל לדרכי טיפול יישומיות.

אני לא בטוחה שהשתכנעתי.