Wednesday, 28 February 2018

Critical Appraisal of a RCT (February 2018)

Critical appraisal of a RCT (February 2018).

Exercise has been shown to be as effective as surgery for subacromial pain syndrome (SAPS) in a number of different studies [1,2,3].  While this is encouraging for proponents of exercise based physiotherapy, it also raises more questions than it answers.  What type of exercise? How much of it? For how long? Does it matter if it's painful or not? How, why, or even does either treatment actually help?  This month we examine a RCT designed to try to shed some light some of these questions by investigating whether one type of exercise (in this case non-painful eccentric training of the shoulder external rotators) produces better results than another (general exercise) in this patient group.

This is the third in a series of blogs (the first two can be found here and here).  Again, this blog will not provide a systematic or comprehensive critical appraisal of the chosen paper, or employ one of the many critical appraisal tools available, but will highlight what we consider to be three important elements to consider when trying to interpret the results of the trial and apply them to our patients.  This blog will consider various aspects of RCT design and implementation, an overview of which can be found here.

This month we will consider the potential sources of bias and the degree of confidence or doubt we have in results of the following RCT:

Shoulder external rotator eccentric training versus general shoulder exercise for subacromial pain syndrome: a randomized controlled trial. The International Journal of Sports Physical Therapy, 12(7), 1121-33.

This study was designed to investigate whether eccentric training of the shoulder external rotators (ETER) or general exercise (GE) produces better clinical outcomes in those with SAPS.   This study defined SAPS based on the presence of at least three of the following; a positive Neer, Hawkins-Kennedy, or empty can test, painful resisted external rotation, palpable tenderness of the supraspinatus or infraspinatus insertion, or a painful arc of abduction. The primary outcome measure was the Western Ontario Rotator Cuff Index (WORC), a patient reported outcome measure which considers physical symptoms, sports and recreation, work, lifestyle, and emotions.   Secondary outcome measures were a Numerical Pain Rating Scale (NPRS) for best, worst, and average pain, isometric strength, active range of movement, Y balance test, and Global Rating of Change (GROC).  48 participants were randomised into two groups (25 in the ETER group and 23 in the GE group) and underwent a six week exercise program including four visits to a physical therapist (the study was conducted in the USA).  The ETER group performed non-painful eccentric external rotation (3 sets of 15 with a 3 second eccentric phase), resisted scapula retraction (2 sets of 10), and a cross body stretch (3 reps with 30-45 second holds).  The GE group performed active flexion and abduction with no resistance (2 x 10 reps of each), and the same resisted scapula retraction and stretching excises as the ETER group. All outcomes were measured at baseline, 3 week, 6 weeks, and 6 months.  The study found that ETER produced statically significant superior results compared to GE at 3 weeks, 6 weeks, and 6 months according to WORC score, NPRS score, and isometric muscle strength. There were no statistically significant differences in active range of movement, Y balance, or GROC. The authors conclude that eccentric training may be efficacious to improve self-reported pain, function and strength in those with SAPS.

It is often easier to find faults when critically appraising a study, so let's start with the strengths of this RCT.  It asked a clinically relevant question and used a design appropriate to answer it, and the study protocol was published before the trial began which guards against bias.  Additionally, the interventions and outcome measures were well described.  This may sound simple but the quality of description of interventions and outcome measures in RCT’s, especially those including therapeutic exercise, is often poor limiting interpretation, application, and replication of results [4,5].

However there are a few aspects of the trial which we must consider before deciding how confident we can be in the reported superiority of ETER exercises over GE.  The three areas we feel are most important to consider are:

1.     Differential dropouts and statistical analysis.
2.     The choice of comparator.  
3.     Methods of randomisation and baseline differences between groups.

1. Differential dropout rates and statistical analysis.

Dropouts and the resultant incomplete data cause investigators a number of problems when it comes to analysing and interpreting RCT results. One of the key characteristics of RCT’s that reduces their susceptibility to bias is randomisation.  Random allocation of participants to different treatment groups aims to ensure that the groups are comparable at baseline in relation to both known and unknown factors that might influence the outcome of the trial.  This increases our confidence that any differences in treatment effectiveness are related to the intervention of interest rather than any baseline differences.  In order to maintain this control of bias, the random treatment assignment must be preserved through the whole trial including in the statistical analysis.  This form of statistical analysis is known as an intention-to-treat (ITT) analysis.  This analyses participants based on the group to which they were assigned at randomisation irrespective of what treatment they actually received or whether they completed the trial. This is considered the most appropriate method of analysis when comparing the effectiveness of treatments in RCT’s [6].

The authors in this study did not employ an ITT approach because they were concerned that the asymmetrical dropouts between groups may cause type I error (finding a significant difference when one does not exist). Instead they analysed only participants that completed the trial which is termed a completed cases analysis.  Whether asymmetrical dropouts cause error in the results of an RCT depends on why the data are missing (rather awkwardly termed ‘missingness’) and how it is handled in the analysis [7].  If those that dropout do so completely at random, then a completed cases analysis is reasonable because the two groups available for analysis are still based on chance alone.  However, if the data is not missing completely at random, a completed cases approach analyses a non-random subset of those that entered the trial and compromises the initial randomisation process. In this trial the differential dropouts between groups (39% in the GE group dropped out compared to 12% in the ETER group) increases suspicion that that the reasons for this may not have been completely random [8].  Using a completed cases analysis therefore increases doubt that the differences in treatment outcome can be attributed to the intervention of interest confidently.

Author Response:
Thank you, this is a great point and one that we deliberated over for quite some time.  There are pros and cons with ITT and we had concerns that week 3 between group comparisons could be inflated with the 2 subjects in the control group reporting a slight worsening in subjective outcomes and subsequently dropping from the trial early.  This will end up lending to your comparator group interventions discussion but generally speaking my biggest concern was that the active range of motion exercises used as a comparator may have slightly increased symptoms in the early phase of treatment for some subjects in the control group and carrying over week 1 data could falsely inflate between group differences in favor of the experimental group.  If the 6 month dropouts were the only issue using ITT to carry over week 6 data to 6 month data would be an easier decision but the early dropouts from the control group were a big factor in the decision.

2. The choice of comparator.

To accurately test how effective a treatment is it needs to be compared to something.  Studies where treatment outcomes are measured without a comparison group can show that a patient had a particular treatment and got better, but cannot show that they got better because of the particular treatment.  Controlled studies (both RCT’s and other non-randomised controlled studies) use a control group to demonstrate what would have happened if the participants had not had the treatment of interest, either by doing nothing (no treatment control), making patients think that they have had the treatment of interest but without administering the active components (placebo control), or by comparing to another treatment (active control). In this case an active control was chosen.  While this is reasonable because the alternative to using eccentric exercises would be to provide an alternative exercise-based treatment, the most appropriate comparator would be representative of current practice (so we know whether changing to this ‘new’ or ‘different’ treatment is better than what we already do). This study uses range of movement exercises (with one resisted exercise that was standardised across groups) to represent general exercise.  The authors themselves identify that this may not be representative to a typical exercise program used in clinical practice.  Unless this reflects our current practice, it makes it very difficult to know what these results mean.

The choice of control also introduces doubt as to whether we can be sure that it was the type of exercise that was the decisive factor in determining the results of this trial.  Both groups performed exercises that involved both concentric and eccentric phases.  This makes it less clear whether this was a true comparison between two distinct types of exercise. The control group also performed lower dose and lower resistance exercises than the ETER group.  Previous studies have suggested that exercise protocols that include resisted exercise may be more effective that those that do not [9], and that higher dose exercise may be more effective than lower dose exercise [10]. Even if we accept the reported differences in outcomes between the two groups, can we be confident that it was the type of exercise that caused them?

Author response:  This is another great point.  I am not confident that the comparator exercise in this trial was representative of what a PT would do in practice.  Simply having a patient actively move the shoulder through an elevation movement without load may not be a typical general exercise program.  Its possible that the various differences between exercise programs, ie load, specific isolated movement, arm position etc could be the reason for between group differences rather than the fact that the experimental group utilized an eccentric exercise.

3. Methods of randomisation and baseline differences between groups.

As described, the benefit of randomisation is that it theoretically balances both known and unknown factors that could potentially influence the outcome of the trial between the groups. This increases our confidence that any difference in outcome is due to the intervention of interest and not some other known or unknown difference between groups.  In this trial the researchers randomised patients by asking them to blindly place a pen on a table of random numbers. Manual randomisation methods such as this (or the use of a coin toss, drawing lots, of shuffling cards) introduce more doubt than more robust methods like using computer generated or remotely generated random numbers because either the participant or investigator could theoretically influence the process.  For example what happened if a participant landed equidistant between 2, 3, or even 4 numbers in the table? This raises an important point in the assessment of the risk of bias; we are not saying that the participants or investigators did unduly influence the randomisation process in this trial, just that the method used increases doubt because they theoretically could have done so.  We as readers will never know for sure if the results were unduly influenced, and that is why we asses for the risk of bias rather than actual bias itself.

If there was bias in the randomisation process this would mean that there were systematic differences between the two treatment groups.  However, the fact that there were systematic differences between the two groups does not necessarily mean that there was bias in the randomisation process. Randomisation can only maximise the probability that known and unknown factors are balanced between the two groups; it cannot not guarantee that this is the case. The bigger the sample size the more likely they are to be balanced (the reasons for this were discussed in a previous blog here).   In this study there were statistically significant differences in favour of the ETER group in strength (ABD/ER ratio) and Y balance. There were also non-significant (but not necessarily non-important) differences in all other baseline strength measurements, most range of movement measurements, best pain, and younger age.  We do not really know how, why, or even if exercise really does help patients with SAPS so we cannot know how, why, or if these baseline differences affected treatment outcomes.  If it is feasible that younger, stronger, patients with better range of movement  and balance are more likely to benefit more from exercise based treatment,  then we have to consider that it could have been differences between groups rather than the differences in treatment effectiveness that caused the differences in outcomes.

Author response:  I also agree with this, hindsight is 20/20, If we did a similar trial again the use of computer generated randomization would be much preferred.  The topic of baseline variables that could be affiliated with improved outcomes is very important.  I would love to have collected more baseline variables in a larger sample and run a regression off the responders to determine the patient characteristics that are consistent with a positive outcome.  In this case we are examining the between group mean but some participants have dramatic improvements over others.  It would be interesting to know which patients respond best to heavy load exercises and which do not respond as favorably.

Conclusion

This study reports that eccentric training of the external rotators of the shoulder produces significantly better results that general exercise in those with SAPS. However, as with any RCT, it is important to appraise the methods used before accepting the results.  Differential dropouts between groups and the way the data were analysed might increase risk of bias and hence decreases our confidence in the reported results, and baseline differences between groups and the choice of comparator increases doubt as to whether any differences in outcome can be specifically linked to the type of exercise performed. 

Author response:  One last point about this topic is that the progression of exercise mode, load and volume dosing is critically important.  The level of tissue irritability is also an important factor to help dictate exercise prescription and in clinical practice I wouldn’t arbitrarily prescribe eccentric exercises to any patient with chronic sub-acromial pain.  Progressions in arm position, type of movement (ie. Isometric vs light isotonic vs eccentric) and load/dose increases respective of patient tolerance and baseline strength will be important to integrate into future trials.  A pragmatic design that allows the clinician to manipulate these exercise prescription variables based on patient presentation will be important in future studies.  Thank you all for your interest and review of this topic.
Eric Chaconas

_______________________________________________________ 
Thanks for reading, hope it was useful; thoughts gratefully received.

Paul Regan, Chris Littlewood, Tomas Parraguez, Brian Cho, Sijmen Hacquebord


[1]
Haahr JP, Østergaard S, Dalsgaard J, Norup K, Frost P, Lausen S, Holm EA, Andersen JH, (2005). Exercises versus arthroscopic decompression in patients with subacromial impingement: a randomised, controlled study in 90 cases with a one year follow up.  Annals of Rheumatic Diseases, 64(5), 760-4.

[2]
Haahr JP, Anderson JH, (2006).  Exercises may be as efficient as subacromial decompression in patients with subacromial stage II impingement: 4–8-years’ follow-up in a prospective, randomized study.   Scandinavian Journal of  Rheumatology, 35(3), 224–228.

[3]
Ketola S, Lehtinen JT, Arnala I, (2017).  Arthroscopic decompression not recommended in the treatment of rotator cuff tendinopathy: a final review of a randomised controlled trial at a minimum follow-up of ten years.  The Bone and Joint Journal, 99-B(6), 799-805.

[4]
Hoffmann TC, Glasziou PP, Boutron I, Milne R, Perera R, Moher D, Altman DG, Barbour V, Macdonald H, Johnston M, Lamb SE, Dixon-Woods M, McCulloch P, Wyatt JC, Chan AW, Michie S, (2014).  Better reporting of interventions: template for intervention description and replication (TIDieR) checklist and guide.  British Medical Journal, 348:g1687.

[5]
Page P, Hoogenboom B, Voight M, (2017).  Improving the reporting of therapeutic exercise interventions in rehabilitation research.  International Journal of Sports Physical Therapy, 12(2):297-304.

[6]
Higgins JPT, Green S (editors). Cochrane Handbook for Systematic Reviews of Interventions Version 5.1.0 [updated March 2011]. The Cochrane Collaboration, 2011. Available from http://handbook.cochrane.org.

[7]
Bell ML, Kenward, MG, Horton, NJ, (2013).  Differential dropout and bias in randomised controlled trials: when it matters and when it may not.  British Medical Journal, 346:e8668.

[8]
Moher D, Hopewell S, Schulz KF, Montori V, Gøtzsche PC,  Devereaux, PJ, Elbourne D,  Egger M, Altman DG, (2010).  ConSoRT 2010 explanation and elaboration: updated guidelines for reporting parallel group randomised trials.  British Medical Journal,340:c869


[9]
Littlewood C, Malliaras P, Chance-Larsen K, (2015).  Therapeutic exercise for rotator cuff tendinopathy: a systematic review of contextual factors and prescription parameters.  International Journal of Rehabilitation Research, 38(2), 95-106.

[10]
Østerås H, Torstensen TA, Østerås B, (2010). High-dosage medical exercise therapy in patients with long-term subacromial shoulder pain: a randomized controlled trial. Physiotherapy Research International, 15(4), 232-42.




Monday, 15 January 2018

Critical Appraisal of a RCT (January 2018) - Spanish version

Evaluación crítica de un ensayo controlado aleatorio (ECA)
Spanish Version
Otro día, otro ECA relacionado con el hombro, o eso parece. El hombro parece ser un tema candente en este momento por los cientos de ECA publicados en los últimos años. A primera vista, esto debería ser algo bueno, pero en realidad puede ser bastante confuso, especialmente cuando no se reciben mensajes claros y consistentes.

Entonces, teniendo esto en cuenta, este es el segundo blog de una serie (aquí) que evaluará críticamente los ECA publicados relacionados con el hombro con el objetivo de comprender cómo estos podrían ser relevantes para la práctica.

Antes de comenzar, para aquellos de ustedes que no están muy familiarizados con los ECA, en un blog anterior se discutió su diseño básico y su justificación, aquí. Este podría ser un punto de partida útil si algunos de los términos utilizados parecen desconocidos o confusos.

El blog de este mes se refiere a: Turgut et al. (2017). Efectos del entrenamiento con  ejercicios de estabilización escapular en la cinemática escapular, la discapacidad y el dolor en el Pinzamiento subacromial: Un ECA. Archivos de Medicina Física y Rehabilitación, 98, 1915-23. (DOI: 10.1016/j.apmr.2017.05.023)

En pocas palabras, este ECA se diseñó para evaluar si el estiramiento y ejercicios de fortalecimiento junto con ejercicios de estabilización escapular adicionales eran mejores que solo estiramiento y ejercicios de fortalecimiento en pacientes clasificados con:

•       - Arco doloroso durante la flexión o abducción del hombro,
•       - Dolor con rotación externa y abducción resistida,
•       - Disquinesia escapular basada en la evaluación observacional combinada con una reducción del dolor en el hombro durante el movimiento durante el test de asistencia escapular.

Los investigadores plantearon la hipótesis de que el grupo que recibió ejercicios adicionales de estabilización escapular mejoraría la posición y el movimiento de su escápula y reportaría menos dolor y discapacidad que el grupo que no recibió los ejercicios adicionales de estabilización escapular. En resumen, los autores informan que no hay diferencias significativas en el dolor y discapacidad entre los dos grupos después de un programa de entrenamiento de 12 semanas. Sin embargo, informan que observaron cambios estadísticamente significativos en la cinemática escapular en el grupo que recibió ejercicios específicos de estabilización escapular en comparación con los que no lo hicieron. Por lo tanto, la cinemática escapular pareció mejorar, pero esto no se asoció con mayores reducciones en el dolor o mejoras en la función.

Dado que los estudios previos no informaron cambios en la cinemática escapular mientras que los pacientes informaron una reducción del dolor y una mejor función [1] y dado que el papel de la disquinesia escapular en el hombro es incierto [2], parece un hallazgo interesante.


Apreciación critica:
En lugar de llevar a cabo una evaluación crítica sistemática e integral, para fines de blog, nos centraremos en algunos aspectos clave que pueden ayudarnos a juzgar si podemos confiar en los hallazgos de un ECA o si debemos ser cautelosos o incluso rechazar los hallazgos. Por lo tanto, tenga esto en cuenta y siéntase libre de agregar al debate como mejor le parezca.

Con respecto a Turgut et al, hay cuatro áreas en las que nos centraremos:
1.     Diferencias entre las características de los grupos
2.     Diferencias en la dosis de ejercicio recibida por los dos grupos
3.     Medición de la disquinesia escapular
4.     Tamaño de la muestra e incertidumbre

1.     Diferencias entre las características de los grupos:
Una característica de un ECA bien realizado es que los dos o más grupos que son creados sean similares al inicio de la prueba en términos de los factores que conocemos, como por ejemplo, edad, altura, peso, etc., pero también los factores que no conocemos o son difíciles de caracterizar, por ejemplo, el perfil genético. Esto es importante porque si queremos concluir al final del ECA que una intervención, es decir, el ejercicio focalizados en la escapular es mejor que otra intervención como el ejercicio general de fortalecimiento, entonces debemos estar seguros de que las únicas diferencias verdaderas entre los grupos son la intervención que ellos recibieron. Si esto no es así, no podemos estar seguros de que las diferencias que observamos se deben a la intervención y no a algún otro factor. Si esto parece confuso, se dan ejemplos en el blog anterior (aquí) para explicar más esto.

Ahora, esto podría no parecer importante en este ECA porque Turgut y cols. No informaron diferencias estadísticamente significativas entre los dos grupos después del programa de entrenamiento de 12 semanas. Pero, en una inspección más cercana, es evidente que las características de base de los dos grupos son diferentes; el grupo que recibe los ejercicios adicionales de estabilización escapular (el grupo de intervención) es en promedio seis años más joven (33.4 comparado con 39.5 años) y en promedio tiene un Índice de Masa Corporal (IMC) dos puntos más bajo (23.7 comparado a 25.8 - equivalente a 7kg de diferencia para un hombre de 1,8 m).

¿Por qué esto podría ser importante? Se reconoce que la edad puede estar asociada con un peor pronóstico y que la evaluación de la disquinesia escapular es difícil, es decir, es inherentemente poco confiable. No está claro si la diferencia de edad en este ECA es relevante, pero es probable que el IMC promedio más alto del grupo de control (como se refleja en los criterios de selección para el ECA) dificulte la evaluación de la disquinesia escapular y por lo tanto sea más difícil determinar si ha sucedido un cambio.

No está claro por qué se produjo el desequilibrio en las características iniciales. Turgut et al utilizan un método válido para generar su secuencia aleatoria, pero no informan cómo oculto su asignación, lo que suscita preocupación. Otros factores, por ejemplo, el dolor y la discapacidad, parecen estar bien equilibrados. La razón otra vez podría estar relacionada con el tamaño de muestra pequeño. Para demostrar esto, toma dos monedas y un amigo. Los dos tiran la moneda diez veces. ¿Obtienen el mismo número de cara o cruz entre sí? ¿Tienes cinco caras y cinco cruces? Probablemente no, porque este es un proceso aleatorio. Ahora intente voltear las monedas 20 veces; es probable que aún obtenga una cantidad diferente de caras y cruces entre sí, pero es probable que obtenga un número más equilibrado de caras y cruces. Ahora intente dar vuelta la moneda 30 veces y el impacto de aumentar el tamaño de la muestra, es decir, el número de volteos, se volverá más claro a medida que el número de caras y cruces se acercan con más volteos. 

2.     Diferencias en la dosis de ejercicio recibida por los dos grupos
Se ha informado una relación de respuesta a la dosis de ejercicio para pacientes que se quejan de dolor en el hombro, pero que aún pueden mover su brazo [3,4]. Esto es importante en el contexto de este ECA que busca evaluar si un tipo específico de ejercicio, es decir, ejercicios adicionales de estabilización escápular, confiere mejores resultados clínicos. En este contexto, debemos considerar si algún resultado se debe al tipo específico de ejercicio o simplemente debido a que se hace más ejercicio.

Turgut et al. estandarizó los ejercicios de estiramiento entre los grupos pero el grupo de intervención realizó diferentes ejercicios de fortalecimiento para el grupo control (se agregó un paso a cada uno) y una cantidad diferente de ejercicio resistido debido a la adición de los ejercicios de estabilización escápular (un mínimo de 240 repeticiones y un máximo de 480 repeticiones 3 veces por semana en comparación con un mínimo de 90 repeticiones y un máximo de 180 repeticiones 3 veces por semana en el grupo de control). Por lo tanto, cualquier diferencia observada entre los dos grupos podría deberse a la dosis adicional de ejercicio en lugar de deberse al tipo específico de ejercicio.

3.     Medición de la disquinesia escapular:
Como ya se mencionó, la medición de la disquinesia escapular es difícil e inherentemente poco confiable. Turgut et al. Utilizó un sistema de seguimiento electromagnético e informó la evidencia que respalda la confiabilidad y la validez, pero con errores de medición estándar que varían entre 3.37⁰ y 7.44⁰ y un cambio mínimo detectable que varía de 7.81⁰ a 17.27⁰. Dado que 8⁰ fue considerado como una diferencia asimétrica importante por los autores, el desafío de medición es claro de ver. Pero, esta limitación es apropiadamente reconocida por Turgut et al. en su sección de limitaciones y no es inmediatamente evidente lo que podrían haber hecho de manera diferente con respecto a la herramienta de medición utilizada; aun así, esta es una limitación importante.   

Sin embargo, dados estos problemas relacionados con la medición, una característica del diseño que podría ser útil es cegar al evaluador de resultados. El cegamiento es cuando los participantes / pacientes, los médicos y aquellos que evalúan los resultados de la investigación desconocen qué tratamiento recibió el paciente y se lo conoce como cegamiento simple, doble y triple, respectivamente. El cegamiento del evaluador de resultados protege contra el sesgo de medición. Sesgo de medición es un riesgo donde la medición no es objetiva, por ejemplo, vivo o muerto, y donde el evaluador de resultados podría influir en la medición consciente o inconscientemente, tal vez porque tienen preferencia por una de las intervenciones. Por ejemplo, si los propios investigadores tuvieran la hipótesis de que la adición de ejercicios de estabilización escapular daría lugar a mejores resultados clínicos para realizar la medición, es posible sugerir que podría haber un riesgo de sesgo de medición.

El cegamiento de la evaluación de resultados no fue informado por Turgot et al, pero debería haber sido factible, aunque muchos estudios se hacen con limitaciones de recursos que podrían haber evitado el empleo de personal adicional. A pesar de esto, la falta de cegamiento del evaluador de resultados en este ECA es potencialmente otra limitación.

4.     Tamaño de la muestra e incertidumbre:
Si bien es posible generar hallazgos a partir de un ECA pequeño que se consideran con validez interna, es decir, podemos confiar en ellos, lo que luego se vuelve difícil es generalizar con algún grado de certeza esos hallazgos a la población en general. Recuerde, a menudo en la investigación intentamos inferir los hallazgos de nuestra muestra de investigación a la población en general, es decir, con respecto a Turgot et al, de los 30 participantes en el ECA a la población más amplia de pacientes con este tipo de dolor de hombro. Cuanto menor es la muestra, más inciertos estamos que los hallazgos sean generalizables para la población en general porque simplemente tenemos menos información. Esta incertidumbre a menudo ahora se presenta como un intervalo de confianza del 95%, es decir, el rango de valores dentro del cual tenemos un 95% de certeza de que radica el verdadero valor poblacional (reconociendo que no estaremos 100% seguros a menos que investiguemos la totalidad de la población, lo cual generalmente no es posible). Por ejemplo, un ECA podría concluir que la diferencia entre los dos grupos en el ensayo fue de dos puntos en una escala analógica visual del dolor a favor del grupo de intervención con un intervalo de confianza del 95% de -2 a +4. Estas estadísticas significan que en el ECA, la diferencia observada entre los grupos fue de dos puntos. Pero, si tuviéramos que repetir este estudio, la diferencia real podría ser dos puntos a favor del grupo de control (-2) o hasta cuatro puntos a favor del grupo de intervención (+4). En este ejemplo, vemos que el intervalo de confianza cruza el cero, es decir, el punto donde no hay diferencia entre los dos grupos y, en consecuencia, este resultado se consideraría como no estadísticamente significativo.

Es posible que esté familiarizado con la significación estadística con respecto al valor p con p> 0.05 considerado como no estadísticamente significativo. Esto significa, en base a los datos de muestra, que no podemos rechazar la hipótesis nula, la cual no establece ninguna diferencia entre los grupos. Es importante leer esta declaración con cuidado porque no es lo mismo que decir que los dos grupos son iguales.

Dado que el intervalo de confianza del 95% brinda un rango de valores que son más fáciles de interpretar, esto es ahora preferido según lo que solicitan guías de reporte. Desafortunadamente, Turgot et al solo nos presentan valores de p que sugieren que no hay una diferencia estadísticamente significativa entre los dos grupos en términos de dolor de hombro y discapacidad al inicio, después de seis y 12 semanas. Pero, observando más de cerca, observamos que la diferencia entre los dos grupos en términos del puntaje total SPADI (Índice de dolor y discapacidad del hombro) es de siete puntos por seis semanas y de 13 puntos por 12 semanas a favor del grupo que recibió los ejercicios adicionales de estabilización escapular (10 puntos se consideran cambios clínicamente significativos en el SPADI). Entonces, ¿por qué Turgot et al informan que no hay diferencia? Una razón podría ser que el número de participantes en el ensayo es muy pequeño (15 en cada grupo), la variable de datos y, por lo tanto, no hay pruebas suficientes para rechazar la hipótesis nula que no hay una verdadera diferencia entre los dos grupos, debido a la información limitada proporcionada por el pequeño número de participantes. Por lo tanto, la falta de una diferencia estadísticamente significativa no se debe a que los resultados del ECA sugieran que los dos grupos tienen el mismo efecto, sino que no hay evidencia suficiente a partir de los datos - esta es una diferencia vitalmente importante, posiblemente indicativa de un error de Tipo II (alguna lectura adicional posterior al blog para usted).  

Conclusión:
Entonces, debido a la preocupación por la diferencia en las características basales, la diferente dosis de ejercicio, el posible error de medición y el riesgo de sesgo, no podemos confiar en que los cambios informados en la cinemática escapular sean válidos y atribuibles a la adición de ejercicios de estabilización escapular. Además, debido al pequeño tamaño de la muestra, no podemos estar seguros de que el informe de que ninguna diferencia estadísticamente significativa entre los dos grupos infiera que la efectividad clínica de las intervenciones en este ECA son similares. Por lo tanto, en base a esta evaluación crítica, se recomienda que los hallazgos de este ECA se traten con precaución; actualmente no está claro cómo estos resultados desarrollan nuestra comprensión sobre cómo, o incluso si, la cinemática escapular cambia en respuesta a la intervención. Tampoco está claro cómo los resultados de este ECA nos ayudan a desarrollar nuestra comprensión de los componentes importantes de un programa de ejercicios.

Es de esperar que haya algunos puntos de aprendizaje en este blog, aunque una implicación clara para la futura investigación relacionada con la fisioterapia es que los ECA que buscan determinar la efectividad de diferentes intervenciones necesitan una muestra suficiente para generar recomendaciones de tratamiento con confianza y sean diseñados para aumentar la confianza de que cualquier diferencia observada se puede atribuir a una variable de interés.

Gracias por leer, espero que haya sido útil; pensamientos recibidos con gratitud

Translated by Tomas Parraguez on behalf of:

Chris Littlewood, Tomas Parraguez, Brian Cho, Sijmen Hacquebord, Paul Regan

[1]       Bury J, West M, Chamorro-Moriana G, Littlewood C. Effectiveness of scapula-focused approaches in patients with rotator cuff related shoulder pain: A systematic review and meta-analysis. Man Ther 2016;25:35–42. doi:10.1016/j.math.2016.05.337.
[2]       Littlewood C, Cools AMJ. Scapular dyskinesis and shoulder pain: the devil is in the detail. Br J Sports Med 2017;0:bjsports-2017-098233. doi:10.1136/bjsports-2017-098233.
[3]       Osteras H, Torstensen T, Haugerud L, Osteras B. Dose-response effects of graded therapeutic exercises in patients with long-standing subacromial pain. Adv Physiother 2009;11:199–209.
[4]       Littlewood C, Malliaras P, Chance-Larsen K. Therapeutic Exercise for rotator cuff tendinopathy: A systematic review of contextual factors and prescription parameters. Int J Rehabil Res 2015;38.