Biology of Sport

Full text

2026 vol. 43
Original paper

Assessing the contributions of technical error and biological variability to error of measurement in reliability studies

  1. Health, Physical Activity and Sports Technology (Health-Tech), Physical Education and Sports, Faculty of Education, University of Alicante, 03690 Alicante, Spain
Biol Sport. 2026;43:1009–1018
Data publikacji online: 2026/03/16
Article file
76_05317_Article.pdf
Confronting perimenopausal women’s knowledge of coronary heart disease with their health behaviours. Controversial role of hormone replacement therapy in the protection of coronary heart disease

INTRODUCTION

Reliability refers to the reproducibility of outcomes from a measurement procedure when administered to the same individuals across multiple occasions. Better reliability represents more precise individual measurements and therefore better tracking of changes in research or practical situations. The most important measure of reliability is a standard deviation (SD) known as the typical or standard error of measurement, which is calculated from the changes within subjects between test and retest measurements [1]. A critical but often overlooked aspect of reliability is the fact that this SD consists of random biological variability, which each subject exhibits every time they are measured, combined with random technical error, which the measuring device adds to every measurement. Traditionally, these two sources of variability have been conflated into a single composite value representing the overall measurement error. Separating technical error from biological variability would allow assessment of device reliability independent of subject variability, which inevitably differs between types of subjects (young adults, athletes, the elderly, and so on).

The technical error can be estimated separately for some kinds of measurement, for example, the concentration of a biomarker in blood samples. Researchers can split the samples and analyze the splits by treating them like test and retest measurements. The resulting error of measurement is the technical error free of biological variability between samples and is often reported in the methods section as a coefficient of variation. This method is not appropriate for measurements of human performance or other behaviors, because it is not possible to split participants’ actions. Estimating the two components of measurement error is nevertheless possible by including simultaneous measurements with two devices during each test. Errors of measurement can be calculated across devices on each testing occasion. In this case, biological variability disappears from these errors of measurement, which now represent the combination of the technical errors of each device. Errors of measurement can also be calculated across testing occasions on each device, as in the usual reliability studies. Now the errors of measurement include the technical error and biological variability. From these estimates of standard error of measurement, it is possible to extract the estimates of the separate contributions of biological variability and technical errors.

One method of extracting the estimates is via a mixed model, as demonstrated by Pueo et al. [2]. Mixed models produce estimates of standard deviations for sources of variability along with their associated confidence intervals, representing sampling uncertainty. The use of mixed models can be problematic for sports practitioners lacking access to the specialized knowledge of a statistics consultant. Alternatively, simple algebraic equations can be used in a spreadsheet to estimate all the relevant standard deviations, but they cannot be used to estimate the uncertainties in the SDs.

The present study introduces the spreadsheet analysis. To deal with the uncertainties, we have used bootstrapping, by adapting a spreadsheet introducing the concept [3]. This approach is investigated through simulations in a computer program and compared with mixed modeling. Both approaches are then applied to real data from jump-height measurements. The aim is to provide a simple trustworthy method to estimate and evaluate device technical error and biological variability.

MATERIALS AND METHODS

Derivation of the spreadsheet method

The spreadsheet method is based on an algebraic partitioning of measurement error. The derivation requires a two-device, two-occasion study design, which yields variances from four key change scores. Figure 1 illustrates the composition of these variances from the constituent components of biological variability and technical error.

FIG. 1

Variances of measurement error with two devices (A, B) on two testing occasions (1, 2) and variances of the change score between tests and of the difference score between devices. The variances consist of biological variability in the tests (SDBiol12 and SDBiol22) and technical error of two devices (SDTechA2 and SDTechB2).

/f/fulltexts/BS/57748/JBS-43-57748-g001_min.jpg

The technical errors and biological variability can be derived from the observed values of the variances of the change scores between the tests for each device and of the variances of the change scores between the devices for each test, as follows. The observed variances of the change scores between tests for Device A and Device B are given by:

(1)
SDA2=SDBiol12+SDBiol22+2SDTechA2
(2)
SDB2=SDBiol12+SDBiol22+2SDTechB2

The observed variances of the change scores between devices for Test 1 and Test 2 are given by:

(3)
SD12=SDTechA2+SDTechB2
(4)
SD22=SDTechA2+SDTechB2

Adding (3) to (4) and rearranging:

(5)
SD12+SD222=SDTechA2+SDTechB2

Subtracting (2) from (1) and rearranging,

(6)
SDA2SDB22=SDTechA2-SDTechB2

Adding (5) to (6) and rearranging,

(7)
SDTechA2=SDA2SDB22+SD12+SD2222
(8)
SDTechB2=SDB2SDA22+SD12+SD2222

The variance representing biological variabilities on each Trial (SDBiol12,SDBiol22) cannot be estimated separately, but their average (SDBiol12,SDBiol22)/2) is given by combining and rearranging (1) with (7) or (2) with (8):

(9)
SDBiol12+SDBiol222=SDA2+SDB22SD12+SD2222

Assuming that the variance of the biological variability does not change between trials because no habituation or fatigue effects are present, the average variance is simple SDBiol12. The spreadsheet produces this estimate of the biological variability.

Subjects

To illustrate a practical application of this methodology, a study was conducted with thirty-one sports sciences students (22 males and 9 females). They were recreational athletes with varying abilities and training backgrounds. Subject characteristics (mean ± SD) for the males were age, 23.2 ± 1.7 years; height, 176 ± 7 cm; body mass, 75.0 ± 10.6 kg; those for the females were age, 21.8 ± 1.0 years; height, 165 ± 5 cm; body mass, 57.0 ± 5.3 kg. Exclusion criteria included lower-limb injuries and current medication use. Participants were instructed to abstain from alcohol and caffeinated beverages for 24 hours prior to testing. To control for circadian rhythm effects, each participant performed all jumps at the same time of day. This study was approved by the Human Research Ethics Committee of the University of Alicante (IRB No. UA-2019-02-25). All participants provided written informed consent before taking part in the study.

Instrumentation

A photoelectric bar system (OptoJump, Microgate, Bolzano, Italy) and a jump mat (Axon Jump, Bioengineering Sports, Buenos Aires, Argentina) were used simultaneously to measure each jump test, referred to as Devices A and B, respectively. Both devices detect take-off and landing events, and they use the ballistic kinematic equation to estimate jump height [4]. The equation used is h = t2g/8, where h represents jump height (m), t is the flight time (s), and g is the acceleration due to gravity (9.81 m/s2).

The OptoJump system is an optical measurement tool designed to calculate temporal parameters during movement. It consists of two bars, each 1 m in length, equipped with infrared light-emitting diodes (LEDs) that transmit and receive signals. When an athlete jumps, any interruption in the communication between the bars is detected, allowing the system to measure flight time with a precision of 1 ms. The OptoJump system has been previously validated for jump height measurement [5, 6] and has been widely used in jump-performance research [7, 8]. Device B, the Axon Jump mat, consists of a 1.0 × 0.8 m platform that functions as a mechanical pressure switch and is connected to a laptop. The mat detects pressure changes when the athlete takes off and lands, enabling the calculation of flight time with a resolution of 1 ms. The Axon Jump system has been validated for jump height assessment [9] and has been utilized in various studies focused on jump-height performance [10, 11].

Procedures

This observational study involved repeated measurements of maximum jump height during a single test session using two devices. Prior to jump execution, participants underwent a standardized 5-min warm-up on a cycle ergometer (Cardgirus Pro Medical, Alava, Spain) set at an 80-W power load and a cadence of 60 to 65 rpm. Following the warm-up, participants completed a series of familiarization jumps to practice proper technique and maintain balance during landing. Each participant then performed two countermovement jumps (CMJ), with a 1-min rest between repetitions. For correct execution, participants kept their hands on their hips, flexed their knees to a 90-degree angle, and executed the jump in a single continuous movement to achieve maximum height. Knee flexion was visually monitored in the sagittal plane, and any improperly performed jumps were repeated to ensure only successful trials were included in the analysis. The study was conducted at the Motion Analysis Laboratory (0001P1006) of the University of Alicante under controlled environmental conditions of approximately 22°C and 60% relative humidity.

Statistical analyses

The statistical analysis proceeded in three stages: a simulation study to evaluate and adjust the spreadsheet method, a simulation study to evaluate and adjust mixed models, and the application of the spreadsheet and mixed modeling to the practical jump-height data.

The spreadsheet method was evaluated with simulations using SAS statistical software implemented on-line in SAS Studio (SAS 9.4; SAS Institute, Cary, NC). The SAS programs are extensively annotated and provided as supplementary material. We generated 4000 independent datasets mimicking reliability studies with a given sample size and pre-defined true population values for means and SDs. The context for the simulations was measurement of V˙O2max, with population values of between- and within-athlete SDs (5.0% and 2.0% respectively) that could represent high-level athletes tested with devices that could have trivial and marginally moderate-large technical errors (0.3% and 3.0% respectively). For each simulated dataset, we applied the spreadsheet’s algebraic method. To quantify the uncertainty of the spreadsheet estimates, each dataset was bootstrapped with 1800 resamples (the maximum that could be included in the spreadsheet) to derive 90% confidence limits via the percentile method [3]. Bootstrapped variances were found to be biased in small samples, so the derived SDs were adjusted for this bias with an empirically derived factor. The SDs also showed the usual downward small-sample bias arising from taking the square root of variance, which was adjusted with a factor [12]. Width of bootstrapped confidence intervals for the means and SDs were found to be too narrow in small samples and were therefore adjusted with empirically derived factors to improve coverage. The spreadsheet incorporates these empirical adjustments, allowing for direct inspection of the formulas used to correct for bias and improve confidence interval coverage.

The performance of the spreadsheet was assessed by estimating bias (the difference between the true and simulated-mean values), uncertainty (the mean half-width of the 90% confidence interval), and coverage (the proportion of confidence intervals that included the true value, nominal being 90%). To assess the practical magnitude of the resulting errors, all outcomes were standardized. The spreadsheet and its simulation allow for standardization either with an external SD, or with either of two between-subject SDs provided by the data: the present SD (the observed SD in either trial, free of technical error) or the pure SD (the observed SD free of technical error and biological variability). For this study, standardization was performed using an external SD to permit meaningful evaluation of pure and present between-subject SDs. Sample values of the external SD were generated via a random variable with a chi-squared distribution. Bias arising from standardization was adjusted using the Becker factor [13] for means and an empirical modified Becker factor for SDs. Cohen’s modified and augmented thresholds were used to assess bias and uncertainty in means [14], and half these thresholds were used to assess SDs [15]. Thresholds for coverage were determined for 90% confidence intervals, which have a nominal expected error rate of 10% (10% of confidence intervals do not include the true value). Thresholds for count ratios (0.9, 0.7, 0.5, 0.3 and 0.1 for small, moderate, large, very large and extremely large reductions; the inverse of these for increases) [16] were applied to this nominal rate to determine the magnitudes for nominal coverage and for coverage substantially below and above nominal. Allowing for slight adjustment of some rounded values, nominal coverage is 89–91%; small, moderate, large, very large and extremely large over-coverage are 92–93%, 94–95%, 96–97%, 98–99%, and 100%; corresponding values for under-coverage are 88–86%, 85–80%, 79–67%, 66–1%, and 0%.

Simulated data were also used to evaluate two linear mixed models implemented with Proc Mixed in SAS. Simulations with poor or failed convergence of the mixed models were detected via zero standard error for the residuals and were excluded from further analysis. The proportion of such simulations was at most 3.5%, when the true technical errors were identical and sample size was the smallest (10). In the first mixed model, the technical error of the first device was estimated via the residual variance, while that of the second was estimated via the residual plus the extra variance of a dummy random-effect variable. This model was used when the true technical errors were different; it was also investigated with the residual estimating the larger technical error (and therefore estimating negative extra variance for the smaller technical error), but it performed less well (data not shown). In the second model, two residuals estimated the technical errors directly; this model was used when the true technical errors were identical, for which it performed better than the first model. In both models, a within-subject random effect estimated biological variability, and fixed effects for trial and device identities estimated mean changes between trials and mean differences between devices. The square roots of the variances and their confidence limits were adjusted for bias with the Gurland and Tripathi factor [12]. Confidence limits for differences and means of the SDs were estimated via parametric bootstrapping of the variances (their square roots being combined, after adjustment for bias), since the sampling distributions of such combinations of SDs are unknown. Parametric bootstrapping was also used for the confidence limits of the SDs themselves, since these were generally more accurate than those provided by the assumption of T distributions for the variances. The performance of the mixed models was assessed using standardization, as for the spreadsheet. The standard error for the standardized mean effects was augmented with the standard error arising from uncertainty in the standardizing SD, using a formula derived via first-order calculus to combine the standard errors. A similar formula was derived for the standardized variances. Bias arising from standardization was adjusted using the Becker factor for means and an empirical modified Becker factor for variances [13].

The jump-height data were analyzed with the spreadsheet and with the one-residual mixed model. Data were log-transformed before analysis, and standardization was performed with a value for an exterior SD [17]. Qualitative magnitudes and probabilistic inferences [14, 16, 18] replaced the analysis of bias, uncertainty and coverage.

RESULTS

Simulation performance

Table 1 shows the performance of the spreadsheet and mixed-model methods when simulating studies of sample size 30 with devices with unequal technical errors (true raw SDs of 0.30 and 3.0). Bias was trivial for all estimated parameters with both methods; the greatest bias occurred with the smaller of the technical errors, where the spreadsheet underestimated and the mixed model overestimated the true value. Uncertainties ranged from trivial to moderate and were similar for the spreadsheet and mixed model, although the spreadsheet uncertainty for the mean change between trials was a little wider. The spreadsheet method produced confidence interval coverage ranging from nominal (90–91%) to moderate over (94%), while coverage with the mixed model ranged from small under (86%) to very large over (98%). When the mixed model was run with the residual estimating the larger technical error, precision and coverage were a little worse for several SDs (data not shown).

TABLE 1

Performance of the spreadsheet (SS) and mixed model (MM) to estimate various means and standard deviations (SD) in 4000 simulations of reliability studies of two devices (A, B) on two occasions (Trial 1, Trial 2) with a sample size of 30, with chosen true raw values (shown in bold), and with standardized values estimated with an external standardizing SD drawn with a sample of size 30 from a population raw SD of 5.0 units. Bias, uncertainty and coverage were assessed with standardized values. Bias (the difference between true and estimated values) was trivial for all estimates. The raw values interpreted as percent units would be realistic in some settings.

MeasureMethodRaw valuesStdized valuesUncertaintyaCoverageb (%)


TrueEst., ± CLTrueEst., ± CL
Mean Trial 2 – Trial 1SS2.02.0, ± 1.10.400.40, ± 0.24small90, nominal
MM2.0, ± 0.90.40, ± 0.19trivial89, nominal

Mean Device B – Device ASS3.03.0, ± 0.70.600.60, ± 0.19trivial90, nominal
MM3.0, ± 0.60.60, ± 0.18trivial91, nominal

Pure between-subject SDSS5.05.0, ± 1.31.001.00, ± 0.36moder.92, small over
MM5.0, ± 1.21.01, ± 0.34moder.89, nominal

Within-subject SDSS2.02.0, ± 1.00.400.40, ± 0.22small94, moder. over
MM1.9, ± 0.60.39, ± 0.16small86, small under

Present between-subject SDSS5.45.4, ± 1.21.081.07, ± 0.36moder.92, small over
MM5.3, ± 1.11.08, ± 0.32moder.86, small under

Technical error ASS0.30.1, ± 1.60.060.02, ± 0.34moder.90, nominal
MM0.6, ± 2.10.11, ± 0.44moder.94, moder. over

Technical error BSS3.03.0, ± 0.60.600.60, ± 0.20small92, small over
MM2.9, ± 0.70.60, ± 0.19small92, small over

Technical error B – ASS2.72.9, ± 2.10.540.57, ± 0.45moder.90, nominal
MM2.4, ± 1.00.48, ± 0.26small90, nominal

Typical error ASS2.02.0, ± 0.50.400.40, ± 0.14small92, small over
MM2.0, ± 0.60.41, ± 0.16small92, small over

Typical error BSS3.63.6, ± 0.80.720.72, ± 0.24small92, small over
MM3.5, ± 0.60.71, ± 0.18small83, moder. under

Typical error B – ASS1.61.6, ± 0.80.320.31, ± 0.18small90, nominal
MM1.5, ± 0.60.30, ± 0.18small95, moder. over

Typical error A – within-subject SDSS0.00.0, ± 0.80.000.00, ± 0.16small94, moder. over
MM0.1, ± 1.30.03, ± 0.29small88, small under

Typical error B – within-subject SDSS1.61.6, ± 0.60.320.32, ± 0.15small94, moder. over
MM1.6, ± 0.70.32, ± 0.19small98, v.large over

Note: Stdized, standardized; Est., estimate provided by the method; CL, 90% confidence limits; moder., moderate; v.large, very large.

a Uncertainty refers to the standardized half-width of the confidence interval; magnitudes are defined by thresholds for means (< 0.20, trivial; 0.20–0.60, small) and for SDs (< 0.10, trivial; 0.10–0.30, small; 0.30–0.60, moderate).

b Coverage is the percentage of simulations where the confidence interval contained the true value; magnitudes for coverage are defined as: 80–85%, moderate under; 86–88%, small under; 89–91%, nominal; 92–93%, small over; 94–95%, moderate over; 96–97%, large over; 98–99%, very large over.

When the simulations were performed with identical devices (true raw technical error SD of 3.0 for both), bias was again trivial with either method (Table 2). Uncertainties were similar and small, with the exception of moderate for the within-subject SD. Coverage was under nominal only for the mixed-model’s mean typical error; coverage otherwise ranged from small to moderate over for the spreadsheet and from nominal to large over for the mixed model.

TABLE 2

Performance of the spreadsheet (SS) and mixed model (MM) as in Table 1, but with equal technical errors (3.0 raw units), to illustrate uncertainty and coverage when a study is performed with identical noisy devices and when relevant estimates of sums and differences of SDs are averaged. Bias was trivial for all estimates.

MeasureMethodRaw valuesStdized valuesUncertaintyaCoverageb (%)


TrueEst., ± CLTrueEst., ± CL
Technical error ASS3.03.0, ± 0.90.600.60, ± 0.24small93, small over
MM3.0, ± 0.90.61, ± 0.25small93, small over

Technical error BSS3.03.0, ± 0.90.600.60, ± 0.24small93, small over
MM3.0, ± 0.90.61, ± 0.25small93, small over

Technical error (A+B)/2SS3.03.0, ± 0.60.600.60, ± 0.19small93, small over
MM3.0, ± 0.60.61, ± 0.19small90, nominal

Within-subject SDSS2.02.0, ± 1.60.400.40, ± 0.35moder.94, moder. over
MM2.0, ± 1.50.40, ± 0.34moder.96, large over

Typical error ASS3.63.6, ± 0.80.720.72, ± 0.24small92, small over
MM3.6, ± 0.90.73, ± 0.22small89, nominal

Typical error BSS3.63.6, ± 0.80.720.72, ± 0.24small92, small over
MM3.6, ± 0.90.73, ± 0.22small89, nominal

Typical error (A+B)/2SS3.63.6, ± 0.60.720.72, ± 0.22small92, small over
MM3.6, ± 0.70.73, ± 0.18small87, small under

Typical error A – within-subject SDSS1.61.6, ± 1.40.320.32, ± 0.30small93, small over
MM1.6, ± 1.20.33, ± 0.29small96, large over

Typical error B – within-subject SDSS1.61.6, ± 1.40.320.32, ± 0.30small94, moder. over
MM1.7, ± 1.20.33, ± 0.29small95, moder. over

Typical error (A+B)/2 – within-subject SDSS1.61.6, ± 1.30.320.32, ± 0.27small93, small over
MM1.7, ± 1.10.33, ± 0.25small96, large over

Note: Stdized, standardized; Est., estimate provided by the method; CL, 90% confidence limits; moder., moderate.

a Uncertainty refers to the standardized half-width of the confidence interval; magnitudes are defined by thresholds for SDs (< 0.10, trivial; 0.10–0.30, small; 0.30–0.60, moderate).

b Coverage is the percentage of simulations where the 90% confidence interval contained the true value. Magnitudes for coverage are defined as: 86–88%, small under; 89–91%, nominal; 92–93%, small over; 94–95%, moderate over; 96–97%, large over.

Bias with sample sizes of 10 and 50 was trivial with all measures and models. Uncertainty and coverage with these sample sizes are summarized in Table 3. The models obviously performed better with the larger sample size, reaching trivial or small uncertainty and nominal coverage for means across all measures and models. Uncertainties for the spreadsheet were similar to those for the mixed model. There was more uncertainty in the SDs than in the means; with the larger sample size, the uncertainty in the SDs was still considerable (small or moderate), and there was still substantial (small) uncertainty with the change in the mean when technical errors were both 3.0. Coverage was better for the means than that for the SDs. Coverage for the spreadsheet was usually better than that for the mixed model, since the spreadsheet never showed under-coverage, and its over-coverage was often less than that of the mixed model.

TABLE 3

Performance of the spreadsheet (SS) and mixed model (MM) as in Tables 1 and 2, to illustrate uncertainty and coverage when studies are performed with sample sizes of 10 and 50, and with unequal or equal technical errors. Bias was trivial for all estimates.

Sample SizeTechATechBMethodMeansSDs


UncertaintyCoverageUncertaintyaCoverageb
100.33.0SSsmallnom., small overmod. – v.largesmall over – v.large over
MMsmallsmall undermod. – largemod. under – v.large over
3.03.0SSsmallnom., small overmod. – v.largemod. – v.large over
MMsmallsmall under, nom.mod. – v.largesmall under – x.large over
500.33.0SStrivialnom.small – mod.nom. – small over
MMtriialnom.small – mod.mod. under – v.large over
3.03.0SSsmall, trivialnom.smallnom. – mod. over
MMsmall, trivialnom.smallsmall under – large over

Note: TechA, technical error of Device A; TechB, technical error of Device B; moder., moderate; v.large, very large; x.large, extremely large.

a Uncertainty refers to the standardized half-width of the 90% confidence interval. Magnitudes are defined by thresholds for means (< 0.20, trivial; 0.20–0.60, small) and for SDs (< 0.10, trivial; 0.10–0.30, small; 0.30–0.60, moderate; 0.60–1.2, large; 1.2–2.0, very large).

b Coverage is the percentage of simulations where the 90% confidence interval contained the true value. Magnitudes for coverage are defined as: 80–85%, moderate under; 86–88%, small under; 89–91%, nominal; 92–93%, small over; 94–95%, moderate over; 96–97%, large over; 98–99%, very large over; 100%, extremely large over.

Jump-height reliability analysis

The analyses of the practical jump-height experiment are presented in Table 4. There was a small likely substantial increase in mean jump height on the second trial, but the difference in means between the devices was clearly trivial. The between- and within-subject standard deviations were clearly substantial, although their standardized uncertainties ranged from small (± 0.14, ± 0.13) to moderate (± 0.41 – ± 0.44). Spreadsheet and mixed-model estimates for these means and SDs were practically identical.

TABLE 4

Analyses with the spreadsheet (SS) and a mixed model (MM) of data from the reliability study of jump height measured with two devices (Device A, OptoJump; Device B, Axon jump mat) on two occasions (Trial 1, Trial 2). Standardization was performed with an external standard deviation (SD = 12.0%, sample size = 241). Qualitative magnitudes and probabilistic inferences were assessed with the standardized values.

MeasureMethodPercent valuesStandardized valuesQualitative magnitudeaProbabilistic inferenceb


Est., ± CLEst., ± CL
Mean Trial 2 – Trial 1SS3.8, ± 2.60.32, ± 0.22small**
MM3.8, ± 2.50.38, ± 0.21small**

Mean Device B – Device ASS−0.6, ± 1.1−0.05, ± 0.10trivialooo
MM−0.6, ± 1.0−0.05, ± 0.09trivialoooo

Pure between-subject SDSS18.0, ± 4.81.44, ± 0.44v.large****
MM18.5, ± 4.51.50, ± 0.37v.large****

Within-subject SDSS5.4, ± 1.40.46, ± 0.14moder.****
MM5.6, ± 1.20.48, ± 0.11moder.****

Present between-subject SDSS18.9, ± 4.51.50, ± 0.44v.large****
MM19.5, ± 4.31.57, ± 0.35v.large****

Technical error ASS1.6, ± 2.80.13, ± 0.25small
MM0.0, ± 2.50.00, ± 0.22trivial

Technical error BSS4.5, ± 0.90.38, ± 0.10moder.****
MM4.7, ± 0.70.41, ± 0.07moder.****

Technical error B – ASS2.9, ± 3.50.25, ± 0.30small**
MM4.7, ± 1.20.40, ± 0.11moder.****

Typical error ASS5.6, ± 1.10.47, ± 0.13moder.****
MM5.6, ± 1.20.48, ± 0.11moder.****

Typical error BSS7.1, ± 1.50.60, ± 0.16large****
MM7.4, ± 1.10.63, ± 0.10large****

Typical error B – ASS1.4, ± 0.90.12, ± 0.08small*o
MM1.7, ± 0.60.15, ± 0.05small**

Typical error A – within-subject SDSS0.2, ± 0.60.01, ± 0.06trivialooo
MM0.0, ± 0.30.00, ± 0.03trivialooo

Typical error B – within-subject SDSS1.6, ± 0.50.14, ± 0.05small**
MM1.7, ± 0.60.15, ± 0.05small**

Note: Stdized, standardized; Est., estimate provided by the method; CL, 90% confidence limits; moder., moderate; v.large, very large.

a The qualitative magnitude is based on the standardized point estimate (Est.) and interpreted using the following thresholds. For means: < 0.20, trivial; 0.20–0.60, small. For SDs: < 0.10, trivial; 0.10–0.30, small; 0.30–0.60, moderate.

b ↑↓ Indicate substantial positive and negative effects, respectively; indicates trivial effects. Probabilities are shown for effects with adequate precision at the 90% level. Probabilities of substantial effects:

* , possibly;

** , likely;

*** , very likely;

**** , most likely. Probabilities of trivial effects:

o , possibly;

oo , likely;

ooo , very likely;

oooo , most likely.

Notable differences were observed in the estimation of the technical errors of the two devices. For the OptoJump, the mixed model’s estimate (0.0%) differed substantially from the spreadsheet’s exact sample value (1.6%), but both these estimates had inadequate precision, and their comparison (not performed) would also have been unclear. The estimates for the Axon were almost identical; they were moderate and clearly substantial, and the differences from the OptoJump were small with the spreadsheet (likely substantial) and moderate with the mixed model (most likely substantial).

The technical errors combined with within-subject variability to produce clearly substantial moderate (OptoJump) to large (Axon) typical errors, and the values for the spreadsheet and mixed model were practically identical. There was some evidence that the difference in the typical errors was small (possibly substantial for the spreadsheet, likely substantial for the mixed model). The contribution of the technical error to the typical error was clearly trivial for the OptoJump but small and likely substantial for the Axon, again with practically identical outcomes for the spreadsheet and mixed model.

DISCUSSION

This study introduced and investigated a spreadsheet-based method for partitioning measurement error into its biological and technical components when analyzing data from a two-device, two-occasion reliability study. The components were expressed as SDs derived algebraically from variances of change and difference scores, and uncertainty on the SDs was determined with bootstrapping. Simulation was used to investigate the method’s performance, defined by bias and uncertainty in estimates of means and standard deviations, and by coverage of their confidence intervals. Performance was improved by implementing empirical corrections and was then compared with that of a linear mixed model. Our key finding is that the spreadsheet represents a trustworthy, accessible analytical tool, often superior to mixed modeling in relation to coverage with the small sample sizes common in sport and exercise science. This study also advances the conceptual approach to evaluating device error, in that the assessment of a device’s technical error should also be considered in the context of the subject’s biological variability: a given magnitude of technical error that might be substantial for a measure with low biological variability could be rendered negligible for a measure with high biological variability.

The bias in all means and SDs was trivial with both the spreadsheet and mixed-models, regardless of sample size and device errors (Tables 13). The qualitative magnitude of bias was derived by applying established magnitude thresholds to the standardized values, a process that makes our findings transferable but also contextdependent (for the simulations, measurement of V˙O2max). The finding of trivial standardized bias indicates that our empirical corrections were successful in addressing the small-sample bias inherent in the spreadsheet estimates of SDs. However, bias may become substantial in a population with a smaller standardizing between-subject SD, particularly for the smaller of two technical errors estimated with the spreadsheet and mixed model.

The uncertainty of the estimates, represented by the confidencelimit half-width, was similar for both the spreadsheet and the mixed model across all conditions. As expected, uncertainty was smaller with the larger sample sizes, although it remained small to moderate for most SDs even with a sample size of 50 subjects, which is more than sport scientists would normally use in a study of reliability. The substantial uncertainty in the estimates is a critical consideration for researchers interpreting these values: most observed SDs in a real study with 50 subjects would have to be at least small to qualify for adequate precision, and they would have to be at least moderate to qualify for clearly substantial. As with bias, these qualitative labels are based on standardized values; the reported uncertainties should therefore be interpreted as specific to a population with a betweensubject SD similar to that of our simulated population.

Differences between the two methods emerged in the coverage of the confidence intervals. Being unaware of published thresholds for what constitutes poor coverage, we have applied thresholds previously suggested for count ratios to the nominal 10% error rate [16, 19]. The findings represent clear evidence for the spreadsheet’s superior performance in providing trustworthy estimates of uncertainty: the spreadsheet method consistently achieved either nominal (89–91%) or small to moderate over-coverage (92–95%) across all conditions and sample sizes; importantly, it never demonstrated under-coverage (Tables 13). In contrast, the mixed model frequently exhibited substantial under-coverage, particularly with smaller samples, falling as low as 86%. Over-coverage makes inferences more conservative, but under-coverage indicates that the confidence intervals are too narrow, leading to an over-confident and potentially misleading interpretation of the results. The superior performance of the spreadsheet is likely due in part to the empirical corrections we devised to improve coverage, and no such attempt was made with the mixed model. Confidence intervals in the spreadsheet were also derived with non-parametric bootstrapping, which does not rely on the assumption of normality for the sampling distributions of variances, an assumption that is violated in small samples. Calculation of confidence intervals with the mixed model requires this assumption.

The analysis of the practical jump-height data further highlighted a key advantage of the spreadsheet’s direct algebraic approach over the mixed model’s iterative algorithm. In the estimation of technical error for the OptoJump device, the mixed model produced a boundary estimate of 0.0% (Table 4), a value that seems implausible for this electronic device. The spreadsheet, in contrast, calculated the exact, non-zero sample value of 1.6%. Such boundary estimates in mixed models are a known issue, particularly with small samples or low variance components, where the estimation algorithm can fail to converge on a realistic value within the parameter space [20]. In these scenarios, the spreadsheet provides a more robust and practically meaningful sample estimate. The actual estimates for the technical error of the Optojump (small with the spreadsheet, trivial with the mixed model) were nevertheless unclear with both methods, which is consistent with the findings of the simulations with a sample size of 30. Nevertheless, the comparison of the two technical errors provided good evidence (spreadsheet) and strong evidence (mixed model) for greater technical error with the Axon jump mat. Furthermore, there was very good evidence that the Optojump contributed negligibly to the typical error, which defines the precision of measurements in practice, whereas there was good evidence that the jump mat contributed substantially to the typical error. Both devices work by estimating jump height from flight time (the time when the subject’s feet are not in contact with the jump mat), so the greater technical error in the jump mat must arise from mechanical activation in the timing circuit.

Our findings indicate that for a two-trial two-device design with small samples, the mixed model’s performance is suboptimal in relation to coverage, so an empirical correction similar to that for the spreadsheet could be investigated. The mixed model’s iterative algorithm would also benefit from the increased information provided by additional trials and/or simultaneous measurement by more than two identical devices, scenarios that are unwieldy or impractical for the spreadsheet. Future research could therefore investigate the performance of mixed models in this context. The mixed model could also be applied to samples bootstrapped directly from the original data, thereby avoiding the assumption of normality for estimation of confidence intervals.

An important caveat to the analyses presented here is the assumption that the technical error of measurement arising from the devices is due entirely to “reliability” error; that is, the error varies randomly from measurement to measurement within subjects, and there is no additional “validity” error, which would vary randomly between subjects. Validity error with measurement of jump height would arise, for example, from a difference in jumping style between subjects, such as consistent differences in the extent of knee flexion on landing. Jump height calculated from flight time by the Optojump and jump mat would be affected equally by such differences, so the validity SD representing such differences between subjects will have contributed to the between-subject SD in our analyses. Jump height estimated from video analysis of the excursion of a marker on the subject’s trunk would not be affected by knee flexion, and a comparison of this method against the Optojump or jump mat would require a statistical model that included SDs representing validity errors of each device. In a preliminary investigation of such a model, we have found that the SDs cannot be estimated algebraically with a spreadsheet. The mixed model can provide all the estimates, but adequate precision of the validity SDs requires larger sample sizes (several hundred, if the SDs are trivial or small), and of course, the validity SDs would be estimated only to the extent that the devices have functionally dissimilar methods of measurement.

CONCLUSIONS

This study has introduced a spreadsheet that provides accessible trustworthy analysis of reliability data taken simultaneously with two devices. By estimating the contributions of biological variability and technical error to the error of measurement, our method provides a nuanced framework that allows researchers to make more informed judgments about a device’s utility for a specific population and measurement context. To facilitate its widespread adoption and to maximize the practical impact of this work, the validated spreadsheet has been made publicly available to the research community. The practical example highlights the relevance of this approach in contexts demanding precise measurement.

Acknowledgment

The authors gratefully acknowledge Dr. Jose M. Jimenez-Olmedo for valuable assistance in conducting the experimental procedures.

Conflict of interests

The authors have no conflicts of interest to declare.

Author contribution

Research concept and study design (B. Pueo, W. Hopkins), literature review (B. Pueo, W. Hopkins), data collection (B. Pueo), data analysis and interpretation (B. Pueo, W. Hopkins), statistical analyses (B. Pueo, W. Hopkins), writing of the manuscript (B. Pueo, W. Hopkins), or reviewing/editing a draft of the manuscript (B. Pueo, W. Hopkins). All authors have read and approved the final version of the manuscript, and agree with the order of presentation of the authors.

Availability of data and materials

The Excel spreadsheet tool (rely2devices.xlsx), which includes the raw jump-height data from the practical application presented in this study, is freely and publicly available for download from the University of Alicante institutional repository at the following permanent handle: http://hdl.handle.net/10045/161081 and in https://sportsci.org/resource/stats/rely2devices.xlsx. Users of the spreadsheet are kindly requested to cite the present journal article as the source of the method.

REFERENCES

1 

Hopkins WG. Measures of reliability in sports medicine and science. Sports Med. 2000; 30(1):1–15.

2 

Pueo B, Lipinska P, Jiménez-Olmedo JM, Zmijewski P, Hopkins WG. Accuracy of jump-mat systems for measuring jump height. Int J Sports Physiol Perform. 2017; 12(7):959–63.

3 

Hopkins WG. Bootstrapping inferential statistics with a spreadsheet. Sportscience. 2012; 16:12–5.

4 

Bosco C, Luhtanen P, Komi PV. A simple method for measurement of mechanical power in jumping. Eur J Appl Physiol Occup Physiol. 1983; 50(2):273–82.

5 

Castagna C, Ganzetti M, Ditroilo M, Giovannelli M, Rocchetti A, Mazi V. Concurrent validity of vertical jump performance assessment systems. J Strength Cond Res. 2013; 27(3):761–8.

6 

Glatthorn JF, Gouge S, Nussbaumer S, Stauffacher S, Impellizzeri FM, Maffiuletti NA. Validity and reliability of Optojump photoelectric cells for estimating vertical jump height. J Strength Cond Res. 2011; 25(2):556–60.

7 

Peng HT, Zhan DW, Song CY, Chen ZR, Gu CY, Wang IL, et al. Acute effects of squats using elastic bands on postactivation potentiation. J Strength Cond Res. 2021; 35(12):3334–40.

8 

Prieske O, Chaabene H, Puta C, Behm DG, Büsch D, Granacher U. Effects of drop height on jump performance in male and female elite adolescent handball players. Int J Sports Physiol Perform. 2019; 14(5):674–80.

9 

Cleveland JD, Patterson J. Assessment of vertical leap using a Vertec and Axon jump mat system. Med Sci Sports Exerc. 2010; 42(5):370.

10 

Carvalho FLP, Carvalho MCGA, Simão R, Gomes TM, Costa PB, Neto LB, et al. Acute effects of a warm-up including active, passive, and dynamic stretching on vertical jump performance. J Strength Cond Res. 2012; 26(9):2447–52.

11 

Torres-Banduc M, Ramirez-Campillo R, Andrade DC, Calleja-González J, Nikolaidis PT, McMahon JJ, et al. Kinematic and neuromuscular measures of intensity during drop jumps in female volleyball players. Front Psychol. 2021; 12:724070.

12 

Gurland J, Tripathi RC. A simple approximation for unbiased estimation of the standard deviation. Am Stat. 1971; 25(4):30.

13 

Becker BJ. Synthesizing standardized mean-change measures. Br J Math Stat Psychol. 1988; 41(2):257–78.

14 

Hopkins WG, Marshall SW, Batterham AM, Hanin J. Progressive statistics for studies in sports medicine and exercise science. Med Sci Sports Exerc. 2009; 41(1):3–12.

15 

Smith TB, Hopkins WG. Variability and predictability of finals times of elite rowers. Med Sci Sports Exerc. 2011; 43(11):2155–60.

16 

Hopkins WG. Magnitude-based decisions as hypothesis tests. Sportscience. 2020; 24:1–16.

17 

Pueo B, Hopkins WG, Penichet-Tomas A, Jimenez-Olmedo JM. Accuracy of flight time and countermovement-jump height estimated from videos at different frame rates with MyJump. Biol Sport. 2023; 40(2).

18 

Hopkins WG. Replacing statistical significance and non-significance with better approaches to sampling uncertainty. Front Physiol. 2022; 13:962132.

19 

Hopkins WG. Linear models and effect magnitudes for research, clinical and practical applications. Sportscience. 2010; 14:49–58.

Copyright: Institute of Sport. This is an Open Access article distributed under the terms of the Creative Commons CC BY License (https://creativecommons.org/licenses/by/4.0/). This license enables reusers to distribute, remix, adapt, and build upon the material in any medium or format, so long as attribution is given to the creator. The license allows for commercial use.
Share
without publication fees