💡
Each example includes: research question → data → formula → step-by-step → result → APA-7 report.
📋 1 · Descriptive Statistics
Exam Scores: Central Tendency & Spread
Descriptive
Research Question:
A professor records exam scores for 10 students. Describe the distribution.
| Student | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 |
|---|---|---|---|---|---|---|---|---|---|---|
| Score | 72 | 85 | 90 | 68 | 78 | 92 | 75 | 88 | 65 | 82 |
-
1Mean (x̄):
-
2Sample Standard Deviation:
-
3Median:
-
4SE:
79.5
Mean
80.0
Median
8.37
SD
2.65
SE
65–92
Range
APA-7
Descriptive statistics for exam scores (N = 10): M = 79.50, SD = 8.37, Mdn = 80.00, range = [65, 92].
⚖️ 2 · t-Tests
2a · One-Sample t-test
t-test
Research Question:
The national average is 75. Does the class (n=10, M=79.5, s=8.37) differ significantly?
Formula
-
1H₀: μ = 75 vs H₁: μ ≠ 75 (two-tailed, α = .05)
-
2
-
3Critical value: t*(df=9, α=.05) = ±2.262 → |1.700| < 2.262
-
4Cohen's d:
1.700
t
9
df
.124
p
0.538
Cohen's d
⚪ Fail to Reject H₀ — p = .124 > .05
APA-7
A one-sample t-test indicated that the class mean (M = 79.50, SD = 8.37) did not significantly differ from the national standard of 75, t(9) = 1.70, p = .124, d = 0.54, 95% CI [−1.48, 10.48].
2b · Independent-Samples t-test (Welch)
t-test
Research Question:
Does Drug A reduce pain scores more than Drug B?
| Group | n | Mean | SD |
|---|---|---|---|
| Drug A | 12 | 4.2 | 1.3 |
| Drug B | 10 | 5.8 | 1.9 |
Welch t-statistic
-
1
-
2
-
3Welch-Satterthwaite df:
-
4Cohen's d:
−2.258
t(Welch)
15
df
.039
p
1.021
Cohen's d
🔴 Reject H₀ — p = .039 < .05
APA-7
An independent-samples Welch t-test revealed that Drug A (M = 4.20, SD = 1.30) produced significantly lower pain scores than Drug B (M = 5.80, SD = 1.90), t(15.00) = −2.26, p = .039, d = 1.02, 95% CI [−3.11, −0.09].
2c · Paired-Samples t-test
t-test
Research Question:
Did training improve performance for the same 8 employees?
| i | Before | After | d = After − Before | d − d̄ | (d − d̄)² |
|---|---|---|---|---|---|
| 1 | 60 | 72 | 12 | 2.75 | 7.56 |
| 2 | 55 | 68 | 13 | 3.75 | 14.06 |
| 3 | 70 | 75 | 5 | -4.25 | 18.06 |
| 4 | 65 | 74 | 9 | -0.25 | 0.06 |
| 5 | 58 | 71 | 13 | 3.75 | 14.06 |
| 6 | 72 | 80 | 8 | -1.25 | 1.56 |
| 7 | 63 | 70 | 7 | -2.25 | 5.06 |
| 8 | 68 | 82 | 14 | 4.75 | 22.56 |
| Σ | 81 | 0 | 83.00 |
Cohen's d (paired)
8.320
t
7
df
<.001
p
2.941
Cohen's d
10.125
Mean Diff.
🔴 Reject H₀ — p < .001
APA-7
A paired-samples t-test indicated a statistically significant improvement in performance following training (M_diff = 10.13, SD_diff = 3.44), t(7) = 8.32, p < .001, d = 2.94, 95% CI [7.25, 13.00].
📊 3 · ANOVA
3a · One-Way ANOVA
ANOVA
Research Question:
Do three teaching methods produce different exam scores?
| Traditional (A) | Blended (B) | Online (C) |
|---|---|---|
| 70 | 85 | 75 |
| 72 | 88 | 78 |
| 68 | 82 | 72 |
| 74 | 90 | 80 |
| 71 | 87 | 76 |
| M=71 | M=86.4 | M=76.2 |
ANOVA Decomposition
-
1Grand Mean:
-
2SSBetween:
-
3SSWithin:
-
4F-statistic:
-
5η²:
36.24
F(2,12)
<.001
p
0.858
η²
306.85
MSB
8.467
MSW
🔴 Reject H₀ — F(2,12) = 36.24, p < .001
APA-7
A one-way ANOVA revealed a significant effect of teaching method, F(2, 12) = 36.24, p < .001, η² = .86. Tukey: Blended > Online > Traditional, all ps < .001.
🔢 4 · Non-parametric Tests
4a · Mann-Whitney U Test
Non-parametric
Research Question:
Do satisfaction scores differ between two departments?
| Dept A | 6 | 8 | 2 | 4 | 9 |
|---|---|---|---|---|---|
| Dept B | 7 | 5 | 10 | 3 | 1 |
U Statistic
-
1Rank all 10 combined (1=lowest):
Value 1 2 3 4 5 6 7 8 9 10 Rank 1 2 3 4 5 6 7 8 9 10 Group B A B A B A B A A B -
2
-
3Normal approximation (for reference):
-
4Effect size r:
11
U
.754
p (exact)
0.099
r
⚪ Fail to Reject H₀ — p = .754
APA-7
A Mann-Whitney U test revealed no significant difference in satisfaction scores between Department A and B, U = 11, z = −0.31, p = .754, r = .10.
4b · Kruskal-Wallis H Test
Non-parametric
Research Question:
Do pain levels differ across three clinics?
| Clinic A | Clinic B | Clinic C |
|---|---|---|
| 3,5,4,2,6 | 7,9,8,10,6 | 5,4,6,3,5 |
Kruskal-Wallis H
-
1Rank all 15 values:
(ties averaged) -
2
-
3Compare to χ²(2) = 5.99 → H = 3.34 < 5.99
3.34
H
2
df
.188
p
⚪ Fail to Reject H₀ — p = .188
APA-7
A Kruskal-Wallis test indicated no significant difference in pain levels across clinics, H(2) = 3.34, p = .188, η² = .10.
🔗 5 · Correlation
5a · Pearson Correlation
Correlation
Research Question:
Is there a linear relationship between study hours and GPA?
| i | X (hrs) | Y (GPA) | X−X̄ | Y−Ȳ | (X−X̄)(Y−Ȳ) | (X−X̄)² | (Y−Ȳ)² |
|---|---|---|---|---|---|---|---|
| 1 | 10 | 2.8 | -5 | -0.55 | 2.75 | 25 | 0.30 |
| 2 | 15 | 3.2 | 0 | -0.15 | 0 | 0 | 0.02 |
| 3 | 20 | 3.8 | 5 | 0.45 | 2.25 | 25 | 0.20 |
| 4 | 12 | 3.0 | -3 | -0.35 | 1.05 | 9 | 0.12 |
| 5 | 18 | 3.6 | 3 | 0.25 | 0.75 | 9 | 0.06 |
| 6 | 8 | 2.5 | -7 | -0.85 | 5.95 | 49 | 0.72 |
| 7 | 22 | 4.0 | 7 | 0.65 | 4.55 | 49 | 0.42 |
| 8 | 15 | 3.3 | 0 | -0.05 | 0 | 0 | 0.003 |
| Σ | 120 | 27.2 | 0 | 0 | 17.30 | 166 | 1.843 |
Significance test (t)
95% CI (Fisher-z)
.989
r
.978
r²
16.36
t(6)
<.001
p
🔴 Strong positive correlation, r = .989, p < .001
APA-7
Study hours and GPA were strongly positively correlated, r(6) = .99, p < .001, 95% CI [.94, 1.00]. Study hours accounted for 97.8% of the variance in GPA.
📈 6 · Simple Linear Regression
Predict GPA from Study Hours
Regression
Research Question:
Predict GPA from study hours (same data as correlation example).
OLS Coefficients
R², F, SE
| Coefficient | B | SE | t | p | 95% CI |
|---|---|---|---|---|---|
| b₀ (Intercept) | 1.837 | 0.107 | 17.17 | <.001 | [1.575, 2.099] |
| b₁ (Hours) | 0.1042 | 0.00634 | 16.44 | <.001 | [0.0887, 0.1197] |
0.978
R²
267.0
F(1,6)
<.001
p
0.1042
β (slope)
✅ Every additional hour → GPA +0.104. R² = 97.8%.
APA-7
Simple linear regression: GPA = 1.84 + 0.10 × Hours, F(1, 6) = 267.0, p < .001, R² = .978.
🗂️ 7 · Chi-Square Test
Smoking vs Cancer
Categorical
Research Question:
Is smoking associated with cancer diagnosis (N=200)?
| Cancer: Yes | Cancer: No | Row Total | |
|---|---|---|---|
| Smoker | 50 (E=32) | 30 (E=48) | 80 |
| Non-Smoker | 30 (E=48) | 90 (E=72) | 120 |
| Col Total | 80 | 120 | 200 |
Expected: E = (R × C) / N
χ²
Cramér's V
28.125
χ²
1
df
<.001
p
0.375
Cramér's V
🔴 Reject H₀ — p < .001. Smoking and cancer are significantly associated.
APA-7
A chi-square test of independence indicated a significant association between smoking and cancer, χ²(1, N = 200) = 28.13, p < .001, V = .375.
🔒 8 · Cronbach's Alpha
5-Item Job Satisfaction Scale
Reliability
Research Question:
Is this 5-item Likert satisfaction scale internally consistent?
| Resp. | Q1 | Q2 | Q3 | Q4 | Q5 | Sum |
|---|---|---|---|---|---|---|
| 1 | 4 | 3 | 5 | 4 | 3 | 19 |
| 2 | 2 | 2 | 3 | 2 | 2 | 11 |
| 3 | 5 | 4 | 5 | 5 | 4 | 23 |
| 4 | 3 | 3 | 4 | 3 | 3 | 16 |
| 5 | 1 | 2 | 2 | 1 | 2 | 8 |
| 6 | 4 | 5 | 4 | 4 | 5 | 22 |
| s² | 2.07 | 1.27 | 1.27 | 2.27 | 1.07 | σ²_T=28.56 |
Cronbach's α
0.902
α
5
Items
Excellent
Rating
✅ α = .902 — Excellent internal consistency (α > .90)
APA-7
The internal consistency of the 5-item job satisfaction scale was excellent, α = .90, exceeding the recommended threshold of .70.
🔁 9 · Repeated-Measures ANOVA
Memory Scores: 3 Time Points
RM-ANOVA
Research Question:
Do memory scores change over 3 time points in 5 participants?
| Subj. | T1 | T2 | T3 | P̄ |
|---|---|---|---|---|
| 1 | 4 | 6 | 8 | 6.0 |
| 2 | 5 | 7 | 9 | 7.0 |
| 3 | 3 | 5 | 7 | 5.0 |
| 4 | 6 | 8 | 10 | 8.0 |
| 5 | 2 | 4 | 6 | 4.0 |
| T̄ⱼ | 4.0 | 6.0 | 8.0 | 6.0 (Grand) |
SS Decomposition
📌
In real data: use MindStat for Mauchly's test + GG/HF corrections.
APA-7 (template)
A one-way RM-ANOVA indicated a significant effect of time, F(2, 8) = XX, p < .001, η²_p = .XX. Pairwise Bonferroni comparisons: all periods differ significantly.
⚡ 10 · Power Analysis & Sample Size
Sample Size for Independent t-test
Power
Scenario:
Two-group RCT, d = 0.5 (medium), α = .05, power = .80.
Cohen's n per group
| d | 80% Power | 90% Power | 95% Power |
|---|---|---|---|
| 0.2 (small) | 197 | 265 | 327 |
| 0.5 (medium) | 63 | 85 | 105 |
| 0.8 (large) | 26 | 34 | 42 |
| 1.0 (very large) | 17 | 22 | 27 |
63
n per group
126
Total N
80%
Power
APA-7
An a priori power analysis indicated that 63 participants per group (N = 126) were required to detect d = 0.50 with 80% power at α = .05 (two-tailed).
📐 11 · Normality Testing (Shapiro-Wilk)
11a · Data Consistent with Normality
Normality
Research Question:
Are BP readings normally distributed before applying a t-test?
| i | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 |
|---|---|---|---|---|---|---|---|---|---|---|
| BP | 118 | 120 | 122 | 124 | 124 | 126 | 128 | 130 | 133 | 135 |
Shapiro-Wilk W
-
1Sort ascending, compute x̄ and SS
-
2Coefficients (n=10): a₁=0.5739, a₂=0.3291, a₃=0.2141, a₄=0.1224, a₅=0.0399
-
3
0.972
W
.895
p
10
n
⚪ Fail to Reject H₀ — Normality confirmed, W(10) = 0.97, p = .895
APA-7
Shapiro-Wilk testing confirmed normality, W(10) = 0.97, p = .895.
11b · Detecting Non-Normality
Non-Normal
Research Question:
Are ER waiting times normally distributed?
| i | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 |
|---|---|---|---|---|---|---|---|---|---|---|
| Min | 5 | 8 | 10 | 12 | 15 | 20 | 35 | 52 | 78 | 95 |
-
1
0.825
W
.030
p
<.05
Violation
🔴 Reject H₀ — Data NOT normal, W(10) = 0.83, p = .030 → use non-parametric test
💡
Decision rule: p > .05 → normality assumed; p ≤ .05 → non-parametric or transform data.
▶ Test Normality in MindStat
🔀 12 · Two-Way ANOVA (Factorial Design)
Teaching Method × Class Size Interaction
ANOVA
Research Question:
Does teaching method's effect depend on class size? 2×2 factorial, N=24.
| Small | Large | Row Mean | |
|---|---|---|---|
| Traditional | 68.0 | 70.0 | 69.0 |
| Active | 82.0 | 74.0 | 78.0 |
| Col Mean | 75.0 | 72.0 | 73.0 (GM) |
SS Formulas
-
1Main Effect A — Teaching Method
-
2Main Effect B — Class Size
-
3Interaction A×B
-
4Interpret: Active Learning gains +14 pts in small classes but only +4 pts in large classes — the interaction is meaningful.
| Source | SS | df | MS | F | p | η²p |
|---|---|---|---|---|---|---|
| Method (A) | 492 | 1 | 492 | 16.27 | .001 | .448 |
| Class Size (B) | 60 | 1 | 60 | 1.98 | .174 | .090 |
| A × B | 156 | 1 | 156 | 5.16 | .034 | .205 |
| Error | 605 | 20 | 30.25 | |||
| Total | 1313 | 23 |
✅ Significant A×B interaction, F(1,20) = 5.16, p = .034
APA-7
A 2×2 ANOVA showed a significant Method × Class Size interaction, F(1, 20) = 5.16, p = .034, η²p = .21. Active Learning gained +14 points in small classes but only +4 in large classes.
🏅 13 · Spearman Rank Correlation
Study Hours Rank vs Exam Rank
Spearman
Research Question:
Monotonic relationship between study hours rank and exam rank (n=8)?
| Student | Hours | Rank X | Exam | Rank Y | d = Rₓ−Rᵧ | d² |
|---|---|---|---|---|---|---|
| A | 3 | 1 | 62 | 2 | −1 | 1 |
| B | 4 | 2 | 55 | 1 | 1 | 1 |
| C | 5 | 3 | 68 | 3 | 0 | 0 |
| D | 6 | 4 | 74 | 5 | −1 | 1 |
| E | 7 | 5 | 71 | 4 | 1 | 1 |
| F | 8 | 6 | 85 | 7 | −1 | 1 |
| G | 9 | 7 | 80 | 6 | 1 | 1 |
| H | 10 | 8 | 91 | 8 | 0 | 0 |
| Σ | 0 | 6 |
Spearman's rₛ
t-test for rₛ
.929
rₛ
6.15
t(6)
<.001
p
Large
Effect
🔴 Strong monotonic relationship, rₛ(6) = .929, p < .001
APA-7
A Spearman correlation indicated a strong positive relationship, rₛ(6) = .93, p < .001.
💡
Use Spearman when data is ordinal, non-normal, or contains outliers.
▶ Spearman Correlation in MindStat
📊 14 · Multiple Linear Regression
Exam Score ← Study Hours + Sleep Hours
Multiple Reg.
Research Question:
Can study hours and sleep hours together predict exam score (n=10)?
| i | X₁ (Study) | X₂ (Sleep) | Y (Score) | Ŷ | e = Y−Ŷ |
|---|---|---|---|---|---|
| 1 | 4 | 6 | 58 | 55.3 | 2.7 |
| 2 | 6 | 7 | 64 | 64.7 | −0.7 |
| 3 | 8 | 8 | 72 | 74.0 | −2.0 |
| 4 | 5 | 5 | 55 | 56.1 | −1.1 |
| 5 | 9 | 8 | 78 | 77.4 | 0.6 |
| 6 | 7 | 7 | 70 | 68.1 | 1.9 |
| 7 | 10 | 9 | 85 | 83.4 | 1.6 |
| 8 | 3 | 6 | 52 | 51.9 | 0.1 |
| 9 | 8 | 6 | 68 | 68.9 | −0.9 |
| 10 | 6 | 8 | 65 | 67.2 | −2.2 |
| x̄ | 6.6 | 7.0 | 66.7 |
OLS solution
| Predictor | b | SE | t | p | 95% CI | β (std) |
|---|---|---|---|---|---|---|
| Intercept | 26.34 | 4.18 | 6.30 | <.001 | [17.1, 35.6] | — |
| Study Hours | 3.39 | 0.41 | 8.27 | <.001 | [2.4, 4.4] | .74 |
| Sleep Hours | 2.57 | 0.74 | 3.47 | .010 | [0.9, 4.2] | .32 |
R², Adjusted R², F
.973
R²
.965
Adj R²
125.2
F(2,7)
<.001
p
✅ Model explains 97.3% of variance — study hours (β=.74) and sleep hours (β=.32) both significant.
APA-7
Multiple regression: F(2, 7) = 125.2, p < .001, R² = .973. Study hours (β=.74, p<.001) and sleep hours (β=.32, p=.010) both predicted exam scores.
⚙️ 15 · Logistic Regression
Predict Pass/Fail from Study Hours
Logistic
Research Question:
Does study hours predict pass/fail probability (n=20)?
Logistic Model
-
1Maximum likelihood estimates
-
2Predicted probabilities
Study hrs 6 7 8 9 10 P(pass) .115 .283 .550 .798 .922 -
3Odds Ratio (OR)
-
4Model fit
3.06
OR
<.001
Model p
.72
Nagelkerke R²
85%
Classified
✅ Study hours predicts pass/fail, OR = 3.06, p < .001
APA-7
Logistic regression: Study hours predicted pass/fail, χ²(1) = 14.5, p < .001, OR = 3.06, 95% CI [1.19, 7.87].
🎯 16 · Effect Sizes
Four Effect Size Measures with Worked Examples
Effect Size
Effect size measures the practical importance of a finding, independent of sample size.
① Cohen's d — for t-tests
Cohen's d
Worked example
② Eta-squared η² — for ANOVA
③ Cramér's V — for χ²
④ r (non-parametric effect size)
| Measure | Use with | Small | Medium | Large |
|---|---|---|---|---|
| Cohen's d | t-tests | 0.20 | 0.50 | 0.80 |
| η² (eta-squared) | ANOVA | .01 | .06 | .14 |
| η²p (partial) | Factorial ANOVA | .01 | .06 | .14 |
| Cramér's V | Chi-square | .10 | .30 | .50 |
| r (Pearson/Spearman) | Correlation | .10 | .30 | .50 |
| r (z/√N) | Mann-Whitney | .10 | .30 | .50 |
💡
Cohen's thresholds are benchmarks, not rules — interpret effect sizes in context.
▶ Compute Effect Sizes in MindStat
📏 17 · Confidence Intervals
Three CI Types — Mean, Proportion, Difference
CI
① 95% CI for a Single Mean
Scenario: n=30 BP readings, M=128.4, SD=14.2 → 95% CI
128.4
Mean
±5.3
Margin
[123.1, 133.7]
95% CI
② 95% CI for a Proportion
Scenario: 142/200 patients satisfied (p̂=0.71) → 95% CI
71.0%
p̂
±6.3%
Margin
[64.7%, 77.3%]
95% CI
③ 95% CI for Difference Between Two Means
Drug (M=85.2, SD=7.4) vs Placebo (M=79.6, SD=8.1), n=20 each → 95% CI for M₁−M₂
6.0
Δ Mean
±4.96
Margin
[1.04, 10.96]
95% CI
✅ CI excludes 0 → significant difference. Drug improves score by 1–11 points.
APA-7
95% CI for the mean difference = [1.04, 10.96]; excludes zero → significant at α = .05.
🔄 18 · Mediation Analysis
Mindfulness → Stress → Anxiety (Simple Mediation)
Mediation
Research Question:
Does mindfulness reduce anxiety indirectly through stress? N=90.
| Path | Description | b | SE | t | p |
|---|---|---|---|---|---|
| a | Training → Stress | −6.40 | 1.30 | −4.92 | <.001 |
| b | Stress → Anxiety | 0.62 | 0.11 | 5.64 | <.001 |
| c | Total effect (c) | −7.80 | 1.90 | −4.11 | <.001 |
| c' | Direct effect (c') | −3.83 | 1.82 | −2.11 | .037 |
Indirect Effect = a × b
−3.97
Indirect
[−6.12, −1.94]
Boot CI
50.9%
Mediated
✅ Significant partial mediation: 51% of training's effect on anxiety is through stress reduction.
APA-7
Simple mediation: ab = −3.97, 95% bootstrap CI [−6.12, −1.94]; 51% of total effect mediated through stress. Partial mediation (direct effect c' = −3.83, p = .037 remains).
💡
Full mediation: c' non-significant. Partial: c' remains significant. Use bootstrap CIs, not Sobel test.
▶ Mediation Analysis in MindStat
📉 19 · Survival Analysis (Kaplan-Meier)
Time-to-Relapse: Drug A vs Drug B
Survival
Research Question:
Does Drug A prolong relapse-free survival vs Drug B? (†=censored)
| Drug A (n=12) | 3 | 5 | 8 | 11 | 15 | 20† | 22† | 24† | 24† | 26† | 28† | 30† |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Drug B (n=12) | 2 | 4 | 6 | 8 | 10 | 14 | 18 | 20† | 22† | 24† | 26† | 28† |
-
1KM formula: S(tⱼ) = S(tⱼ₋₁) × (1 − dⱼ/nⱼ)
t (months) nⱼ (at risk) dⱼ (events) 1−dⱼ/nⱼ S(t) Drug A 0 12 0 — 1.000 3 12 1 11/12 0.917 5 11 1 10/11 0.833 8 10 1 9/10 0.750 11 9 1 8/9 0.667 15 7 1 6/7 0.571 -
2Summary statistics
Median Survival S(12 mo) Events/N Drug A 22 months 0.750 5/12 Drug B 14 months 0.500 7/12 -
3Log-rank test
22 mo
Median A
14 mo
Median B
5.02
χ²(1)
.025
p
✅ Drug A significantly prolongs survival: 22 vs 14 months, p = .025
APA-7
Kaplan-Meier analysis: Drug A (median = 22 mo) vs Drug B (median = 14 mo), log-rank χ²(1) = 5.02, p = .025. At 12 months, 75% vs 50% remained relapse-free.
💡
Censored observations (†) contribute data up to their last follow-up. KM correctly handles censoring.
▶ Survival Analysis in MindStat
📐 20 · Sample Size
How Many Participants Do I Need?
Sample Size
Research Question:
Master's student comparing two teaching methods.
| Parameter | Value | Why |
|---|---|---|
| d | 0.5 | متوسط (Cohen 1988) |
| α | 0.05 | معياري |
| Power | 0.80 | الحد الأدنى الموصى به |
Formula
-
1z-values:
-
2n per group:
-
3Attrition buffer:
-
4Total N:
63
Min/group
75
Recruit/group
150
Total N
80%
Power
APA-7
APA-7 power statement.
📋 Open your SPSS file
💻 21 · SPSS Output
Interpreting SPSS t-test Output
SPSS
Research Question:
How to read the SPSS Independent Samples Test table.
| Levene F | Sig. | t | df | p | Δ̄ | SE | CI₋ | CI₊ | |
|---|---|---|---|---|---|---|---|---|---|
| Equal var. | 2.14 | .149 | −3.42 | 48 | .001 | −4.80 | 1.40 | −7.63 | −1.97 |
| Unequal var. | −3.42 | 45.7 | .001 | −4.80 | 1.40 | −7.64 | −1.96 |
-
1Levene's Test: F = 2.14, p = .149 → use Equal variances row.
-
2t, df, p: t(48) = −3.42, p = .001 → significant.
-
3Cohen's d:
-
495% CI [−7.63, −1.97]: CI excludes zero → significant.
−3.42
t(48)
.001
p
−4.80
Δ̄
0.68
Cohen's d
🔴 Reject H₀ — p = .001 < .05
APA-7
APA-7 result.
✍️ 22 · APA-7 Writing
APA-7 Result Sentences
APA-7
Why:
APA-7 rules for reporting statistics.
| الخطأ | الصحيح |
|---|---|
| p = 0.043 | p = .043 |
| p = 0.000 | p < .001 |
| t = 3.42 (df=48) | t(48) = 3.42 |
| No effect size | Add d / η² / r / R² |
| R² = 0.47 | R² = .47 |
-
tt-test:t-test APA sentence.
-
FANOVA:ANOVA APA sentence.
-
rCorrelation:Correlation APA sentence.
-
R²Regression:Regression APA sentence.
p = .043
✓
p < .001
✓
t(48)
✓
η² / d
✓
APA-7
APA-7 methods sentence.
📋 23 · Likert Scale
Likert Scale — Complete Analysis
Likert
Research Question:
5-item Likert scale for 8 employees.
| م | ف1 | ف2 | ف3 | ف4 | ف5 | ∑ |
|---|---|---|---|---|---|---|
| 1 | 4 | 3 | 4 | 5 | 4 | 20 |
| 2 | 2 | 2 | 3 | 2 | 3 | 12 |
| 3 | 5 | 4 | 5 | 4 | 5 | 23 |
| 4 | 3 | 3 | 3 | 4 | 3 | 16 |
| 5 | 4 | 5 | 4 | 5 | 4 | 22 |
| 6 | 2 | 3 | 2 | 3 | 2 | 12 |
| 7 | 5 | 4 | 5 | 5 | 4 | 23 |
| 8 | 3 | 4 | 3 | 3 | 4 | 17 |
-
1Descriptives:
-
2Item mean:
-
3Cronbach's α:
-
4Interpretation: 3.63 falls in High range (3.41–5.00).
18.13
M total
3.63
M item
0.90
α
HIGH
Level
APA-7
APA-7 Likert result.
📊 24 · Multiple Regression
Multiple Regression — Step by Step
Multiple Reg.
Research Question:
Predict job performance from 3 predictors.
| م | Y | X₁ | X₂ | X₃ |
|---|---|---|---|---|
| 1 | 72 | 3 | 40 | 3.5 |
| 2 | 85 | 7 | 60 | 4.2 |
| 3 | 90 | 10 | 80 | 4.8 |
| 4 | 68 | 2 | 30 | 3.0 |
| 5 | 78 | 5 | 55 | 3.8 |
| 6 | 92 | 12 | 90 | 5.0 |
| 7 | 75 | 4 | 45 | 3.6 |
| 8 | 88 | 8 | 70 | 4.5 |
| 9 | 65 | 1 | 25 | 2.8 |
| 10 | 82 | 6 | 65 | 4.0 |
Model
-
1Coefficients:
المتنبئ b SE β t p b₀ 48.2 4.1 — 11.76 <.001 X₁ 1.83 0.42 .52 4.36 .003 X₂ 0.21 0.09 .28 2.33 .052 X₃ 3.94 1.21 .38 3.26 .014 -
2Model fit:
-
3Strongest predictor: X₁ β = .52 → strongest predictor.
-
4Interpret b₁: +1 year experience → +1.83 performance points.
.973
R²
71.9
F(3,6)
<.001
p
β=.52
X₁ best
🟢 Model significant — p < .001
APA-7
APA-7 multiple regression result.
🌀 26 · Mauchly's Sphericity Test
Mauchly's Test of Sphericity
RM-ANOVA
Research Question:
Do we meet the sphericity assumption before an RM-ANOVA?
| P | A | B | C | D |
|---|---|---|---|---|
| 1 | 8 | 7 | 1 | 6 |
| 2 | 9 | 5 | 2 | 5 |
| 3 | 6 | 2 | 3 | 8 |
| 4 | 5 | 3 | 1 | 9 |
| 5 | 8 | 4 | 5 | 8 |
| 6 | 7 | 5 | 6 | 7 |
| 7 | 10 | 2 | 7 | 2 |
| 8 | 12 | 6 | 8 | 1 |
Mauchly's W
-
1Orthonormal contrasts: Transform to k−1 = 3 orthonormal contrast scores.
-
2Covariance matrix S:
-
3W:
-
4χ² approximation:
-
5Decision: p = .044 → violated → apply GG/HF correction.
0.1362
W
11.41
χ²
5
df
.044
p
🔴 Sphericity violated — W = 0.136, χ²(5) = 11.41, p = .044.
APA-7
Mauchly's test indicated the sphericity assumption was violated, χ²(5) = 11.41, p = .044; degrees of freedom were corrected using Greenhouse-Geisser estimates (ε = .62).
💻 27 · Reading SPSS ANOVA Output
Interpreting SPSS ANOVA Output
SPSS
Research Question:
How to read the SPSS ANOVA and homogeneity tables.
| Source | SS | df | MS | F | Sig. |
|---|---|---|---|---|---|
| Between | 613.73 | 2 | 306.87 | 36.24 | .000 |
| Within | 101.60 | 12 | 8.47 | ||
| Total | 715.33 | 14 |
-
1Levene's Test first: Levene Sig. > .05 → standard ANOVA valid.
-
2F and Sig.: F(2,12) = 36.24, and .000 means p < .001, not p = 0.
-
3Effect size η²:
-
4Post-hoc: Tukey 'Multiple Comparisons': pairs with Sig. < .05 differ.
36.24
F(2,12)
<.001
p (from .000)
0.858
η²
🔴 Significant — F(2,12) = 36.24, p < .001, η² = .86.
APA-7
A one-way ANOVA showed a significant effect, F(2, 12) = 36.24, p < .001, η² = .86. (SPSS's .000 is reported as p < .001.)
📊 28 · Welch's ANOVA
Welch's ANOVA — Unequal Variances
ANOVA
Research Question:
Does course format affect scores despite wildly unequal variances?
| A | B | C |
|---|---|---|
| 58 | 50 | 78 |
| 62 | 100 | 82 |
| 59 | 55 | 79 |
| 61 | 95 | 81 |
| 60 | 75 | 80 |
Standard F
Welch's F
-
1Means and variances: B's variance is 200× A's or C's.
-
2Levene's test first: Variances significantly unequal → prefer Welch.
-
3Standard ANOVA (for contrast): Misleading 'not significant' — B's variance inflates the pooled error.
-
4Welch's ANOVA:B is down-weighted to near-zero; F jumps from 3.14 to 182.99.
-
5Decision: Report Welch's F — the standard F would be a Type II error here.
182.99
F_Welch(2,7.12)
<.001
p (Welch)
3.14
F standard
.080
p standard
🔴 Significant by Welch's ANOVA — F(2, 7.12) = 182.99, p < .001.
APA-7
Because Levene's test indicated unequal variances, a Welch's ANOVA was conducted, F(2, 7.12) = 182.99, p < .001; the standard ANOVA alone would have misleadingly suggested no effect, F(2, 12) = 3.14, p = .080.
🧪 29 · Levene's Test
Levene's Test — Variance Homogeneity
Assumption
Research Question:
Do novice and expert typists have equal reaction-time variance?
| Novice | Expert |
|---|---|
| 812 | 430 |
| 940 | 465 |
| 705 | 410 |
| 1050 | 455 |
| 880 | 398 |
| 690 | 448 |
| 965 | 421 |
Levene's F
-
1Means and spread: Experts are faster AND far more consistent.
-
2Transform then ANOVA on z: Levene's test = an ANOVA run on |deviations|.
-
3F:
-
4Decision: Violated → use Welch's t-test (MindStat's unconditional default).
12.67
F(1,12)
.0039
p
134.8 / 24.6
SD ratio
🔴 Variances significantly unequal — F(1,12) = 12.67, p = .004.
APA-7
Levene's test indicated the homogeneity-of-variance assumption was violated, F(1, 12) = 12.67, p = .004; a Welch's t-test was used for the comparison.
🔢 30 · Wilcoxon Signed-Rank Test
Wilcoxon Signed-Rank Test
Non-parametric
Research Question:
Did coping scores change after the program, without assuming normality?
| # | Before | After | Diff |
|---|---|---|---|
| 1 | 60 | 66 | +6 |
| 2 | 58 | 62 | +4 |
| 3 | 65 | 74 | +9 |
| 4 | 72 | 70 | −2 |
| 5 | 55 | 62 | +7 |
| 6 | 63 | 66 | +3 |
| 7 | 59 | 64 | +5 |
| 8 | 68 | 76 | +8 |
| 9 | 61 | 62 | +1 |
W+ / W−
-
1Rank |d|:
-
2Sum ranks by sign: Only employee #4 regressed.
-
3Exact test: Exact distribution used (n<50, no ties).
-
4Decision: p = .012 → reject H₀ → significant change.
2
W
9
n
.012
p
🔴 Significant — W = 2, p = .012 (exact).
APA-7
A Wilcoxon signed-rank test showed a significant increase in coping scores, W = 2, p = .012 (exact, two-tailed).
🔢 31 · Friedman Test
Friedman Test — Non-parametric RM
Non-parametric
Research Question:
Do the three study techniques differ in recall, without assuming normality?
| P | Re-reading | Practice Testing | Spaced Practice |
|---|---|---|---|
| 1 | 12 | 15 | 17 |
| 2 | 14 | 16 | 18 |
| 3 | 10 | 14 | 16 |
| 4 | 16 | 18 | 19 |
| 5 | 13 | 15 | 17 |
| 6 | 17 | 14 | 16 |
| 7 | 15 | 17 | 14 |
| 8 | 12 | 15 | 17 |
Friedman χ²
-
1Rank within each row: Not every participant follows the overall pattern.
-
2Rank sums:
-
3χ²:
-
4Decision + post-hoc: Reject H₀; Spaced Practice > Re-reading survives post-hoc correction.
6.25
χ²(2)
.044
p
8
N
🔴 Significant — χ²(2) = 6.25, p = .044.
APA-7
A Friedman test showed a significant difference across techniques, χ²(2) = 6.25, p = .044, N = 8; spaced practice beat re-reading on post-hoc (p_adj < .001).
🗂️ 32 · McNemar's Test
McNemar's Test — Paired Proportions
Categorical
Research Question:
Did training change the pass rate for the same 20 candidates?
| After: Pass | After: Fail | |
|---|---|---|
| Before: Pass | 6 (a) | 2 (b) |
| Before: Fail | 9 (c) | 3 (d) |
McNemar χ²
-
1Discordant pairs: b=2 (regressed), c=9 (improved).
-
2Uncorrected:
-
3Continuity-corrected: The correction changes the conclusion here.
-
4Decision: Corrected result is not significant — the uncorrected test overstates it.
3.27
χ² (cc)
.070
p (cc)
4.5
OR
20
N
🟢 Not significant (corrected) — χ²(1) = 3.27, p = .070.
APA-7
A McNemar's test with continuity correction showed no significant change in pass rate, χ²(1) = 3.27, p = .070 (OR = 4.50); the uncorrected test would have overstated the evidence.
🔗 33 · Partial Correlation
Partial Correlation
Correlation
Research Question:
Does exercise still relate to blood pressure once age is controlled for?
| P | X | Y | Z |
|---|---|---|---|
| 1 | 1 | 148 | 58 |
| 2 | 2 | 145 | 62 |
| 3 | 2 | 150 | 45 |
| 4 | 3 | 138 | 60 |
| 5 | 4 | 142 | 50 |
| 6 | 4 | 130 | 40 |
| 7 | 5 | 128 | 48 |
| 8 | 6 | 135 | 35 |
| 9 | 6 | 120 | 42 |
| 10 | 7 | 118 | 30 |
| 11 | 8 | 115 | 33 |
| 12 | 9 | 110 | 28 |
Partial r
-
1Zero-order correlations: Classic confound pattern: all three pairs are correlated.
-
2Partial r:
-
3Significance test (df=n−3):
-
4Decision: Relationship survives controlling for age, though attenuated.
−.846
Partial r
−4.765
t(9)
.001
p
−.939
Zero-order r
🔴 Significant — partial r(9) = −.846, p = .001.
APA-7
A partial correlation controlling for age showed exercise remained significantly related to blood pressure, r(9) = −.846, p = .001 (zero-order r = −.939).
🔗 34 · Point-Biserial Correlation
Point-Biserial Correlation
Correlation
Research Question:
How strong is the association between group and outcome, expressed as a correlation?
| P | Group | Score |
|---|---|---|
| 1 | 0 | 72 |
| 2 | 0 | 68 |
| 3 | 0 | 75 |
| 4 | 0 | 70 |
| 5 | 0 | 74 |
| 6 | 0 | 69 |
| 7 | 1 | 82 |
| 8 | 1 | 88 |
| 9 | 1 | 79 |
| 10 | 1 | 85 |
| 11 | 1 | 90 |
| 12 | 1 | 84 |
Point-biserial r
-
1Group means:
-
2Compute r:
-
3Significance test:
-
4Same test as t-test: Identical t/df to the independent t-test on the same data.
.904
r_pb
6.704
t(10)
<.001
p
.818
r²_pb
🔴 Significant — r_pb(10) = .904, p < .001.
APA-7
A point-biserial correlation showed a strong association, r_pb(10) = .904, p < .001 (Treatment M = 84.67 vs. Control M = 71.33).
🔬 35 · Exploratory Factor Analysis
Exploratory Factor Analysis
Reliability
Research Question:
Do the 6 items form one factor or two?
| P1 | P2 | P3 | C1 | C2 | C3 | |
|---|---|---|---|---|---|---|
| P1 | 1.00 | .65 | .62 | .20 | .18 | .22 |
| P2 | .65 | 1.00 | .68 | .15 | .20 | .17 |
| P3 | .62 | .68 | 1.00 | .22 | .19 | .21 |
| C1 | .20 | .15 | .22 | 1.00 | .64 | .60 |
| C2 | .18 | .20 | .19 | .64 | 1.00 | .66 |
| C3 | .22 | .17 | .21 | .60 | .66 | 1.00 |
KMO & Bartlett
-
1Gatekeeper checks: Both gatekeepers pass — proceed to extraction.
-
2Eigenvalues: Exactly 2 factors survive Kaiser's rule.
-
3Varimax rotation:Clean 2-factor structure — pay and coworker satisfaction are separate.
Item F1 F2 h² P1 .124 .853 .742 P2 .082 .887 .793 P3 .130 .866 .766 C1 .851 .109 .735 C2 .878 .105 .781 C3 .858 .121 .750 -
4Decision: Score as 2 subscales, not 1 total.
.744
KMO
555.13
Bartlett χ²
2
Factors
75.1%
Var. explained
🔴 Two-factor structure confirmed (KMO=.744, Bartlett p<.001).
APA-7
An EFA (PCA extraction, varimax rotation) confirmed a clean 2-factor structure, KMO=.744, Bartlett χ²(15)=555.13, p<.001, explaining 75.1% of variance.
📊 36 · ANCOVA
ANCOVA — Adjusting for a Covariate
ANOVA
Research Question:
Does method affect post-test scores after adjusting for pre-test?
| A: Pre | A: Post | B: Pre | B: Post |
|---|---|---|---|
| 60 | 80 | 59 | 68 |
| 65 | 75 | 63 | 75 |
| 58 | 85 | 61 | 66 |
| 70 | 83 | 66 | 80 |
| 62 | 73 | 60 | 64 |
| 64 | 87 | 68 | 78 |
Adjusted mean
-
1Group means: Groups started nearly level; ANCOVA still gives the correct adjustment.
-
2Pooled slope:
-
3Adjusted means:
-
4F-test:
-
5Decision: Method A > Method B after adjustment.
80.38 / 71.95
Adj. means
6.60
F(1,9)
.030
p
.423
η²_p
🔴 Significant — F(1,9) = 6.60, p = .030.
APA-7
A one-way ANCOVA with pre-test as covariate showed a significant effect of method, F(1, 9) = 6.60, p = .030, η²p = .42 (Method A adjusted M = 80.38 vs. Method B M = 71.95).
💻 37 · SPSS Regression Output
Interpreting SPSS Regression Output
SPSS
Research Question:
How to read the SPSS regression Model Summary, ANOVA, and Coefficients tables.
| Model | B | SE | Beta | t | Sig. |
|---|---|---|---|---|---|
| (Constant) | 15.20 | 4.10 | 3.71 | .001 | |
| Experience | 1.85 | 0.32 | .612 | 5.78 | <.001 |
| Education | 2.10 | 0.68 | .331 | 3.09 | .005 |
Model Summary: R=.856, R²=.733, Adj. R²=.713 | ANOVA: F(2,27)=37.03, p<.001
-
1Check the ANOVA row: F(2,27)=37.03, p<.001 → model is useful overall.
-
2R²: Report Adjusted R² with 2+ predictors.
-
3Each predictor: B is the unstandardized effect, holding the other predictor constant.
-
4Compare via standardized β: Standardized β enables cross-predictor comparison.
.733
R²
37.03
F(2,27)
<.001
p
β=.612
Strongest predictor
🔴 Model significant — R² = .733, p < .001.
APA-7
A multiple regression significantly predicted salary, F(2, 27) = 37.03, p < .001, R² = .733; experience (β=.612) and education (β=.331) were both significant predictors.
💻 38 · SPSS Correlation Output
Interpreting SPSS Correlation Output
SPSS
Research Question:
How to read the symmetric SPSS bivariate correlation matrix.
| Study Hours | Exam Score | ||
|---|---|---|---|
| Study Hours | Pearson Corr. | 1 | .742** |
| Sig. (2-tailed) | .000 | ||
| N | 20 | 20 | |
| Exam Score | Pearson Corr. | .742** | 1 |
| Sig. (2-tailed) | .000 | ||
| N | 20 | 20 |
-
1The matrix is symmetric: Read the correlation once, not twice.
-
2Read the asterisks: ** = p < .01; Sig. = .000 means p < .001.
-
3Judge strength: r=.742 is a large effect by Cohen's guidelines.
-
4Correlation ≠ causation: Correlation shows co-variation, not causation.
.742
r
<.001
p
18
df
🔴 Significant, large effect — r(18) = .742, p < .001.
APA-7
A Pearson correlation was significant and large, r(18) = .742, p < .001 (SPSS's .000 is reported as p < .001).
💻 39 · SPSS Reliability Output
Interpreting SPSS Reliability Output
SPSS
Research Question:
How to read the Reliability Statistics and Item-Total Statistics tables together.
| Item | Corr. Item-Total | α if Deleted |
|---|---|---|
| Q1 | .68 | .77 |
| Q2 | .71 | .76 |
| Q3 | .24 | .85 |
| Q4 | .65 | .78 |
| Q5 | .70 | .76 |
Reliability Statistics: Cronbach's α = .812, N of Items = 5
-
1Headline alpha: α=.812 is 'good' by common guidelines.
-
2Scan for weak items: Q3's .24 is below the .30 rule-of-thumb threshold.
-
3Confirm with α-if-deleted: Only Q3's α-if-deleted exceeds the overall α.
-
4Decide and report: Keep 5 items (α=.812) or drop Q3 (α=.85) — disclose the choice.
.812
α (5 items)
.24
Q3 corr.
.85
α if Q3 deleted
🟢 Good overall (α=.812); Q3 is a candidate for removal.
APA-7
The scale showed good reliability, α = .812; Q3's weak item-total correlation (r=.24) means dropping it would raise α to .85.
🧭 40 · How to Choose the Right Statistical Test
Choosing the Right Statistical Test
Guide
How to use this page:
Find your question below, follow the rule to a test, or use the interactive Advisor.
1 · Comparing groups (continuous outcome):
2 · Relationships between variables:
3 · Categorical / binary data:
4 · Scale and measurement quality:
5 · Checking assumptions first:
6 · Planning a study and reporting effects:
7 · Reading SPSS output and writing up results:
🟢 Still unsure? Use the interactive Test Advisor below.
▶ Open in MindStat