Cesaly runs a real research question through Quanta: do three teaching methods change exam scores? The full transcript is below the video, lightly edited for reading.
Hi, this is Cesaly with ReliCheck. Today I want to show you something that, the first time it clicks, kind of changes how you see your own data. We're going to take one real research question, the messy human kind you actually care about, and get a real answer to it using a one-way ANOVA in Quanta, our Mac app for folks doing social, behavioral, and health science research. I'll run the whole thing in front of you and explain every number as it comes up, so you don't just watch me do it. You leave knowing how to do it yourself.
Let me make this real, because ANOVA only matters once it's attached to something you'd actually want to know. Picture three instructors teaching the same course. One lectures the way most of us were taught. One blends the lecture with online work. The third throws out the script and runs an active classroom, students up out of their seats solving problems together. Same course, same final exam, 90 students in all. The grades come in and the active classroom is on top. The blended group did well too. And here's that moment every one of us knows: you're staring at those averages and you want to believe the teaching made the difference, but there's that quiet voice in the back of your head going, what if it's just luck? 90 students, 90 different lives, 90 different weeks. Maybe that gap shows up no matter how you teach. That question, is this real or is it just noise, is the whole reason ANOVA exists, and it's exactly what we're about to settle together in Quanta.
Before I touch a single button, let me give you the idea, because it's simpler than the name makes it sound. You've probably used a t test. A t test compares two groups. ANOVA just does that for three or more all at once. You might be thinking, why not run a bunch of t tests instead? Lecture against blended, blended against active, and so on. Here's the problem: every test you run rolls the dice on a false alarm, and the more you run, the more likely one of them lies to you. ANOVA sidesteps that. It asks one clean question first: is anything going on among these groups at all? One test, one honest answer.
The way it works is right there in the name, analysis of variance. It takes all the spread in your exam scores and sorts it into two buckets. One bucket is the spread between the groups, the gap between those class averages. That's the signal. The other is the spread inside each group, just students being students. That's the noise. The number ANOVA hands you, called F, is simply the signal divided by the noise. When F is big, the differences between your groups tower over the everyday scatter inside them, and that's your sign the teaching actually did something.
On the overview page, I'll click Import your data and bring in my file. Notice this is just a plain CSV, the kind that falls out of any survey tool or spreadsheet. Quanta also reads Excel, SPSS, Stata, even ReliCheck survey packages, so whatever you've already got, you can bring it straight in. And look what Quanta tells me the second it lands: 90 rows, three columns, zero missing cells, and that little green check, structure looks aligned. Before I've done a thing, it's already confirmed my data came in clean. That's the kind of quiet reassurance you want at the start, because a messy import is where so many analyses quietly go wrong.
Now look at the table itself. This is our actual study. 90 students, one per row, three columns. There's a student ID, and we'll leave that one alone, it's just a name tag. There's Method, telling us which of the three classrooms each student sat in. And there's Exam Score, the thing we actually care about. Here's a habit worth picking up: look at the shape of this. One row per student, one column for the group they're in, one column for their score. People call that long format, and it's exactly how ANOVA likes to see the world. If your groups were sitting in separate columns right now, you'd reshape first. But us? We're ready to go.
Before we run anything, let's actually look at our scores. I'll click Descriptives and Explore and pick Exam Score. Right away Quanta gives me the lay of the land: 90 valid cases, zero missing, a mean of about 79, a median of 80 sitting right on top of each other. That's already a good sign. When the mean and the median agree like that, your data is probably nice and symmetric. And look at the skewness, .06. That's about as close to perfectly balanced as real data ever gets. You can see it in the histogram and that tidy little box plot too. No long tail dragging off to one side.
But I don't want to just eyeball it, and neither should you. So I'll scroll down to the normality card. This is the assumption a lot of tests, ANOVA included, quietly lean on: that your scores follow a roughly normal, bell-shaped pattern. Quanta runs the Shapiro-Wilk, and here's the result: W of .986, p of .461. Now here's the part that trips people up, so stay with me. With a normality test, a big p value is the good news. A small p would mean your data departs from normal, but .461 is comfortably large, so there's no significant departure. Our scores are normal enough. And just to be thorough, Quanta runs a second check, the Lilliefors, right beside it, and it agrees. Both green lights. The normality box is checked and we can head to the test itself with confidence.
I'll head to Compare Means, and in the center panel you'll notice four kinds of compare-means analysis to choose from: the t test, the paired t test, the one-way ANOVA, and ANOVA. Here we want the one-way ANOVA. I set the outcome to Exam Score, the grouping factor to Method, and once we pick our variables, that's it. The program populates the analysis right in the center panel.
Our F is 18.49. Remember, F is signal over noise, so a number sitting that far above one is telling us loud and clear that the gaps between these classrooms dwarf the scatter inside them. Right beside it are the degrees of freedom, 2 and 87. The two is your number of groups minus one. The 87 is your students minus your groups, 90 minus 3. You always report both, because F doesn't mean much without them. Then the p value. It's way under .05, so far under that Quanta just shows it as less than .001. Here's what that's actually saying: if these three methods truly made no difference at all, you'd see a gap this big by pure chance less than one time in a thousand. That's strong evidence the method matters. And the number of cases, 90, confirms the analysis ran over all 90 students in the study.
Here's what I love about Quanta. It doesn't just hand you an F and walk away. It lays out the whole story in plain cards, top to bottom, and reads it back to you. Let's go down the stack. First, the plain-language result. It tells you in words that there's a statistically significant difference between the groups, and it even names the test it picked and why. Right under it, the analysis decision card: ReliCheck used the classic one-way ANOVA because the homogeneity check didn't flag unequal variances. That's the app showing its work.
Next, the descriptives group by group. Lecture at 74, blended at 80, active at almost 85, 30 students each. The very picture we hoped for. And right below, the homogeneity check, whether our groups have a similar spread. Levene comes back at .93, Brown-Forsythe at .91, both miles from trouble. So our variances are even, and that's exactly why Quanta reached for the classic ANOVA. Then the omnibus tests, the full between-groups and within-groups table behind our F. And see how Quanta runs the backups right alongside: Welch, Brown-Forsythe, Kruskal-Wallis. Every one of them agrees, every one significant. When all the versions point the same way, you can trust the answer.
Now this card right here is the one that earns its keep, the post-hoc comparisons. Because here's the thing the F never told us. It said the three methods aren't all the same, but it would not say which ones differ. Maybe one classroom ran away with it and the other two tied. Maybe all three are truly apart. The post-hoc is where we find out. Quanta runs Tukey HSD here, the Tukey-Kramer version that handles unequal group sizes, with familywise-adjusted p values and intervals straight from the studentized range distribution. In plain terms, it compares every pair while holding your overall error rate in check, so you don't undo the very thing the ANOVA just protected you from.
Let's read the three rows, because there's more here than a yes or no. Active versus blended: the means are about five points apart, the adjusted p is .017, significant. And look at that last column, Hedges' g, the effect size for the pair: .72, a solid gap. Active versus lecture: 11 points apart, the biggest jump on the board, p under .001, and a Hedges' g of 1.56. That's a huge effect. And blended versus lecture: about six points, adjusted p of .04, Hedges' g of .84. So every single pair clears the bar, and Quanta hands you the 95% confidence interval on each one too, so you can see the range the true difference likely lives in. This is the strong version of the story: active beats blended, blended beats lecture, and that ladder is real at every rung, not just at the top and bottom.
Then the overall effect sizes card, and this is the part I really want you to carry with you. Eta squared at .30. A small p tells you an effect is there. It does not tell you how big. This says about 30% of the difference in exam scores rides on which classroom a student was in. In education research, that's large. So this isn't just real, it's big enough to care about.
And see those assumption warnings near the bottom? Quanta isn't hiding them. It's nudging you to glance at a sensitivity check before you report. That's the app keeping you honest. And last, the report-ready summary: the whole finding in clean sentences, ready to lift into your writeup.
Let's take a breath and look at what we just did. We started with a real question: do these three teaching methods actually change exam scores? We picked the test, and Quanta laid out the whole story for us. The F, the assumptions, the pairwise comparisons, the effect size. We found a strong, real difference: F of 18.49, p under .001, and an effect big enough to matter. Every pair of methods truly differs. And we wrote it up honestly, as an association, not a cause.
That's a one-way ANOVA start to finish in Quanta. And if it clicked for you, that's the feeling we're after. I'm Cesaly. Thanks for spending this time with me. Next time, we'll take this very same data and add a second factor, so we can ask a richer question: not just whether the method matters, but whether it matters differently for different kinds of students. I'll see you there.
Transcript lightly edited for reading. ReliCheck Quanta is statistical software for the Mac with a 30-day free trial: relicheck.com/quanta. The full analysis list is on the analyses page and the validation record on the validation page.