Reading a Solution vs. Watching One: Which Helps More on a Hard Problem?
You're stuck on part (c), it's late, and you paste the problem into ChatGPT. Back comes a complete, correct, well-organized answer: nine steps, each with a short justification, some of it in LaTeX that didn't render properly, all of it on your screen at once.
You read it. It makes sense while you're reading it. You close the tab and cannot reproduce step four.
That specific failure — comprehension that evaporates on contact with a blank page — is common enough that it's worth asking whether the format of an explanation, and not only its content, is doing something. The answer from learning research is yes, with conditions that are more interesting than the marketing version.
<br>
1. What the research actually supports (and what it doesn't)
Richard Mayer's work on multimedia learning is the standard reference here, and it gets over-quoted in one direction. The modality principle does find that people learn better from graphics with spoken narration than from graphics with on-screen text — the reasoning being that narration and visuals use separate processing channels, so you're not asking your eyes to read and look at the same time.
Here's the part that usually gets dropped, and it matters for exactly the case above: the modality effect has boundary conditions, and worked mathematical solutions sit inside several of them. The principle is documented as unlikely to apply when the material is long and complex, contains technical terms or symbols, is not in the learner's first language, or is paced by the learner rather than the system. A multi-step derivation full of notation, that you want to pause and re-read, is close to a worst case for the narration-beats-text claim.
So the honest version isn't "video is better." It's that a different principle is doing the work: segmenting — breaking an explanation into learner-paced chunks rather than delivering it whole. A nine-step solution presented as nine steps you advance through is easier to hold than the same nine steps arriving simultaneously, regardless of whether the words are spoken or written.
That reframing changes what to look for. Not "is there a video," but: does the explanation arrive in pieces I control, and can I see the work being produced rather than presented finished?
<br>
2. Where a wall of text actually breaks down
Four specific situations, all common in STEM homework:
Notation that doesn't render. LaTeX in a chat window is a lottery. Subscripts collapse, fractions flatten, matrices become unreadable rows. You end up decoding the formatting before you can decode the mathematics, and that decoding consumes exactly the working memory the problem needs.
Steps whose order matters more than their content. In a derivation, why this move now is often the whole lesson. Text presents all the moves at once, which makes the sequence invisible — you see the destination and the path simultaneously, and the path is what you needed.
Diagrams described instead of drawn. "Consider the free-body diagram with tension resolved into components" is a sentence. A free-body diagram is a picture. Prose descriptions of spatial reasoning are a translation you have to undo.
Anything you'd want to pause. Reading is technically self-paced, but a wall of text invites skimming, and skimming a solution is how you arrive at "it made sense" without arriving at "I can do it."
Sovi's Video Explanation under Ask Sovi is aimed at that shape: you upload a photo of the problem or type it, and get back a walkthrough of about a minute in which the solution is worked through in sequence, with the notation shown rather than typed into chat formatting. The brevity is part of the design — a minute is roughly what one stuck step is worth, and it's short enough that you'll actually rewatch it. The point isn't that watching is easier than reading; it's that a sequenced explanation preserves the order of moves, which is the part a wall of text loses.
<br>
3. Where text is genuinely better
Being straight about this, because the honest answer is "often."
| Situation | Better format | Why |
|---|---|---|
| You need one specific step, not the whole solution | Text | Scannable. You can jump to step 4 without watching steps 1–3 |
| Conceptual explanation with no notation | Text — ChatGPT, Gemini and Claude are excellent here | Fast, re-promptable, no formatting problems |
| You want to argue with the explanation | Text | You can quote a line back and ask why |
| Multi-step derivation with heavy notation | Video / segmented | Sequence and symbols both survive |
| Something spatial or diagrammatic | Video / segmented | Drawn beats described |
| Reviewing something you already understood once | Text | Faster, and re-derivation isn't the goal |
| Non-native English and dense technical prose | Depends — try both | The modality research specifically flags this case as unsettled |
Note the second row. For pure conceptual explanation — "what actually is a p-value" — a conversational assistant is hard to beat, because you can immediately say "simpler," "give me an analogy," "now quiz me," and iterate. That back-and-forth beats production quality every time, and it's the highest-value use of general AI in studying.
<br>
4. The habit that makes either format work
Format helps. It doesn't substitute for the thing that actually converts an explanation into an ability:
- Attempt first, badly. Ninety seconds of real attempt before you ask anything. This turns "explain question 6" into "I set up the equation but can't resolve the tension," and a targeted explanation lands in a gap that already exists.
- Predict the next step before it arrives. In a segmented explanation you can do this literally — pause, guess, then advance. Where your guess and the solution diverge is the thing you didn't know, isolated from the seven things you did.
- Close everything and redo it cold. Blank page, from the problem statement. People skip it, and it's the only part that produces retention. If you can't reproduce it, you watched rather than learned — go back to the step you couldn't predict.
- Do the neighbouring problem. Textbooks cluster similar problems deliberately. Solving the next one unaided is the actual proof.
None of that is format-specific. It's what makes any explanation worth the time you spent on it.
<br>
5. Four ways explanations get wasted
- Consuming the whole thing when you needed one step. If you already have steps 1–3, don't re-watch or re-read them. Find the joint.
- Mistaking fluency for understanding. A well-produced explanation feels clearer than a confusing one whether or not you absorbed it. The blank page is the only honest test.
- Never verifying. Whatever the format, check the result: units, a limiting case, substituting the answer back. Multi-step work is where errors hide, and a confident explanation looks identical to a correct one.
- Watching passively. Video invites spectating in a way text doesn't. Pause, predict, and write things down, or the format advantage disappears entirely.
<br>
Reference
- Mayer's Principles of Multimedia Learning — 12 principles overview (NYU)
- Cognitive Theory of Multimedia Learning — Modality Principle and its boundary conditions
<br>
Frequently Asked Questions
Q1: Are video solutions better than text explanations for math?
Not universally, and the research is more specific than the marketing. Mayer's modality principle supports narration over on-screen text under certain conditions, but it's documented as unlikely to apply to material that's symbol-heavy, technical, or learner-paced — which describes most worked mathematics. What does help is segmenting: receiving an explanation in steps you advance through rather than all at once. Look for that property rather than for video specifically.
Q2: Why is LaTeX from ChatGPT so hard to read?
Chat interfaces render mathematical markup inconsistently, so fractions, subscripts and matrices often arrive flattened or broken. The cost isn't cosmetic — decoding malformed notation uses the same working memory you need for the actual problem. If you're getting unreadable output, ask for the solution described in words plus a clean final expression, or use a format that shows the notation properly.
Q3: When should I just use ChatGPT, Gemini or Claude instead?
Whenever the thing you're stuck on is conceptual rather than procedural, and whenever you want to argue with the answer. General assistants are excellent at explaining an idea at whatever level you ask for, re-explaining it differently on request, and quizzing you afterward. That back-and-forth is genuinely hard to beat, and it doesn't need any special format.
Q4: I understand every solution I read and still fail exams. What's wrong?
You're studying by recognition and being tested on production, which are different acts. Reading a solution feels like understanding because it is understanding — of a passive kind that doesn't survive a blank page. The fix is mechanical rather than motivational: after every problem you needed help on, close everything and redo it from the problem statement. If you can do that twice, a few days apart, you'll have it in the exam room.
Continue Your Learning with Sovi.AI
Sovi.AI is your free AI study buddy for step-by-step explanations, document-based learning, and AP exam prep. Put what you just read into practice with the tools below:
- Ask Sovi — upload a photo to open the Ask Sovi chat and get a clear, step-by-step AI homework explanation across math, science, and writing.
- AI Study — upload your draft or source PDFs to outline arguments, generate cheatsheets, and revise faster.
- AP Test Prep — drill timed AP questions with full mock exams and unit-level practice across every AP subject.
- Practice Resources — browse expert-verified study guides across Math, Biology, Chemistry, History, and more.
Looking for more guides like this one? Visit the Sovi.AI Blog for writing tips, grammar walkthroughs, and study strategies.