Storytelling through Captions
Captions that carry feeling, not just words.
The team behind "Magic Words."
Product & Design Lead
Full Stack Engineer
Full Stack Engineer
Disney's promise is emotional — but that promise doesn't reach everyone equally.
The gasp at the reveal. The lump in your throat at the goodbye. The chill as the score swells. Disney doesn't just tell stories — it engineers emotional experiences, frame by frame, line by line.
For DHH viewers, neurodivergent viewers, and anyone watching without sound, the emotional signal gets stripped out before it ever reaches them.
Captions give you the words, but they don't give you the emotion or depth.
Tone of voice, the tremor in a whisper, the score swelling, the flinch at a scream — doing enormous emotional work underneath the dialogue.
The exact same words — uniform white text, same font, same size — no matter what's happening emotionally on screen.
Two viewers can watch the exact same scene — one hearing it, one reading it — and have measurably different emotional experiences of the same story.
That gap is invisible to most of the audience, which is exactly why it's been able to persist this long.
Captions shouldn't just carry information — they should carry feeling.
Instead of uniform white text, captions shift in color, weight, motion, and rhythm to match the emotional intensity of the scene. The words stay accurate — nothing about meaning changes. The feeling just isn't stripped out in transit.
Shifts with the emotional tone of the line — a whispered secret renders differently than a scream.
Intensifies with emphasis — text gets heavier as the moment gets more intense.
Mirrors the pacing of the scene — steady for calm, urgent for tension.
Anchor scene: The Incredibles — Violet and Dash escape the jungle ambush.
Same word. Same scene. One version captures the fear — the other doesn't.
Deaf and hard-of-hearing viewers are the center of this idea — not an afterthought.
For a DHH viewer, captions are the entire channel — not a backup to tone of voice, music, and sound design. When that channel is flat, they get a structurally different, quieter, less urgent version of the story. A scream and a sigh currently look identical on screen. Magic Words closes that gap directly, at the level where it actually lives.
A huge and growing share of all viewing happens with the sound off — on a plane, on transit, in any shared space. Every one of those viewers has the same problem as a DHH viewer, they just don't think of it as a caption problem. Magic Words is a genuine upgrade to how anyone experiences a story without sound.
This isn't a niche edge case — it's a mainstream viewing experience.
of mobile video is watched with the sound off
Verizon Media / Publicis Media
more likely to watch a full video when captions are on
Verizon Media / Publicis Media
Americans report some trouble hearing
NIDCD
A working prototype built in one hackathon night — and the pipeline that scales it across the catalog.
Pulled dialogue, timing & speaker data from 3 repos — including Web Player — into one hardcoded JSON.
Ran our real Magic Words Claude skill on that JSON to tag emotion, intensity & delivery.
HiVE subtitle-renderer applies emotion styles (color, animation, weight) per phrase during playback.
New subtitle file lands in S3, triggers EventBridge.
Claude (Bedrock) runs the Magic Words skill — tags emotion + delivery.
Validated JSON stored in S3 alongside subtitle assets.
CloudFront serves JSON from the edge on playback start.
Web Player fetches the JSON, passes it to the HiVE subtitle renderer.
HiVE styles each phrase by emotion — color, animation, weight.
S3 • EventBridge • Step Functions • Lambda • Bedrock (Claude) • DynamoDB • CloudFront
A few directions this could grow in.
Move from core categories to nuanced intensity, mixed emotions, and finer-grained delivery.
User-adjustable intensity, reduced-motion, and color-blind-friendly modes.
Scale from one hand-built demo scene to a transcript-driven pipeline across the catalog.
Magic Words — Storytelling through Captions