AccessibilityTestingUpdated Jul 27, 2026
SceneCast
SceneCast transcribes audio or video into text. Most transcripts stop at the words, losing who spoke, the sounds in the room, and where one scene ends and the next begins.
Why this matters
Understanding a piece of media is not only about the words in it. A transcript that preserves speaker changes and non-speech sound gives someone the same footing as a person who can hear it.
- Current idea
- SceneCast transcribes audio or video into readable, downloadable text. Alongside the speech it identifies who is speaking, tags sound effects and music, and marks scene changes, so the transcript reflects what happened rather than only what was said.
- Accessibility notes
- Speakers, sound effects, music, and scene markers are distinguished by label rather than color alone, so the transcript survives being read aloud by a screen reader or exported to plain text. Output is downloadable, and the transcript view is keyboard navigable with strong contrast in both light and dark.
- What was learned
- A transcript can capture too much as easily as too little. The trade-off between completeness and readability is the whole product.
Development updates
Progress in public
Speaker identification, sound-effect tagging, and scene-change markers released.
Core speech-to-text pipeline and downloadable transcript view tested.
Feedback
Have thoughts on this project?
Head to the Ideas board to share a related barrier or idea. Attribution is optional.
Open the Ideas board