CASE STUDY / CREATIVE TOOLS / 2026
Frame
Caption
Studio.
A caption workflow built around the part automation can't finish: the human edit.
Captions need
a second pair of eyes.
I’m interested in videography and cinematography. Automatic transcription can give an editor a starting point, but it can still get names, punctuation, or timing wrong. I wanted the review step to be part of the tool, where I could watch the source and adjust the captions together.
This project focuses on a clear post-production handoff: take a media file, generate a draft transcript, correct it, and leave with a subtitle file ready for a video editor or platform.
From raw speech
to editable subtitles.
Choose video or audio in MP4, MOV, M4V, MP3, M4A, WAV, or WebM format. The app accepts files up to 800 MB.
Choose a tiny, base, or small faster-whisper model. Speech is transcribed on the computer through the local Python server.
Play the media, jump to a caption, then correct its words and timestamps. Add or remove segments when needed.
Validate the timestamps and download an SRT or WebVTT subtitle file for the next editing step.
A first run with
sample audio.
The repository includes a short spoken sample, so you can try the complete workflow without finding your own footage. Run the app locally, open examples/first-run.wav, select the Tiny model, review the draft against the recording, and export an SRT file. The first run downloads the transcription model.
A few choices
that shaped the tool.
Keep footage on the computer.
The server binds to 127.0.0.1, accepts the media into a temporary file, transcribes it, then removes that file. The app has no account or database. The first model download does require an internet connection.
Treat model output as a draft.
Instead of exporting automatic captions immediately, the interface puts editable text and timestamps beside playback. Selecting a caption takes the viewer to its start time, so checking it against the source is less cumbersome.
Export interoperable files.
SRT and WebVTT can be used in other editing and publishing workflows. The app checks caption timing before it creates the download, so invalid segments can be fixed first.
Small, local
architecture.
The browser holds the current editing session in memory. The application does not send your footage to a transcription API; the chosen model is downloaded on first use.
What works.
What’s next.
The result connects transcription to a usable review and export workflow. It is a local tool for producing subtitle files, rather than a hosted service or a video renderer.
Current boundaries
- Caption quality depends on speech clarity, background audio, accents, and names.
- Long model segments may need manual splitting.
- Edits are lost if the tab closes before export.
- Long footage can take substantial time to process on a CPU and needs space for a temporary copy.
Next iteration
- Use word timestamps to suggest readable caption lengths.
- Save a local draft so edits survive a refresh.
- Add a clearer processing state for long files.
- Gather feedback from people who caption their own videos before adding more features.
The 800 MB figure is the configured upload limit; it is not a claim that every large file transcribes quickly or that the app has been benchmarked across formats.