Papers/2608.12549
🧪 Test?View on arXiv

StrAD: A Streaming Method and Benchmark for Audio Description Generation for Long-form Videos

Not specified in the content

streamingaudio descriptionvideo captioningaccessibility
2608.12549
Builder Relevance
80%
Aug 14

Abstract

The paper presents StrAD, a benchmark and method for generating audio descriptions for long-form videos, addressing accessibility for blind and low-vision individuals.

Reality Card

Core Claim

StrAD is the first streaming approach for generating audio descriptions on the fly without requiring ground-truth timestamps.

Method / Result

StrAD-FT sets the state of the art on CMD-AD with a CIDEr score of 36.3.

Limitations

Both StrAD-FT and StrAD-Zero exhibit limitations in temporal localization and narrative coherence.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers