# CLIPSonic: Text-to-Audio Synthesis with Unlabeled Videos and Pretrained Language-Vision Models

> Research article (2023 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2023) · cited 17× · AI/ML

**Wikidata**: [openalex:W4386764084](https://www.wikidata.org/wiki/openalex:W4386764084)  
**Source**: https://4ort.xyz/entity/clipsonic-text-to-audio-synthesis-with-unlabeled-videos-and-pretrained-language-vision-models
