# A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions

> Research article (2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024) · cited 15× · AI/ML

**Wikidata**: [openalex:W4402716204](https://www.wikidata.org/wiki/openalex:W4402716204)  
**Source**: https://4ort.xyz/entity/a-picture-is-worth-more-than-77-text-tokens-evaluating-clip-style-models-on-dense-captions
