Other
OpenAI
About this model
CLIP (Contrastive Language-Image Pre-Training) is a neural network trained on a variety of (image, text) pairs. It can be instructed in natural language to predict the most relevant text snippet, given an image, without directly optimizing for the task, similarly to the zero-shot capabilities of GPT-2 and 3. We found CLIP matches the performance of the original ResNet50 on ImageNet “zero-shot” without using any of the original 1.28M labeled examples, overcoming several major challenges in computer vision.
Tags
Related Models
Similar AI models you may like
Other
Other
【TOOL】ComfyUI Installer
⭐ 0.0
⬇ 21,669
Other
NoobAI
Danbooru/e621 autocomplete tag lists incl. aliases (+ Krita AI support)
⭐ 0.0
⬇ 10,429
Other
ZImageTurbo
Z-Image Uncensored Text Encoder - Abliterated Huihui Qwen3 4B v2 (Q_8 GGUF)
⭐ 0.0
⬇ 8,763
Other
Other
assDetailer, aDetailer model for butts, rears, pants, and panties
⭐ 0.0
⬇ 6,194
Other
Other
Mask aDetailer - Face detailer for Eyes, Eyebrows, and Nose
⭐ 0.0
⬇ 6,126
Other
Other
FLAN-T5-XXL (Text-Encorder only)
⭐ 0.0
⬇ 5,917
Other
Flux.1 D
【FLUX】ComfyUI Installer
⭐ 0.0
⬇ 5,106
Other
Other
2x-AnimeSharpV4
⭐ 0.0
⬇ 4,666