Ask a Question

Prefer a chat interface with context about you and your work?

Pic2Word: Mapping Pictures to Words for Zero-shot Composed Image Retrieval

Pic2Word: Mapping Pictures to Words for Zero-shot Composed Image Retrieval

In Composed Image Retrieval (CIR), a user combines a query image with text to describe their intended target. Existing methods rely on supervised learning of CIR models using labeled triplets consisting of the query image, text specification, and the target image. Labeling such triplets is expensive and hinders broad applicability …