Two models may share an enclosure and differ only in a label. Retrieving them together is unsurprising; treating them as interchangeable products is the error. This exercise uses deliberately simple owned graphics, not customer photography or an invented commercial case.
Materials and fixed environment
The fixture contains four 384-by-384 images: two red circles labeled AX-220 and AX-110, a blue square labeled BX-220 and a green triangle labeled CX-110. These are teaching graphics, not product photographs. We executed openai/clip-vit-base-patch32 on CPU with a pinned repository revision.

model-results.json retains raw logits and complete rankings; benchmark.json identifies the images. Model weights are not redistributed in the package. Reproduction downloads the same revision from the model publisher's repository.
What the five queries returned
| Query | First candidate | Interpretation |
|---|---|---|
| a red circle | red-circle-110 | Two images satisfy appearance; no model specified |
| a blue square | blue-square-220 | Correct appearance candidate in this fixture |
| a green triangle | green-triangle-110 | Correct appearance candidate in this fixture |
| AX-220 red circle | red-circle-220 | Named model ranked first here |
| AX-110 red circle | red-circle-110 | Named model ranked first here |
This set does not show a top-rank failure for the named queries. We do not invent one. The ambiguous red-circle query nevertheless demonstrates that appearance does not uniquely identify a product. Blur, cropping, real machinery and cross-language image queries were not evaluated, so these results are not catalog accuracy estimates.
A similarity score is not a probability
A CLIP logit supports comparison within a particular setup. It is not the probability that a product is correct. Preprocessing, model revision and candidates affect scores. A value near 31.99 must not be reported as 31.99% recognition confidence, and scores from unrelated models should not be added directly.
The Sentence Transformers image-search guide explains shared image-text representations. Our implementation loads a pinned CLIP component through Transformers. Its model card identifies the model source; this is not a proprietary Zhihe Growth foundation model.
Design a B2B relevance key
Separate visual relevance, identity match and attribute support. A visually relevant candidate may be useful for human review; it does not support a procurement conclusion until identity and factual evidence match. Each image needs a model, configuration, provenance and permission record.
| Layer | Positive requirement | Negative example |
|---|---|---|
| Appearance | Requested color, shape or scene | Green triangle for a blue-square query |
| Identity | Specified model is established | Same enclosure, different supply |
| Specification | Source supports the field | Guess tank capacity from appearance |
| Suitability | Conditions and restrictions match | Recommend wet-area use from a photo |
A practical benchmark adds similar models, unreadable labels, blurred images and no-answer questions. Split development and validation by independent entities. Four simple graphics do not establish generalization to a real catalog.
Combine candidates with explicit gates
Retrieve by textual model ID and visual similarity separately, then filter by identity and attach evidence. Use calibrated scores or a defined rank-fusion procedure rather than arbitrary addition. A visual match with a conflicting identifier should be escalated, not allowed to overwrite the identifier.
This package executes encoding and ranking, not a complete product-recognition service. media_labs.py reproduces the inputs and outputs. The candidate-combination design is an implementation approach, not a claim that the feature is deployed in a customer backend.
Zhihe Growth's application
Zhihe Growth can help Chinese exporters associate images with explicit product entities, versions and approved descriptions. Acceptance should distinguish “found a similar image” from “confirmed this model is suitable” and check whether mismatches are blocked.
Continue with evidence locations, unit and condition checks and hybrid retrieval. The technical FAQ provides direct answers. The advanced package preserves all graphics, code and rankings. This experiment does not evaluate commercial AI image search or contribute to customer citation rates.