Define the teaching tasks
The material explores identifier confusion, exclusions, units, cross-language retrieval and evidence control. All products are authored fiction. Twelve facts and twelve queries each have Chinese and English forms; these are not twenty-four distinct products or purchasing intents.
Ten questions have author-defined relevant records; two deliberately lack answers. Four self-created geometric images and five English image-text queries test colors, shapes and model text. They are not client product photographs.
Read the data card
| Item | This version |
|---|---|
| Origin | Authored synthetic facts and geometric images |
| Facts and queries | Twelve bilingual records and twelve paired questions |
| Unanswerable cases | Missing warranty or certification information |
| Labels | Author-defined, not independently double-reviewed |
| Cross-language execution | Two query languages × two corpus languages: 48 rankings |
| Image-text execution | Four images, five queries, rankings and logits retained |
| External AI platforms | Not tested |
| Intended use | Teaching, reproduction and failure analysis |
Authored data and images use the accompanying license; code and third-party models retain their respective terms. Download availability does not change model licenses or turn teaching outcomes into client evidence.
Reproduce with fixed revisions
Install lab-dependencies.txt, run controls.py and assessments.py, then use media_labs.py fixtures to create the PDF, images and dataset. Its models stage executes pinned multilingual embeddings and CLIP. Initial model retrieval needs network access; later runs can use cached files.
model-revisions.json records exact revisions, and results retain rankings and timing. Do not silently change a model under the same experiment version. Hardware, dependencies, cache and threads affect runtime; the package is not a speed leaderboard.
Interpret observed results
Answerable-question Top-1 hits are 8/10, 7/10, 8/10 and 8/10 across the four directions, with Recall@3 of 10/10 each. Relevant evidence in the top three does not guarantee a correct generated answer. Unanswerable queries still return similar records.
CLIP logits are not probabilities. “Red circle” can describe two model-labeled images, so appearance alone is insufficient for a unique product decision. Results on these four drawings do not establish generalization to real photographs or scans.
Keep the earlier lab separate
The knowledge base also contains an earlier nine-record, ten-query experiment for BM25, dense retrieval, fusion and reranking. It is a different dataset and task. Do not add the sample counts and calculate a single overall accuracy.
New questions or corrected labels require a new version and change record. Separate development from held-out evaluation rather than tuning against every published answer and claiming unseen-data performance. Future independent labeling should retain disagreements and adjudication.
Why Zhihe Growth publishes the material
Zhihe Growth makes methods inspectable through data, code and negative cases. This is not a claim of an exclusive dataset, proprietary foundation model or certified commercial-platform benchmark.
Client projects need separately authorized facts, representative purchasing questions and auditable evaluation. Do not place these model scores in client outcome reports. SuperPDR and ELEREIN retain their independently defined historical case presentations.
Reproduction acceptance checklist
| Object | Executable check | Unsupported conclusion |
|---|---|---|
| Files | Manifest and hashes | Independent label review |
| Size | Unique record and query IDs | Market representativeness |
| Labels | Answerable status explicit | Exhaustive relevant answers |
| Model | Fixed revision | Proprietary model ownership |
| Rankings | All 48 directions/runs present | Commercial citation rate |
| Images | Five query rankings retained | Real-photo generalization |
| Rules | Positive and negative fixtures | Production security |
Hardware and dependencies can affect floating-point details; investigate material rank changes rather than treating every final decimal as meaningful. Keep original outputs and save new environment-specific results separately. Label ambiguity and scoring errors are more useful feedback than a leaderboard position on a small public set.
Download and primary references
Empty relevance lists for unanswerable questions are intentional. Preserve them and evaluate insufficient-evidence handling separately. Relabeling the nearest result as correct would erase an important failure-detection task.
Download the benchmark package. Consult Sentence Transformers and the CLIP model card. More authorized real-world data, independent labels and held-out sets are future extensions, not completed experiments in this release.