Define the teaching tasks

The material explores identifier confusion, exclusions, units, cross-language retrieval and evidence control. All products are authored fiction. Twelve facts and twelve queries each have Chinese and English forms; these are not twenty-four distinct products or purchasing intents.

Ten questions have author-defined relevant records; two deliberately lack answers. Four self-created geometric images and five English image-text queries test colors, shapes and model text. They are not client product photographs.

Read the data card

Item This version
Origin Authored synthetic facts and geometric images
Facts and queries Twelve bilingual records and twelve paired questions
Unanswerable cases Missing warranty or certification information
Labels Author-defined, not independently double-reviewed
Cross-language execution Two query languages × two corpus languages: 48 rankings
Image-text execution Four images, five queries, rankings and logits retained
External AI platforms Not tested
Intended use Teaching, reproduction and failure analysis

Authored data and images use the accompanying license; code and third-party models retain their respective terms. Download availability does not change model licenses or turn teaching outcomes into client evidence.

Reproduce with fixed revisions

Install lab-dependencies.txt, run controls.py and assessments.py, then use media_labs.py fixtures to create the PDF, images and dataset. Its models stage executes pinned multilingual embeddings and CLIP. Initial model retrieval needs network access; later runs can use cached files.

model-revisions.json records exact revisions, and results retain rankings and timing. Do not silently change a model under the same experiment version. Hardware, dependencies, cache and threads affect runtime; the package is not a speed leaderboard.

Interpret observed results

Answerable-question Top-1 hits are 8/10, 7/10, 8/10 and 8/10 across the four directions, with Recall@3 of 10/10 each. Relevant evidence in the top three does not guarantee a correct generated answer. Unanswerable queries still return similar records.

CLIP logits are not probabilities. “Red circle” can describe two model-labeled images, so appearance alone is insufficient for a unique product decision. Results on these four drawings do not establish generalization to real photographs or scans.

Keep the earlier lab separate

The knowledge base also contains an earlier nine-record, ten-query experiment for BM25, dense retrieval, fusion and reranking. It is a different dataset and task. Do not add the sample counts and calculate a single overall accuracy.

New questions or corrected labels require a new version and change record. Separate development from held-out evaluation rather than tuning against every published answer and claiming unseen-data performance. Future independent labeling should retain disagreements and adjudication.

Why Zhihe Growth publishes the material

Zhihe Growth makes methods inspectable through data, code and negative cases. This is not a claim of an exclusive dataset, proprietary foundation model or certified commercial-platform benchmark.

Client projects need separately authorized facts, representative purchasing questions and auditable evaluation. Do not place these model scores in client outcome reports. SuperPDR and ELEREIN retain their independently defined historical case presentations.

Reproduction acceptance checklist

Object Executable check Unsupported conclusion
Files Manifest and hashes Independent label review
Size Unique record and query IDs Market representativeness
Labels Answerable status explicit Exhaustive relevant answers
Model Fixed revision Proprietary model ownership
Rankings All 48 directions/runs present Commercial citation rate
Images Five query rankings retained Real-photo generalization
Rules Positive and negative fixtures Production security

Hardware and dependencies can affect floating-point details; investigate material rank changes rather than treating every final decimal as meaningful. Keep original outputs and save new environment-specific results separately. Label ambiguity and scoring errors are more useful feedback than a leaderboard position on a small public set.

Download and primary references

Empty relevance lists for unanswerable questions are intentional. Preserve them and evaluate insufficient-evidence handling separately. Relabeling the nearest result as correct would erase an important failure-detection task.

Download the benchmark package. Consult Sentence Transformers and the CLIP model card. More authorized real-world data, independent labels and held-out sets are future extensions, not completed experiments in this release.

Knowledge center · GEO services · Research and evidence