“Is this machine suitable?” may involve selection, technical validation or budget approval. Label the observable request. Do not infer the buyer's employer, purchasing power or personal characteristics. Unknown is a useful value when the question does not provide enough information.

Define an operational dictionary

A question such as “Can this machine handle oily floors, and what is the minimum order?” contains suitability verification and procurement conditions. Preserve the original and identify both tasks rather than letting the final quantity phrase erase applicability. Leave the user's role unknown when it is not supplied. Adjudication should identify supporting text and whether multiple labels or subquestions are appropriate. Declare whether reporting counts original questions or subtasks; splitting alone can change category proportions.

Use one primary intent for stratified reporting and additional controlled labels for risks or constraints. Each definition needs inclusion rules, exclusions and its closest confusing category.

Label Includes Excludes Preferred evidence
discover Concepts, categories and tasks A named model's specification Definition and scope
validate Compatibility, conditions and parameters Price-only comparison Manual and test conditions
compare Alternatives and supplier trade-offs One model's factual lookup Like-for-like matrix
procure MOQ, lead time and quotation General technical learning Configuration-specific terms
support Faults, maintenance and consumables Initial category discovery Maintenance and service records

Branded versus unbranded is a separate dimension. A named model voltage question is technical validation; an open question about equipment for indoor floors is task discovery. Combining them obscures differences in both test purpose and expected visibility.

Control multilabel decisions

Require a primary label and allow a bounded list of secondary risks such as electrical, surface or delivery. Do not create a new synonym for every difficult question. When suitability and quotation appear together, record both and explain which unresolved prerequisite controls the primary label.

Fill the role only when explicit. “As a distributor” supports that role; an MOQ question alone does not. Record region, language, model and budget from explicit evidence. Unresolved cases enter review rather than being silently overwritten by the latest annotator.

Use positive and negative examples

Synthetic question Primary label Secondary concern Common error
Can AX-220 operate in wet areas? validate Environment and safety Treat as a generic introduction
Which option has lower three-year cost? compare Cost and time horizon Compare purchase price only
May I order four cartons of DW-1? procure MOQ and packaging Check pieces per carton only
Why is the machine not recovering water? support Fault and downtime Recommend a new product immediately

These examples are authored teaching cases, not customer inquiries. Correct categorization does not establish that evidence exists to answer the question.

Separate independent annotation from adjudication

Train reviewers on the dictionary, then ask them to label the same questions independently. Freeze their original labels before discussing disagreements. Retain both labels, the rule invoked, adjudicated result and dictionary version. A changed definition may require relabeling affected historical examples.

The material package includes review-templates.json with independent-review fields left empty. The execution record's annotation_demo contains two deliberately constructed label sequences. It tests scoring logic; it is not evidence that two human reviewers completed annotation. Five of its six labels agree, giving raw agreement of 5/6. Cohen's kappa adjusts for agreement expected from the marginal label distributions.

Understand the measurement limits

With imbalanced categories, repeatedly choosing the largest category can produce misleading agreement. Inspect the confusion table and category-level disagreements. For multilabel tasks, specify whether scoring requires exact set agreement or evaluates each label separately.

The scikit-learn reference defines kappa. Our small executable calculation does not establish a universal acceptance threshold. If both sequences contain only the same category, expected agreement is one and the kappa denominator is zero; the implementation returns null rather than manufacturing a perfect score. Agreement also cannot demonstrate that the taxonomy itself is useful.

How Zhihe Growth applies the method

For Chinese exporters, Zhihe Growth can map intent to a knowledge guide, product fact, FAQ or evidence record and distinguish an explanation gap from an evidence gap. Customers should sample high-risk questions during acceptance instead of relying on an automated clustering image.

Continue with topic clustering, rewrite constraints and question decomposition. The short FAQ handles direct questions; the execution record makes the calculation inspectable. These labels support an editorial workflow, not a claim about proprietary platform classification.

Knowledge center · GEO services · Research and evidence