Decide whether AI is worthwhile for a small feedback batch, define useful categories, review ambiguous comments and compare the full effort with manual coding.

Direct answer: For a small, one-off batch, read and categorise the comments yourself before adding an AI tool. Use AI when the categories are already clear, the information is approved for the service, and reviewed classification takes less effort than the manual alternative. Keep multiple labels and an uncertain category available, and do not use a positive or negative sentiment score as a substitute for understanding what the customer needs.

The useful outcome is not a colourful chart. It is a decision such as correcting delivery instructions or investigating a payment failure. Classification that hides the reason for a complaint can make your reporting look more organised while making that decision harder.

AI becomes more plausible when you repeat the same bounded task often enough to repay the setup. Even then, label quality depends on your definitions and review. This article narrows the broader knowledge-work workflow to one job: deciding whether assisted labelling is worthwhile for a small feedback set.

Applies to: low-consequence analysis of short feedback comments, not automated decisions about individual customers, eligibility, complaints outcomes or regulated advice.

Start with the category disagreement review

The category disagreement review is an editorial method for deciding which disagreements matter before you count categories. You define the intended action, label a small set yourself, compare the AI's suggestions, and inspect cases where the labels would send you towards different actions.

Before using real comments, preserve the original in its authorised location and confirm the purpose for which you can use it. Names are not the only identifying information: order details, locations or an unusual incident may identify someone. Removing a name does not establish that the remaining text is anonymous.

The UK ICO's AI guidance treats minimising personal information as part of system design and use. Check the applicable account, supplier terms and organisational approval before uploading customer text. Requirements vary by country, sector and contract; ask your data protection adviser when the permitted use is unclear. ICO guidance on security and data minimisation in AI.

For a first rehearsal, write fictional comments instead. Do not paste confidential feedback into a second AI service to ask whether the first upload would have been safe.

Define labels around the decision

Write one sentence explaining what you intend to do with the result. “Find which checkout instructions need investigation” is narrower than “understand customer satisfaction”. That sentence determines your labels.

For an illustrative online shop, a compact set might distinguish delivery information, payment process, product description, and returns instructions. Each label needs an inclusion rule and an exclusion example. A late parcel can be a delivery-service problem without being a problem with the delivery instructions. Do not mix those questions merely because the same word appears.

Keep the original comment alongside the proposed label, using a non-identifying row reference. Add a short supporting phrase and a note where context is missing. The phrase must actually occur in the source. An invented explanation is not evidence for a label.

A comment can describe more than one issue. In technical terms this is multilabel classification: the categories need not be mutually exclusive. The distinction is described in the scikit-learn documentation on multilabel classification. You do not need that software or a coding project to apply the idea in a spreadsheet.

Avoid creating dozens of labels for a few dozen comments. If two labels imply the same next action and reviewers cannot distinguish them consistently, combine them. If a rare comment describes a serious problem, preserve it explicitly even when its category count is one.

Compare interpretations, not just agreement percentages

Read a varied selection before asking AI to label the remainder. Include short praise, mixed feedback, an unclear complaint and a comment containing several issues. Write your labels and reasons without seeing the AI answer first.

Then ask the tool to use only your definitions, retain row references and leave uncertain cases unresolved. A possible instruction is: “Return the allowed labels, the exact supporting phrase and any missing context. Multiple labels are permitted. Do not infer a refund request from dissatisfaction alone.” This is an illustrative instruction, not a tested guarantee.

Inspect disagreements in a simple table:

DisagreementWhat to inspectUseful response
Two plausible labelsWhether the comment actually covers both issuesPermit both or clarify the boundary
AI adds an unstated motiveThe exact words that supposedly support itRemove the inference from the label
Your manual label missed an issueWhether the source supports the additional categoryCorrect your own judgement
Neither label fitsWhether this is a new issue or missing contextKeep unresolved or revise the definitions

The manual answer is not automatically correct. If two capable people disagree because the definition is vague, changing the AI tool will not fix the underlying problem.

For a small batch, my default is to inspect every final row, particularly before reporting counts to someone who will act on them. That is an editorial recommendation, not a statistical sampling rule. With a large recurring programme, a qualified analyst may design a more efficient evaluation and sampling approach.

Calculate whether the assistance pays back

Imagine a shop with 65 fictional comments. All timings and outcomes below are illustrative assumptions, not observations about any product.

Manual categorisation takes one minute per comment, plus 12 minutes to define the labels and eight minutes to reconcile the totals:

65 × 1 + 12 + 8 = 85 minutes.

An AI-assisted approach takes 20 minutes to prepare definitions and examples, six minutes to run and arrange the output, and 30 seconds to check each row. Nine ambiguous comments need a further two minutes each. Reconciliation still takes eight minutes:

20 + 6 + (65 × 0.5) + (9 × 2) + 8 = 84.5 minutes.

The difference is half a minute. That does not justify a new subscription or a more complicated data-handling process. It also leaves almost no margin for correcting a missed issue.

Suppose a later batch with the same definitions needs only five minutes of setup. With all other assumptions unchanged, it takes 69.5 minutes, releasing 15.5 minutes. If maintaining the definitions after a product change takes 18 minutes, that month's benefit disappears again.

Record your actual timings before deciding. Count active preparation, checking, corrections and maintenance, not only the seconds the model spends producing labels. Released time becomes a cash saving only if expenditure falls. Using that time for additional work is a capacity benefit; any resulting additional income needs separate evidence.

Report what the counts can and cannot mean

Reconcile the number of source comments, labelled comments and unresolved comments. If one comment can receive several labels, category totals may exceed the number of comments. Say so beside the chart.

For example, “18 comments mention delivery information” is different from “18 customers had a delivery problem”. Several comments might come from one person, and the text may concern instructions rather than the actual service. Do not silently change the unit of analysis.

Avoid treating the batch as representative of all customers unless the collection method supports that claim. Feedback from people who chose to respond is useful evidence of experiences, not automatically a prevalence estimate. Preserve contradictory and unusual comments rather than suppressing them to create a clearer headline.

Make the decision in one afternoon

  1. Spend 15 minutes defining the decision, permitted data and initial categories. If you cannot identify a real action the analysis will support, stop before choosing a tool.
  2. Label a varied small selection manually, then rehearse AI assistance with fictional or approved material. Compare reasons for disagreement.
  3. Time one complete batch, including review and ambiguous cases. Check counts against the source rows.
  4. Choose manual coding unless the assisted approach produces a useful, repeatable advantage. If you retain AI, set a review point for the next changed product, feedback channel or category definition.

Stop automated classification if serious comments are being hidden, the upload is not authorised, or reviewers cannot explain the labels. Finish the small batch manually and repair the method before repeating it.

Frequently asked questions

Can AI create the categories for me?

It can suggest candidates, but you should define the categories used in the final analysis. Suggestions often reflect broad themes rather than the operational decision you need to make. Read the source comments, test whether each proposed label has a clear boundary, and combine labels that lead to the same action. Keep an unresolved option for material that does not fit. If the batch is exploratory research rather than routine reporting, category development may itself be the main analytical work, and outsourcing that judgement can remove the insight you wanted.

Is sentiment analysis enough for a monthly report?

Usually not if the report is supposed to identify something to improve. A customer can praise the product and criticise the payment process in the same sentence. A single positive label loses the complaint, while a negative label loses the praise. Record the topic and supporting text first, then include sentiment only when it adds a distinct decision-relevant dimension. For a consistent long-running measurement programme, a carefully defined sentiment measure may be useful. It should not replace reading serious, mixed or ambiguous feedback before deciding what to do.

Should I delete duplicate comments before counting them?

First determine what duplicated means and what your count represents. An accidental duplicate export row should not become a second response. Two separate contacts from the same customer may be relevant evidence of repeated friction, even though they are not two different customers. Preserve the original records, document your counting rule and use a working copy for exclusions. If you cannot reliably connect comments to individuals without collecting more personal information, report comment counts instead. Do not claim a count of unique customers that the available evidence cannot support.

What if the comments are in several languages?

Treat language as a review requirement, not a detail the model can silently handle. Keep the original text and use a reviewer who understands the language and context for comments that affect important conclusions. A translation can change politeness, idiom or the apparent severity of a complaint. Do not assume a tool's performance on one language establishes its performance on another. If you lack competent review, report that limitation and avoid confident language-specific comparisons. A smaller, properly understood set is preferable to a complete-looking but unreliable classification.

Can I use the model's confidence score to skip review?

Do not treat a conversational confidence percentage as a calibrated probability that the label is correct. Ask instead for the source phrase and the reason the allowed category applies, then inspect the relationship. A specialised system may provide documented scores, but their interpretation still depends on its evaluation and your workload. For a small batch, reading the rows is usually manageable and gives you more useful evidence than a high-looking score. If review is impossible, narrow the proposed use rather than assuming a score supplies the missing judgement.

What should I do when the categories change next month?

Version the definitions and record when the change occurred. If you compare monthly counts without noticing that a category expanded, a procedural change can look like a change in customer experience. Recode an earlier comparison set using the new definitions if that is necessary and permitted, or clearly mark the break in the series. Keep the previous definitions so another person can understand the earlier report. When changes are frequent, a short narrative explaining the actual issues may be more honest and useful than a trend chart built on unstable labels.

Sources and verification

  • scikit-learn: multilabel classification, checked 11 September 2026 for non-exclusive category terminology; no library installation is required by this article.
  • ICO: security and data minimisation in AI, checked 11 September 2026 for UK data-handling context.
  • The parent guide was read in the supplied website source because its public route could not be retrieved. The comparison is an editorial method with illustrative figures, not a hands-on product evaluation.
Twokq Tech

This article is practical guidance. Apply it in proportion to your tools, evidence, risks, and responsibilities.