Back to Blog
AI

Nobody Reads Your Tens of Thousands of Support Tickets: Building an LLM-Based VOC Classification and Insight Pipeline

A practical guide to building an LLM pipeline that classifies and summarizes scattered support tickets and reviews, then ties them to improvement work. It covers taxonomy design, accuracy validation, privacy, and cost trade-offs.

POLYGLOTSOFT Tech Team2026-10-057 min read1
VOC AnalysisVoice of CustomerLLM ClassificationText AnalyticsCustomer Experience

The Voice of the Customer That Only Piles Up

Support transcripts live in the contact center system, reviews in the store admin, open-ended survey answers in spreadsheets, and sales notes in the CRM. Each channel has its own format and its own owner, so even a simple question such as "what frustrated customers most this month?" is hard to answer.

The usual workarounds are keyword counts and manual sampling. Keyword counts tell you how often "delivery" appeared, but not whether the problem was a delay, damage, or a wrong item. Complaints with no keyword at all, like "it's broken again," are missed entirely. Sampling 100 tickets a month means that an issue behind 0.5% of all tickets shows up, on average, in half a ticket.

Start With the Taxonomy

Before wiring up an LLM, decide what you are sorting into. Labels that mix several attributes, such as "payment error complaint," multiply into hundreds of categories, so the key is to keep the axes separate.

  • Product and feature: which product, and which part of it
  • Issue type: defect, how-to question, feature request, policy complaint, and so on
  • Sentiment: positive, neutral, negative
  • Urgency: act now, routine, for reference
  • Building the taxonomy takes three steps. First, feed 500 to 1,000 recent tickets to an LLM and have it propose candidate categories. Next, the people who handle these tickets merge and split the candidates until each axis has 10 to 30 categories. Finally, write down a definition for each category, with examples of what belongs and what does not. That document becomes the basis of the classification prompt.

    Assembling the LLM Classification Pipeline

    The pipeline has four stages.

  • PII masking: replace names, phone numbers, addresses, and account numbers before any model call.
  • Classification and summary: present the agreed taxonomy and get back a label per axis plus a one-line summary in a structured format.
  • Evidence extraction: return the original sentence that justified each label, so staff can verify the result.
  • Storage: save labels, summary, and evidence together with the model, prompt, and taxonomy versions.
  • Tickets that fit no category should go to "other" and never be forced into a poor match. When "other" exceeds 5% of volume, or similar content keeps recurring in it, a new issue type is emerging. Revise the taxonomy once a quarter, and off-cycle when something big changes, such as a product launch.

    Can You Trust the Labels? Measuring Accuracy

    A dashboard without validation is just plausible-looking numbers. Have two staff members independently label the same 300 to 500 tickets to build an evaluation set, then measure agreement with the model on each axis. Check agreement between the two people as well. If they agree only 85% of the time on an axis, demanding 95% from the model is meaningless, and the result tells you the category definitions are ambiguous.

    Changing the model or the prompt can give the same ticket a different label and break your trend lines. Before any change, run a regression test on the evaluation set, then classify the most recent month with both the old and new versions and compare the counts per category. If the gap is large, reclassify historical data with the new version. If it is small, mark the change date on the charts.

    Turning Analysis Into Action

    Results have to arrive in a form each team can use. Product needs feature requests ranked by feature, quality needs trends by defect type, and sales needs complaint history by account. A spike alert is also valuable: notify the team in chat when a category runs at more than twice its average over the previous four weeks.

    Integration matters even more. Linking improvement tasks in the issue tracker to the related ticket counts and evidence sentences moves prioritization from impressions to data, and lets you confirm after release whether that category actually declined.

    Watch Points: Personal Data and Cost

    Support data contains personal information. If you use an external LLM API, check whether it counts as outsourced processing or a cross-border transfer under privacy law, and confirm in the contract that your inputs are not used for model training. Set retention rules up front as well. In Korea, e-commerce businesses must keep records of consumer complaints and dispute handling for three years, so a safe design keeps the originals for the statutory period and stores the analysis data separately in masked form.

    Cost becomes an easy decision once you run the numbers. Assuming ₩3 per ticket, processing all 30,000 tickets a month costs ₩90,000. A 10% sample saves ₩81,000, but an issue with a 0.5% incidence drops from 150 tickets to 15, too few to read a weekly trend. If spike detection is the goal, full processing is usually the sensible choice, with costs trimmed by routing short tickets to a lighter model.

    How POLYGLOTSOFT Builds VOC Analysis Systems

    POLYGLOTSOFT builds the analysis pipeline on top of your existing support system and CRM, with no need to replace either. We cover the whole flow: collecting and masking data from each channel, taxonomy design workshops, accuracy validation against an evaluation set, team dashboards, and issue tracker integration. If you want to turn a backlog of unread customer feedback into a list of improvements, get in touch with POLYGLOTSOFT.

    Need Technical Consultation?

    Our expert consultants in smart factory, AI, and logistics automation will analyze your requirements.

    Request Free Consultation