Back to Blog
AI

When Users Say Your AI Is Slow: Prompt Caching, Batching, and Streaming to Cut Both Latency and Unit Cost

Complaints that an internal AI is slow are usually solved faster by fixing the architecture than by swapping the model. This post covers how prompt caching, batch processing, streaming, and output-length control cut both latency and cost per request.

POLYGLOTSOFT Tech Team2026-10-057 min read0
Prompt CachingLLM LatencyBatch ProcessingStreamingAI Service Design

Accuracy Passed, but Users Are Leaving

When a company rolls out an internal AI assistant, most of the validation effort goes into answer accuracy. Once it is live, though, the first complaint is often simply "it's slow." Long-cited usability research puts the thresholds at roughly 1 second before a user's train of thought starts to break and 10 seconds before their attention moves elsewhere. An accurate answer does not help adoption if the employee has already gone back to a search box or asked a colleague.

Latency needs to be split into two measures.

  • Time to first token (TTFT): how long until the first character appears. It is driven by input length and by pre-processing steps such as retrieval and tool calls.
  • Total completion time: how long until the last token arrives. It scales almost linearly with output length.
  • The two have different causes, so they need different fixes.

    Prompt Caching: Don't Recompute the Repeated Prefix

    Major LLM APIs can reuse the beginning of a request when it is identical to a previous one, instead of processing it again. Input tokens read from cache are billed at anywhere from about half the list price down to 10% or less, depending on provider and model, and TTFT drops as well.

    Consider a chatbot with an 8,000-token system prompt and policy document, plus a 200-token user question. If cached reads cost 10% of the list price, the input cost of a cache-hit request falls from 8,200 tokens' worth to 1,000 tokens' worth — a reduction of about 88%. Some providers add a surcharge when the cache is first written, so the real saving depends on your hit rate.

    The rule is stable content first, variable content last: tool definitions → system prompt → reference documents → conversation history → the current question. These are the common mistakes that destroy cache hits:

  • Putting the current timestamp or a request ID at the top of the system prompt
  • Letting retrieved documents or tool definitions appear in a different order on each request
  • Placing personalized details such as user name or department ahead of the fixed section
  • Forgetting that cache lifetimes are typically measured in minutes, so a low-traffic service expires its cache between requests
  • Send Non-Urgent Work Through Batch

    Batch APIs from the major providers charge about half the list price in exchange for completion within 24 hours. The test is simple: is someone sitting in front of a screen waiting for this result?

  • Batch candidates: overnight meeting summaries, support ticket classification, document tagging, re-running evaluation sets
  • Separate the real-time path and the batch path with queues so bulk jobs do not eat into the per-minute rate limit that live requests depend on.
  • Give each batch job an idempotency key and build a way to reprocess only the failed items.
  • Streaming and UI Design Change Perceived Speed

    Streaming does not shorten total completion time. What it does is turn waiting time into reading time.

  • Progress indicators: show the stages, such as "Searching documents → 3 found → Writing answer."
  • Intermediate results: display the retrieved sources before the answer itself.
  • Asynchronous handling: for jobs longer than a minute, such as report generation, return a job ID immediately and notify the user by chat or email when it finishes.
  • Right-Size the Model and the Output for Each Task

    Output tokens dominate total completion time. A model generating 50 tokens per second takes 20 seconds for a 1,000-token answer and 6 seconds for a 300-token one.

  • Set a maximum output length per task and specify the format, such as "summarize in three lines."
  • Move detailed explanations behind a "Show more" action so only those who need them pay the wait.
  • Steps such as classification, routing, and information extraction often run well on a smaller model. Build an evaluation set from 100–200 real requests, compare the small model against the large one, and switch the steps that clear your pass threshold.
  • No Measurement, No Improvement: Latency and Cost Metrics

    Averages hide the problem. A p50 of 2 seconds with a p95 of 20 seconds means one request in twenty takes 20 seconds or more — and that is the one users remember.

  • Latency distribution: record p50 and p95 separately for TTFT and total completion time.
  • Cache hit rate: the share of input tokens read from cache. A sharp drop right after a deployment means the prompt prefix changed.
  • Cost per request: aggregate by feature and by department to see where the money goes.
  • Input and output token counts: the baseline data for tracing both latency and cost.
  • How POLYGLOTSOFT Helps Optimize AI Services

    A slow AI service usually needs its structure reviewed before its model is replaced. POLYGLOTSOFT analyzes the request logs of your running internal AI service to diagnose where latency and cost come from, then supports the improvements: prompt restructuring, real-time and batch path separation, streaming UI, and observability dashboards. If speed or cost is holding back your internal AI, contact POLYGLOTSOFT.

    Need Technical Consultation?

    Our expert consultants in smart factory, AI, and logistics automation will analyze your requirements.

    Request Free Consultation