Accuracy Passed, but Users Are Leaving
When a company rolls out an internal AI assistant, most of the validation effort goes into answer accuracy. Once it is live, though, the first complaint is often simply "it's slow." Long-cited usability research puts the thresholds at roughly 1 second before a user's train of thought starts to break and 10 seconds before their attention moves elsewhere. An accurate answer does not help adoption if the employee has already gone back to a search box or asked a colleague.
Latency needs to be split into two measures.
The two have different causes, so they need different fixes.
Prompt Caching: Don't Recompute the Repeated Prefix
Major LLM APIs can reuse the beginning of a request when it is identical to a previous one, instead of processing it again. Input tokens read from cache are billed at anywhere from about half the list price down to 10% or less, depending on provider and model, and TTFT drops as well.
Consider a chatbot with an 8,000-token system prompt and policy document, plus a 200-token user question. If cached reads cost 10% of the list price, the input cost of a cache-hit request falls from 8,200 tokens' worth to 1,000 tokens' worth — a reduction of about 88%. Some providers add a surcharge when the cache is first written, so the real saving depends on your hit rate.
The rule is stable content first, variable content last: tool definitions → system prompt → reference documents → conversation history → the current question. These are the common mistakes that destroy cache hits:
Send Non-Urgent Work Through Batch
Batch APIs from the major providers charge about half the list price in exchange for completion within 24 hours. The test is simple: is someone sitting in front of a screen waiting for this result?
Streaming and UI Design Change Perceived Speed
Streaming does not shorten total completion time. What it does is turn waiting time into reading time.
Right-Size the Model and the Output for Each Task
Output tokens dominate total completion time. A model generating 50 tokens per second takes 20 seconds for a 1,000-token answer and 6 seconds for a 300-token one.
No Measurement, No Improvement: Latency and Cost Metrics
Averages hide the problem. A p50 of 2 seconds with a p95 of 20 seconds means one request in twenty takes 20 seconds or more — and that is the one users remember.
How POLYGLOTSOFT Helps Optimize AI Services
A slow AI service usually needs its structure reviewed before its model is replaced. POLYGLOTSOFT analyzes the request logs of your running internal AI service to diagnose where latency and cost come from, then supports the improvements: prompt restructuring, real-time and batch path separation, streaming UI, and observability dashboards. If speed or cost is holding back your internal AI, contact POLYGLOTSOFT.
