Semantic response cache
The semantic response cache is an opt-in, per-(organization, agent) store of recent first-turn question-and-answer pairs. When a new conversation opens with a question that is semantically near a question the agent already answered well, the runtime serves that stored answer instead of running a full generation — a single embedding lookup replaces the model call, the tool loop, and any knowledge retrieval behind it. This page explains the model so you can decide whether to enable it and predict what it will do. The per-turn behavior of an agent run as a whole is covered in Agent run lifecycle; this page stays on the cache.What it is
For each agent you enable it on, the cache keeps a small ring of recent entries in the platform’s shared cache store. Each entry pairs a query’s embedding — the vector representation of the customer’s question — with the answer the agent produced for it. On an eligible opening message, the runtime embeds the incoming question once, compares that vector against every entry in the ring by cosine similarity, and serves the stored answer from the highest-scoring entry that clears the similarity threshold. A hit answers the turn without invoking the model; a miss runs the normal agent path and, once it completes, stores the fresh question-and-answer pair at the head of the ring for the next near-duplicate question. Only the opening message of a conversation is ever eligible — see Eligibility and exclusions.Why the existing caches did not cover this
Two adjacent caches already exist, and neither answers a reworded question:- The embedding cache avoids re-vectorizing text you have embedded before. It matches identical text only — a reworded question is embedded again. It saves the embedding call, never the generation.
- The exact-match answer cache keys responses on a hash of the message text. Change a single word and it misses. It saves the generation only when the question is character-for-character identical.
Eligibility and exclusions
The eligibility rules are deliberately narrow so a cached single-shot answer can never be served in place of a context-dependent reply:
There is a second guard on the store side, not just the lookup side: an answer whose original turn called tools or escalated to a human is never written to the ring. Tool-using answers are dynamic (the data behind a tool call moves), and an escalated opening is not something you want replayed as a canned reply. Only clean, self-contained answers are stored.
Ring mechanics
Each enabled agent gets one ring, and the ring has hard bounds:- Keyed per organization and agent. The ring’s key is built from your organization identifier plus the agent identifier, so one organization’s cached answers can never be read by another, and two agents in the same organization keep separate rings.
- Newest-first, capped. The ring holds at most 50 entries. When a store would exceed that, the oldest entry is dropped — the ring always holds the most recent 50 near-first-turn pairs for that agent.
- TTL clamped. The entry ring carries a time-to-live, clamped to between 60 seconds and 86,400 seconds (one day). The default is 3,600 seconds (one hour). The whole ring expires together; the TTL re-arms on every store. A stale answer cannot outlive the clamp.
- Threshold clamped. The cosine-similarity threshold is clamped to between 0.8 and 0.9999. The default is 0.95 — deliberately tight, so a “hit” is a near-certain paraphrase. A value you set below 0.8 is raised to the floor; below the floor a hit would be too loose to be a safe paraphrase.
Failure posture
Every cache operation is best-effort. If the cache store is unavailable, slow, or returns malformed data, the turn degrades to a plain cache miss and the normal agent path runs — the same behavior a customer sees with the feature switched off. A failed store at worst drops one ring entry, which the next miss re-generates; a lost update under concurrent opening messages costs one extra generation, never a wrong answer.Enabling the cache
Set the toggle under your organization’s agent settings:enabled is the only required field. threshold and ttlSeconds are optional; omitted or out-of-range values fall back to the defaults (0.95 and 3,600 seconds) clamped into the allowed band. You can change the values at any time — the resolution happens per turn.
A cache hit changes what a turn costs, in time and in money:
Hits are deliberately conservative: a first turn that would have triggered escalation is never served from the cache, and an answer that used tools was never stored. The cache trades a saved generation only when the stored answer was already a complete, self-contained reply.
When to enable it
Enable the cache when your agent’s inbound traffic is FAQ-shaped — where most conversations open with one of a handful of questions phrased a thousand ways. Support-deflection agents, onboarding assistants, and status/FAQ agents are the canonical fit: near-duplicate first questions dominate, and each hit avoids a full generation. Leave it off when first questions are inherently contextual or personal — an agent that opens by recalling the contact’s history, or one whose opening questions routinely fire tools or escalate, will find almost nothing eligible, and the narrow exclusions above already enforce that.Related reading
- Agent run lifecycle — the states every run moves through; a cache hit short-circuits the run before the model executes.
- AI agent architecture — flows, retrieval, tools, and where this cache sits in front of the stack.
- Cost controls — the ceilings that bound what a cache-missed run may spend.
- Outcome-based AI-agent billing — how resolved conversations are billed; a cache hit emits no token charge.