> ## Documentation Index
> Fetch the complete documentation index at: https://docs.orbit.devotel.io/llms.txt
> Use this file to discover all available pages before exploring further.

# Semantic response cache: serving near-duplicate first questions without a model round-trip

> The opt-in per-agent semantic response cache — how a near-duplicate first-turn question can be answered from a cached reply, the eligibility rules that keep context-dependent turns out, the ring's per-agent bounds, and when enabling it is worth it.

# Semantic response cache

The semantic response cache is an opt-in, per-(organization, agent) store of recent first-turn question-and-answer pairs. When a new conversation opens with a question that is semantically near a question the agent already answered well, the runtime serves that stored answer instead of running a full generation — a single embedding lookup replaces the model call, the tool loop, and any knowledge retrieval behind it.

This page explains the model so you can decide whether to enable it and predict what it will do. The per-turn behavior of an agent run as a whole is covered in [Agent run lifecycle](/concepts/agent-run-lifecycle); this page stays on the cache.

## What it is

For each agent you enable it on, the cache keeps a small ring of recent entries in the platform's shared cache store. Each entry pairs a query's embedding — the vector representation of the customer's question — with the answer the agent produced for it.

On an eligible opening message, the runtime embeds the incoming question once, compares that vector against every entry in the ring by cosine similarity, and serves the stored answer from the highest-scoring entry that clears the similarity threshold. A hit answers the turn without invoking the model; a miss runs the normal agent path and, once it completes, stores the fresh question-and-answer pair at the head of the ring for the next near-duplicate question.

Only the opening message of a conversation is ever eligible — see [Eligibility and exclusions](#eligibility-and-exclusions).

## Why the existing caches did not cover this

Two adjacent caches already exist, and neither answers a reworded question:

* **The embedding cache** avoids re-vectorizing text you have embedded before. It matches identical text only — a reworded question is embedded again. It saves the embedding call, never the generation.
* **The exact-match answer cache** keys responses on a hash of the message text. Change a single word and it misses. It saves the generation only when the question is character-for-character identical.

Real FAQ traffic lives between those two: customers ask the same thing in different words ("How do I reset my password?", "I forgot my password", "Password reset steps"). The semantic cache closes the gap by comparing meaning, not text — that is what lets a stored answer serve a phrasing it has never seen.

## Eligibility and exclusions

The eligibility rules are deliberately narrow so a cached single-shot answer can never be served in place of a context-dependent reply:

| Rule | What it does |
| - | - |
| Opt-in per organization | Nothing is cached until you set `settings.agents.semantic_response_cache` on your organization (see [Enabling the cache](#enabling-the-cache)). Off by default. |
| First turn only | Only a conversation with no message history is eligible. Once any history exists, the turn runs the normal agent path, so a cached standalone answer is never substituted for a follow-up that depends on what was said earlier. |
| Per-contact-memory agents excluded | Agents configured to remember the individual contact are excluded — their answers are personalized by who is asking, so one caller's answer must not serve another. |
| Cross-channel agents excluded | Agents enabled for cross-channel continuity are excluded — their answers are bound to a transcript carried over from another channel. |
| Text-only input | Opening messages with attachments are excluded. |
| Escalation always wins | An opening message that matches an escalation trigger you configured, or that reaches for a human ("speak to an agent", "a person, please"), is never cache-eligible — the hand-off instruction beats the cache and the run escalates as it otherwise would. |

There is a second guard on the store side, not just the lookup side: an answer whose original turn **called tools** or **escalated to a human** is never written to the ring. Tool-using answers are dynamic (the data behind a tool call moves), and an escalated opening is not something you want replayed as a canned reply. Only clean, self-contained answers are stored.

## Ring mechanics

Each enabled agent gets one ring, and the ring has hard bounds:

* **Keyed per organization and agent.** The ring's key is built from your organization identifier plus the agent identifier, so one organization's cached answers can never be read by another, and two agents in the same organization keep separate rings.
* **Newest-first, capped.** The ring holds at most 50 entries. When a store would exceed that, the oldest entry is dropped — the ring always holds the most recent 50 near-first-turn pairs for that agent.
* **TTL clamped.** The entry ring carries a time-to-live, clamped to between 60 seconds and 86,400 seconds (one day). The default is 3,600 seconds (one hour). The whole ring expires together; the TTL re-arms on every store. A stale answer cannot outlive the clamp.
* **Threshold clamped.** The cosine-similarity threshold is clamped to between 0.8 and 0.9999. The default is 0.95 — deliberately tight, so a "hit" is a near-certain paraphrase. A value you set below 0.8 is raised to the floor; below the floor a hit would be too loose to be a safe paraphrase.

## Failure posture

Every cache operation is best-effort. If the cache store is unavailable, slow, or returns malformed data, the turn degrades to a plain cache miss and the normal agent path runs — the same behavior a customer sees with the feature switched off. A failed store at worst drops one ring entry, which the next miss re-generates; a lost update under concurrent opening messages costs one extra generation, never a wrong answer.

## Enabling the cache

Set the toggle under your organization's agent settings:

```json theme={null}
{
  "settings": {
    "agents": {
      "semantic_response_cache": {
        "enabled": true,
        "threshold": 0.95,
        "ttlSeconds": 3600
      }
    }
  }
}
```

`enabled` is the only required field. `threshold` and `ttlSeconds` are optional; omitted or out-of-range values fall back to the defaults (0.95 and 3,600 seconds) clamped into the allowed band. You can change the values at any time — the resolution happens per turn.

A cache hit changes what a turn costs, in time and in money:

| | Without the cache (miss) | With the cache (hit) |
| - | - | - |
| Model generation | Full generation runs | Skipped |
| Tool loop | Runs if the query turns out to need tools | Skipped |
| Knowledge retrieval (RAG) | Fan-out over your knowledge base | Skipped |
| Token usage | Billed tokens for the generation | Zero billed tokens (a hit costs nothing) |
| Reported usage | Actual token count | A consistent estimated count, so a cache-served reply never reads as "0 tokens" on usage surfaces |

Hits are deliberately conservative: a first turn that would have triggered escalation is never served from the cache, and an answer that used tools was never stored. The cache trades a saved generation only when the stored answer was already a complete, self-contained reply.

## When to enable it

Enable the cache when your agent's inbound traffic is FAQ-shaped — where most conversations open with one of a handful of questions phrased a thousand ways. Support-deflection agents, onboarding assistants, and status/FAQ agents are the canonical fit: near-duplicate first questions dominate, and each hit avoids a full generation.

Leave it off when first questions are inherently contextual or personal — an agent that opens by recalling the contact's history, or one whose opening questions routinely fire tools or escalate, will find almost nothing eligible, and the narrow exclusions above already enforce that.

## Related reading

* [Agent run lifecycle](/concepts/agent-run-lifecycle) — the states every run moves through; a cache hit short-circuits the run before the model executes.
* [AI agent architecture](/concepts/ai-agent-architecture) — flows, retrieval, tools, and where this cache sits in front of the stack.
* [Cost controls](/agents/cost-controls) — the ceilings that bound what a cache-missed run may spend.
* [Outcome-based AI-agent billing](/concepts/outcome-pricing-model) — how resolved conversations are billed; a cache hit emits no token charge.
