Skip to main content

Inbox AI co-pilot

When a new inbound customer message lands in an open conversation, Devotel Orbit’s AI co-pilot drafts a suggested reply and surfaces it as a card above the composer. Your operator sees the draft, its confidence score, and its reasoning — they can Accept, Edit, or Discard it. Nothing ships to the customer until a person approves or edits the draft and sends it. This page covers the full picture: what the drafting pipeline reads, how the per-workspace kill switch and confidence threshold work, latency and cost, per-agent behavior including manual drawer access, the edge cases the pipeline handles, and how the co-pilot differs from AI deflection (the auto-send sibling on the same settings page).

What the drafting pipeline reads

Every time a new inbound message arrives — and the operator is not already typing in the composer — the co-pilot assembles a draft from four sources:
  1. Prior-thread window. The last 10 conversation turns (inbound and agent messages, newest-first trimmed to 10) provide conversational context so the draft stays on-topic.
  2. Knowledge base scope. A semantic search over your workspace’s knowledge bases returns the top 3 most relevant chunks, reranked for quality. The model is instructed to ground its reply in these citations and must return at least one cite_sources entry when KB chunks were consulted — unsourced drafts fall back with a missing_citations reason so hallucinated facts never ship without provenance.
  3. Long-term memory (when available). If the conversation is linked to a contact, the pipeline pulls up to 3 long-term memory items for that contact — past tickets, preferences, and interaction history stored in the agent memory plane.
  4. Tenant tone guideline. The workspace’s org.settings.ai.tone config shapes the voice and style of every draft.
The model is Anthropic Claude Haiku 4.5, called with temperature: 0.3 and a 600-token output budget. The prompt is JSON-only — every draft returns a structured { draft, suggested_action, confidence, reasoning, cite_sources } envelope validated against a Zod schema before it reaches the operator. PII handling. Before any history is sent to the model, customer names, phone numbers, email addresses, and credit-card numbers are redacted from the transcript. Redaction fails open — if the PII scrubber cannot run, the raw transcript is sent and Sentry captures the failure so the team can investigate. Anti-injection. The prompt wraps every message in XML tags (<customer_message>, <agent_message>, <kb_article>) and the system instruction tells the model that only content inside <customer_message> is customer-controlled. XML-significant characters in message bodies are escaped so a malicious inbound message cannot break out of the tag and inject instructions.

Channels and languages covered

The co-pilot drafts replies on every text-based inbound channel the inbox supports: SMS, MMS, WhatsApp, RCS, Facebook Messenger, Instagram, Telegram, Viber, LINE, KakaoTalk, Zalo, WeChat, email, and the native chat widget. Voice calls are not co-pilot-draftable — the voice AI-agent plane has its own real-time pipeline. The model operates in the default language of the workspace. There is no per-conversation language detection or translation step in the drafting pipeline — the draft is generated in the same language the model is configured to use.

Latency: does the agent wait on a draft?

No. The drafting pipeline runs asynchronously — a new inbound message triggers the draft fetch, but the operator’s composer is never blocked waiting for it. The frontend waits 5 seconds after the operator stops typing before fetching a fresh draft (the typing grace period). If a draft is still generating when the operator starts composing their own reply, the card is suppressed and the in-flight request is aborted at the server. The server caches drafts for 60 seconds in Redis, keyed on (conversationId, last-message-id:message-count). When the operator navigates across the inbox and revisits a conversation with no new messages, the cached draft is returned instantly with no LLM call.

Per-agent behavior: kill switch and manual drawer access

Workspace kill switch

The enabled field on the co-pilot config is a workspace-wide kill switch. When an owner or admin sets it to false:
  • The AI-draft endpoint short-circuits with an empty fallback envelope before any transcript read or LLM call — no tokens are burned.
  • The composer card disappears for every operator in the workspace.
  • The change takes effect immediately (the config cache is invalidated on every PATCH so there is no TTL delay).

Manual co-pilot drawer

Agents can still open the co-pilot drawer manually even when the kill switch is off. The drawer is a separate surface — the operator types a prompt, selects a tone variant (friendly, formal, empathetic, or concise), and gets back three stylistic rewrites of their draft. This is the reply-coach surface, distinct from the auto-draft pipeline. It honors the kill switch: when enabled is false, the drawer’s LLM calls are also gated so no tokens are spent. The drawer itself remains accessible — an operator can open it, see the disabled state, and close it again. The manual drawer is reachable from the “AI” button in the composer toolbar and is documented separately under the reply-coach feature.

auto_approve_eligible visual in the composer

When a draft qualifies for one-click approve-and-send, the composer card shows a green “Auto-approve eligible” badge next to the confidence score. When it does not qualify, the badge reads “Human review” in a neutral tone. Both badges include a tooltip: “Auto-approve eligibility is advisory — confidence is the model’s self-report and is never trusted alone.” The badge is purely informational. It marks a draft as meeting the workspace’s confidence bar — it does not bypass any outbound gate, and the human operator still clicks Accept or Edit and then Send.

Design rationale for require_human_review

Every draft envelope carries require_human_review: true. This field is always true — it is not configurable and there is no path in the co-pilot pipeline where it can be false. The rationale: confidence is the model’s self-reported 0–1 score, and a self-reported number from a model is not a safety guarantee. No matter how high the confidence or how low the threshold, a human always reviews the draft before anything ships to the customer. This is the “human-in-the-loop” contract that distinguishes the co-pilot (assistive) from AI deflection (automated). The auto_approve_eligible flag is the routing decision — it says “this draft meets the bar for one-click approval” — but require_human_review is the safety assertion that says “and a person is still in the loop.”

Telemetry: per-day auto-approve rate visibility

There is no dedicated “auto-approve rate” metric because auto-approve eligibility is advisory and never dispatched. The telemetry that tracks operator behavior on drafts lives in the AI-draft feedback system:
  • POST /ai-drafts/:conversationId/feedback — the frontend fires this whenever an operator accepts, edits, discards, or regenerates a draft. Each feedback row records the action (sent_verbatim, edited, discarded, regenerated), the draft’s confidence score, an optional structural diff of edits, and a suggestion hash for deduplication.
  • GET /ai-drafts/feedback/aggregate?by=agent|template — rollup by operator or by template, returning counts of each action and derived rates: acceptance rate = (sent_verbatim + edited) / total, edit rate, and discard rate.
The feedback aggregate is the co-pilot’s quality dashboard — it surfaces which templates produce the highest acceptance rate, which operators consistently edit or discard drafts, and whether a model change improved or degraded the suggestion mix. It is consumed by the per-agent quality dashboard and the per-template macro-ranking surface. There is also a separate AI-draft outcome endpoint (POST /ai-drafts/:conversationId/outcome) that captures labeled data for model re-tuning — the final disposition, resolution time, and customer satisfaction score for conversations where a co-pilot draft was involved.

Cost: each draft is an LLM call

Every fresh draft generation burns one LLM call. At Haiku 4.5 pricing (~1/Minputtokens, 1/M input tokens, ~5/M output tokens), a single draft costs a fraction of a cent. The pipeline has two mechanisms to keep cost predictable: Rate limiting. The POST /conversations/:conversationId/ai-draft endpoint is gated at 60 requests per minute per tenant. This is a burst ceiling — normal inbox usage with a handful of operators drafting replies stays well under it. Deduplication by fingerprint cache. The 60-second Redis cache, keyed on (conversationId, last-message-id:message-count), means the draft is regenerated at most once per new message. Browsing back to a conversation without new messages hits the cache and costs nothing. The cache is intentionally bypassed when the operator is actively typing (a draft string in the request body) — caching a stale draft while the operator is composing their own reply would be wasteful. Kill switch as cost control. Setting enabled: false prevents ALL LLM calls for co-pilot drafts across the workspace — no tokens are burned until an owner re-enables it. No per-draft spending cap. There is no budget or spending cap specific to co-pilot drafts. The sibling AI-deflection feature has its own monthly resolution cap and hard-pause switch (configured on the same Inbox → Settings → AI Deflection page), but the co-pilot does not share that budget. If you want to cap co-pilot spend, use the kill switch or reduce the number of operators triggering draft generation.

Edge cases

Draft on a closed conversation. The ai-draft endpoint checks conversation existence (returns 404 for missing) but does not gate on conversation status. A draft can be requested on a closed or resolved conversation. In practice, the frontend suppresses the co-pilot card when the conversation is in a terminal state, so this path is rarely exercised — but the API itself does not reject it. Operator starts typing during draft generation. The frontend suppresses the draft card when the composer has content. The in-flight LLM request is aborted at the server via AbortController bound to the client connection close event, freeing the model connection immediately. Race with an agent already typing. If two operators are viewing the same conversation, one operator’s typing does not affect the other’s draft card. Each operator’s browser independently fetches the draft. The fingerprint cache prevents duplicate LLM calls: the first operator’s fetch populates the cache, and the second operator’s fetch (same conversation, same fingerprint) hits the cache within the 60-second window. Outbound message arrives while draft is showing. A new outbound message updates the fingerprint, so the next fetch bypasses the cache and generates a fresh draft. The frontend also suppresses the draft card when the latest thread message is outbound — this prevents the “accept, then re-draft the same thing” loop. LLM unavailable. When the Anthropic API is unreachable or the API key is missing, the endpoint returns a 200 fallback envelope (fallback: true, empty draft, confidence 0) instead of a 5xx error. The frontend renders a gentle “AI assistance is not available” state rather than an error toast. Malformed model output. If the model returns invalid JSON, the parser tries direct parse, then attempts to extract an embedded {...} block from the raw string, and finally falls back with confidence: 0 and a "parse_error" reasoning. Parse failures are reported to Sentry once per deployment (deduplicated reportOnce warn) to avoid noise. Threshold change mid-TTL. When an owner tightens the auto-approve threshold (e.g., from 0.75 to 0.95), the config cache is invalidated immediately. However, a draft that was cached before the change at the old threshold carries the old auto_approve_eligible flag for up to 60 seconds. This is an accepted trade-off — threshold changes are operator-controlled, rare, and the flag is advisory-only. Rate limit hit. At 60 requests per minute per tenant, the rate limiter returns 429 TOO_MANY_REQUESTS. The frontend shows a retry-after banner. Normal inbox usage stays well under this ceiling.

Comparison with AI deflection (auto-send)

The co-pilot and AI deflection are sibling features on the same Inbox → Settings → AI Deflection page, but they serve different purposes: Both features honor the same kill switch pattern: enabled: false stops all LLM calls. Both are per-workspace, tenant-owned configuration — nothing is platform-wide. The co-pilot is for teams where a human stays in the loop; deflection is for teams comfortable with fully automated KB-grounded replies.

Configure in the dashboard

The Inbox → Settings → AI Deflection page is the self-serve console for both features. You need the owner or admin role to open it.
  1. Open Inbox → Settings → AI Deflection.
  2. Scroll to the Co-pilot mode section.
  3. Set the values you want:
    • AI co-pilot enabled — the org kill switch. Turn it off to stop drafting across the workspace; agents can still open the co-pilot drawer manually.
    • Auto-approve confidence threshold — a 0–1 value with two-decimal steps. Drafts at or above it qualify for one-click approve-and-send; drafts below it always require explicit review.
  4. Save. The page validates the threshold to the 0–1 range before it saves, and the new values apply to live drafts immediately.

Configure via the API

For automation — provisioning a workspace, driving config from a script, or wiring the settings into your admin tooling — use the API. The console above and the two endpoints below read and write the same config; this section is the way to manage it without a browser session.

Read the current settings

Read the effective config with GET /inbox/copilot/config:
The response is the normalized, effective config — the exact values the drafting path uses:

Change the settings

Change the config with PATCH /inbox/copilot/config. Send a partial update with only the fields you want to change:
Both fields are optional on each request, and the threshold must sit between 0 and 1 inclusive — anything else returns 422 VALIDATION_ERROR. Non-owner/admin callers get 403 INSUFFICIENT_PERMISSIONS. The response is the effective config after the update, so follow with a GET or read the response body to confirm the new values. Changes are applied to live drafts immediately.

How the threshold applies to each draft

Each draft response includes two routing fields, derived from your config:
  • require_human_review — always true. A person reviews every draft before anything is sent; confidence is displayed as guidance, never as permission to skip review.
  • auto_approve_eligible — true only when all three hold: the co-pilot is enabled, the draft’s confidence meets your threshold, and the suggested action is send. This is the signal the composer uses to show the “Auto-approve eligible” badge.
Only send-shaped suggestions can ever qualify. When the model suggests escalate or end — closing or handing off the conversation — the draft is never eligible, at any confidence level, because those actions change the conversation’s trajectory rather than just replying to it. Confidence values above 1 or below 0 are clamped to the 0–1 range before comparison, and a workspace whose threshold is 0 effectively qualifies every send-shaped draft. A threshold of 1 requires a perfect score.

Common failures

  • Can’t reach the AI Deflection settings page — the dashboard route requires an owner or admin role. Ask an owner to grant the role, or configure via the API with an owner/admin API key instead.
  • 403 INSUFFICIENT_PERMISSIONS on PATCH — the API key or session belongs to a non-admin operator. Use an owner or admin identity.
  • 422 VALIDATION_ERROR on PATCH — the threshold is outside 0–1, or the body includes a field that does not exist. Send only enabled and auto_approve_confidence_threshold.
  • The AI suggest card is empty for every conversation — check enabled on the config; the kill switch returns a gentle empty state rather than an error.