Skip to main content

Troubleshooting: campaign and conference state-machine transition errors

A state-machine violation is a 409 that refused your write because the object you are updating has moved on. The guard did not lose the race — your read did. Read the current status with a GET first, and you will see the transition the code names, or the one it replaced. All five codes on this page close on the same recovery step; never mutate outside the legal transition, never loop the retry. The five surfaces grouped here: A marketing campaign moves through exactly this progression (mirror of the campaign lifecycle concept page):
approved is the terminal state for this gate; the campaign then follows the wider draft → scheduled → sending → running → completed arc that the concept page owns. The two codes below gate only the first three hops — they are the door on a write that assumes the campaign still sits in the state your UI last observed.

Per-transition legality table

Each row names the only transition the write is trying to perform, the only starting states that accept it, and the code you get when the precondition no longer holds: The legality gate runs as an atomic flip: the write reads the current status, decides it is in the allowed list, and flips — all in one transaction. A 409 from this gate means the row’s status moved between your earlier GET and this write, or your GET never happened and your UI is carrying a stale state map.

CAMPAIGN_NOT_DRAFT — submit-for-approval gate

POST /campaigns/:id/approvals opens the supervisor queue on a campaign your trust tier requires sign-off for. The only legal starting states are draft and scheduled — everything else already holds a live disposition the approval flow cannot reset.
Recovery — in this order, never skip the read:
  1. GET the campaign first. Read status back from GET /campaigns/:id before you react. Whatever it says now is the legal truth.
  2. Match on the status. If it moved to pending_approval since you queued the submit, another writer beat you. If it moved to sending or running, a supervisor approved it while your request was in flight. If it moved to cancelled or failed, the approval path is closed.
  3. Re-route. When the current state is one of the legal starts, retry the same POST once. When it is not, close the loop — the duplicate attempt is the bug to fix, not the gate.

What NOT to try

  • Never PATCH the campaign back to draft to force the gate open — an approved campaign in sending/running rejects the rewrite with its own gate, and an approval that already landed is not retractable via the status field.
  • Never loop the same POST — the 409 is deterministic. Each retry appends a fresh rejection to your audit trail without changing the campaign’s actual state.

CAMPAIGN_NOT_PENDING_APPROVAL — approve / reject race

POST /campaigns/:id/approvals/:approvalId/approve (or /reject) closes the supervisor decision. Both flip the campaign out of pending_approval; both refuse with the same code when the campaign has already moved.
Recovery:
  1. The row already moved. Either a second supervisor decided first (their decided_by_user_id is on the approval row), or the integration you are holding took a path the gate does not accept (an operator edited the campaign and reverted it to draft).
  2. List the pending queue again. GET /campaigns/approvals/pending?campaign_id=:id returns the current decision surface. If the row is gone, the decision already landed.
  3. Rely on idempotence. If your goal was “approve this campaign” and the list shows it is no longer pending, treat the goal as reached — the 409 you just saw is the proof another supervisor closed it first.

What NOT to try

  • Never re-approve against a second approval row for the same campaign — the first decision sticks and subsequent approvals are redundant.
  • Never interpret the 409 as a transient fault; the precondition read is the whole point, and looping widens the supervisor audit log.

CONFERENCE_STATE_UNRESOLVED — the lock-a-room gate

POST /voice/conferences/:id/lock refuses when no leg in the room is currently connected. A lock only applies to a live room: legs that are still pending, dialing, or ringing are named as unresolved, legs in any terminal state (failed, no_answer, busy, rejected, disconnected) are already settled, and an empty room resolves to the same refusal. The gate reads the room’s participant roster — the same one the GET /voice/conferences/:id response and the participant_joined / participant_left webhooks present — and refuses whenever zero participants are in connected.
Recovery:
  1. Poll the roster once. GET /voice/conferences/:id returns the participant list; read it before reacting. The unresolved count in error.details tells you whether legs are still dialing or already settled.
  2. Wait for one participant_joined. The conference.participant_joined webhook fires the moment a leg enters the room; the gate is open from that point on.
  3. Then POST the lock once. The delegate only needs the first connected participant — the rest of the room can still be dialing.
The same gate executes on unlock: when you release the lock, the write proceeds unconditionally; the refusal only ever fires on the locking direction. The conference lifecycle runbook covers the room-level event flow this gate reads, and the supervisor take-over runbook covers the take-over side of the same plane.

What NOT to try

  • Never retry the lock in a tight loop while legs are dialing — poll GET /voice/conferences/:id on a backoff (the roster converges in milliseconds once a leg answers).
  • Never treat an unresolved count of zero as a green light — that is the settled-but-never-connected case the gate closes explicitly.

INVALID_STATE_TRANSITION — the generic 409

Whenever a route guards its write with an atomic status flip — the same “read the precondition, flip, commit” shape every state-machine surface uses — the generic INVALID_STATE_TRANSITION code covers the rest of the corpus. The code name is intentionally broader than the per-surface codes above; read it as “the precondition a specific route declared did not pass.”
Recovery is identical across every surface:
  1. Identify the route. The error.message (and the endpoint your request hit) names the transition it tried — match it back to the “From an error code to a runbook” table on the hub.
  2. GET the current state first. Before any retry, re-read the object.
  3. Re-issue once, with the correct state. The route’s status-whitelist flip rejects stale preconditions deterministically; one retry after a fresh read is all that is ever safe.
INVALID_STATE_TRANSITION is the catch-all to route to when a take-over, queue-flow flip, or lifecycle write you control fails with 409 and your surface is not one of the four named above. The Error Code Reference keeps the per-surface message strings; this page owns the recovery shape.

INVARIANT_VIOLATION — the take-over sentinel

A supervisor action — listen, whisper, barge, unlisten — must never resolve to the retired provider branch. When the routing layer resolves a supervisor leg to a path outside the current provider contract, the take-over refuses with 409 INVARIANT_VIOLATION and a route-level log breadcrumb.
Recovery:
  1. Read the leg stamp. GET /voice/calls/:id returns the call metadata the resolver read; the refuse names the provider branch it resolved to in details.
  2. Re-stamp or terminate. A leg stamped with a pre-cutover provider value or an explicit override env needs a re-dial; terminate the supervisor leg and re-dial against the current softswitch routing contract.
  3. Escalate with the breadcrumb. When the refuse persists past a re-dial, the platform-side config your tenant inherited is the cause — open a ticket with the request id and the verb you attempted.
This code is the one refusal customers on the supervisor plane treat as fail-closed: the take-over never silently degrades to a retired path, so the only valid fix is routing the leg through the current softswitch surface. The supervisor take-over runbook carries the per-verb recovery table for the non-409 failure family.

Decision checklist

Run these in order on any of the five codes:
  1. Re-read the object (GET /campaigns/:id, GET /voice/conferences/:id, GET /voice/calls/:id).
  2. Match the current status against the legality table for the surface you called.
  3. Re-issue the same write at most once, after the read reflects the new precondition.
  4. If the precondition still fails, the object is in a disposition the write cannot advance — treat the current status as the terminal answer, and close the loop from your side.

When to escalate

  • The same INVALID_STATE_TRANSITION recurs on multiple surfaces — the gate contract your tenant relies on is ambiguous; open a ticket with the endpoint and the error.details blob.
  • CAMPAIGN_NOT_PENDING_APPROVAL lands repeatedly because more than one supervisor in your tenant holds the queue — the visibility contract in the approvals list is what you lean on; ask support whether the pending-list shape should include a claim row.
  • CONFERENCE_STATE_UNRESOLVED lands when the webhook stream shows at least one participant_joined before your POST — the roster and the webhook stream disagree; file the conference id and the request id together.
  • INVARIANT_VIOLATION returns a verb you did not configure — the legacy provider stamp is in metadata you inherited; the platform team needs the legacy record to look at, not just the request id.

Cross-references