Troubleshooting: conference lifecycle failures
A conference moves through a fixed lifecycle:POST /api/v1/voice/conferences
creates the room, each dial-out leg produces a join or a failure, legs leave
one at a time, and the room closes once with a terminal end_reason. Every
step emits a webhook event, so a conference that “did nothing” always broke at
a specific step — and the events you did and did not receive tell you which
one.
This page maps each conference symptom to the step that failed. For a single
bad leg with audio trouble after a successful join, see
Troubleshooting: voice call quality;
this page works the room-level state machine, not per-call audio. The
AI-agent participant section below
covers the one join path that isn’t a PSTN leg — an AI agent bridged into the
room.
What the conference lifecycle tells you
Every conference emits the same four event types, in this order:-
conference.created— the create call succeeded and the room exists. Carries the requested participant count, so you can compare “how many legs were asked for” against how many ever joined. -
conference.participant_joined— one leg answered and entered the room. Fires once per successful leg, after the far end answers the dial-out. -
conference.participant_left— one leg disconnected, withduration_seconds, a free-texthangup_reason, and the lastsip_response_codefor that leg. -
conference.ended— exactly once, always the last event for a conference. It carries three fields that close the case:
created, N participant_joined, up to N
participant_left, and exactly one ended. Anything that deviates from that
shape is the symptom — find yours in the map below. The full payload for each
event is in Webhook events.
Symptom map
The conference was created, but no participants ever joined
You receiveconference.created, then nothing — no participant_joined, and
eventually conference.ended with final_status: "failed" or a
timeout end_reason.
The room was created but every dial-out leg failed before it could join.
Per-leg failures don’t produce a participant_joined event; they are recorded
on the conference’s participant roster with an error code instead. Fetch the
conference and read the legs:
failed, busy, no_answer, …) and, on
failure, an error code and SIP response code. Match the code against the
cause-and-fix table below — a trunk problem is fixed on
the SIP trunk page, a compliance block on the
contact or the gate, a balance block by topping up.
One subtle case: when every seed participant on the create request fails to
dial, the create can complete as a 201 while the returned data.status is
already failed with CONFERENCE_DIAL_ALL_FAILED. Check
status on the create response before treating a 201 as a healthy room —
an idempotent retry replays the same failed row as another 201.
A participant joined, then dropped seconds later
You receiveparticipant_joined and an immediate participant_left for the
same leg. The dial-out succeeded — the far end answered — but the leg was
torn down right after. Read the participant_left payload:
hangup_reasonis free text from the carrier (Normal Call Clearing,caller hangup,RTP timeout). Match on substrings, never the exact string.Normal Call Clearingwithin a second or two of join usually means the far end’s own equipment (an IVR, a spam screener, a PBX) answered and hung up.sip_response_code: 200with a near-zeroduration_secondsis a clean answer-then-bye; the failure is on the destination side, not the trunk.- An absent or non-
200sip_response_codeon a short leg points back at the carrier path — work it as a trunk issue on the SIP trunk page.
The conference sits in completed but the room never really ended — or a zombie room lingers
A room closes exactly once. If the active participant count reached zero but
no conference.ended arrived, the auto-close did not run; if the room shows
completed in the dashboard but your system still treats it as live, the
ended event was missed on your side.
- Room still live with zero participants: the auto-close rolls up the
room after the last leg disconnects, stamping
end_reason: all_legs_terminated. A room that stays non-terminal with an empty roster needs one thing from you: end it explicitly.POST /api/v1/voice/conferences/{id}/end(the moderator/operator end action) closes the room and firesconference.endedwithend_reason: manual_end. Never wait for the max-durationtimeoutbackstop to do this for you — it exists to close genuinely stuck rooms, not to order your workflow. - Your system missed the closure: subscribe your webhook endpoint to
conference.ended(or*) and replay missed deliveries from the webhook deliveries page — see Troubleshooting: webhook deliveries.conference.endedis the only signal a room closed; polling the conference read endpoint forstatusis the fallback, not the primary path.
There is no end_reason on a closed conference
Older closed conferences — and rooms closed through a path that stamped only
the room row — can carry no end_reason. Treat a missing end_reason as
“reason not recorded,” not as a distinct failure: derive the story from the
roster instead. If every participant is in a terminal state (disconnected,
left, completed), the room closed because everyone left — the same
conclusion all_legs_terminated would have carried. If every participant is
in a failure state (failed, busy, no_answer, rejected), the room
closed because every leg failed.
Conferences created today always stamp an end_reason; only do this
derivation for legacy rows.
Causes and fixes
Conference leg failures carry one of two kinds of codes. Conference-level codes (CONF_*, CONFERENCE_DIAL_ALL_FAILED) come from the create step or
the room as a whole. Per-leg gate codes come from the pre-dial checks on
an individual addParticipant dial and are stored on that participant as its
failure reason.
Distinguish trunk failover from bulk leg failure while you read the
roster: one leg failing with a
5xx SIP code or a trunk code is a route
problem on that leg; every leg failing with the same code points back at
the trunk or account-level gate (balance, send block), not at the individual
destinations. Fix the shared cause once, not one participant at a time.
A leg that fails with no SIP code at all (NO_ANSWER-class) never reached
SIP — it was refused by a pre-dial gate, so the SIP trunk page can’t help;
the gate’s code is the whole story.
AI agent as conference participant
A conference can also carry an AI agent as a native participant — a voice agent bridged into the room next to human legs, not a listener on the side. Use it for an agent that takes over an interaction inside an existing room (a supervisor escalation bot, a compliance mid-call prompt, a follow-up scheduler). The join path is different from a PSTN participant, and so is its failure surface: instead of a recipient number, the room dials an internal SIP leg that the voice edge answers and connects into the room. One operation sequence to hold in mind:- Add —
POST /api/v1/voice/conferences/{id}/ai-agentwith{ "agent_id": "agent_..." }. Accepts an optional caller ID override viafrom; the platform default is used otherwise. Returns the agent attachment row id and the call sid. - Observe —
GET /api/v1/voice/conferences/{id}/ai-agentslists the recent attachment rows for the room. The row’sstatusmovesjoining→activewhen the agent answers; a failed join flips toerror. - Whisper —
POST /api/v1/voice/conferences/{id}/ai-agent/whisperwith anote(up to 1000 characters) the agent picks up on its next turn. - Remove —
DELETE /api/v1/voice/conferences/{id}/ai-agenthangs up the agent’s leg and stamps the rowleft.
participant_joined/participant_left like any other — so
everything above still applies. What changes are the join-time rejections:
The three AI-agent failure classes
A stuck
joining with no recorded error is the one failure class here you
cannot resolve from the response body alone: the join completed, but the
agent’s leg was never connected into the room. Remove and re-add once —
if the same conference id reproduces the stuck row, it’s a routing issue on
the agent-join path, so escalate (criteria below).
Paste-able samples for tickets:
What NOT to try
- Re-join loops. Retrying
POST .../ai-agentin a retry loop against a409 CONF_AI_AGENT_ALREADY_ACTIVEre-fires the same refusal every time; the only path forward is DELETE-then-POST. If your integration adds the agent on a webhook handler, make that handler idempotent: on 409, read the roster and treat an already-activerow as success, not as a reason to POST again. - DELETE without a roster read. Blindly removing “just in case” turns
a
joiningagent you wanted into a404 CONF_AI_AGENT_NOT_FOUNDon your next whisper — and the row you removed has to be re-added. ReadGET .../ai-agentsfirst; the roster is cheap. - Polling whisper for acknowledgement. Whisper is last-write-wins by design — a new note overwrites the previous one, the agent consumes it on its next turn, and the row carries no per-note acknowledgement. Don’t poll the roster for a whisper “status”: there isn’t one. If the agent should act on the note, drive that through your agent’s own escalation, not through repeated whisper calls.
- Re-creating the conference to swap the agent. Ending and recreating
the room to get a different agent in also ends every human leg. Remove
the
activeagent and add the new one in place; the room survives the swap.
When to escalate
Self-serve covers the two codes (fix the operation sequence) and a one-time stuckjoining (remove, re-add once). Escalate to support with
the conference id, the agent id, and the attachment row id from
GET .../ai-agents when:
- the same room keeps returning
409even thoughGET .../ai-agentsshows no row injoining/active(roster and rejection disagree); - the agent row flips to
erroron every add attempt across different conferences, not just one room; - the row stays
joiningpast ~30 seconds after a remove-and-re-add, or the remove-and-re-add cycle has already happened once per room. - the room closes with
final_status: "failed"at the moment you add the agent — that’s the room’s lifecycle, not the join path: work the symptom map above.
When an event stops firing
Conference events follow strict sequencing per conference id:created is
always first, participant_joined only ever precedes the matching
participant_left for that leg, and ended is always last and fires exactly
once. If your integration depends on that order, hold these expectations:
- Per-leg failure never emits
participant_joined. A leg that is refused pre-dial or fails post-dial appears only on the conference roster — it has no join/left pair. “Leg count from webhooks” and “participant count requested” diverge by exactly the failed-leg count; reconcile through the read endpoint, not by counting events. - Failure events are roster data, not webhooks. There is no
conference.participant_failedevent type; a silent failure is read off the conference, not listened for. endedalways closes the series. Any event gap beforeended(a join you expected, a left you didn’t see) is a delivery gap, not a lifecycle gap — the conference service emits in order, and delivery is the layer that retries. Work it on Troubleshooting: webhook deliveries and Troubleshooting: webhook event dedup if a delivery replay shows duplicates instead.- An
endedwith no precedingjoinedevents is valid. It means zero legs ever connected — the create-time failure path above, not a lost event.
status/end_reason once the room closes.
See also
- Troubleshooting: voice call quality — one leg joined fine but the audio is the complaint.
- Troubleshooting: SIP trunk — trunk
registration and upstream failures behind
CONF_TRUNK_UNREGISTEREDandCONF_PROVIDER_ERROR. - Troubleshooting: insufficient balance
— the top-up path for
CONF_INSUFFICIENT_CREDITand its aliases. - Troubleshooting: dialing-window blocked calls — the quiet-hours gate codes a conference leg can carry.
- Troubleshooting: webhook deliveries — when the lifecycle is healthy but a lifecycle event never reached you.
- Webhook events — the full payload contract for every conference event.
- Error codes — the HTTP-level conference error set this page works from.