Frontier on Cloud

Frontier on Cloud › Tests › gemini-live-resume-test

A pending tool call across a lost connection and session resumption

Model
gemini-3.8-live
Endpoint
Gemini Developer API with an API key. Not Vertex AI.
SDK
google-genai 2.25.0 with websockets 16.1.1, Python 3.13
Runs
2026-10-02 evening and 2026-10-03 morning
Sessions
43 in four rounds, 1 to 4 runs per scenario, plus 2 runs of the one-file reproduction
Repository
frontier-on-cloud/gemini-live-resume-test (MIT), numbers at 687b1af
Report to Google
google-gemini/gemini-live-api-examples#60, opened 2026-10-03 with the one-file reproduction
Reddit
Series 2 post with the clip (2026-10-03)

Question

A voice session's WebSocket connection is lost while a side-effecting tool call is pending. Can the client resume the session with its resumption handle and deliver the result, and does the user hear the truth about the booking? None of the Gemini API or Google Cloud pages read on 2026-10-02 says what happens to a tool call pending at a disconnect, how soon a resume can succeed, or which errors a resume can return.

Setup

RoundHow the connection is lostSessions
1The client aborts its own TCP connection, so the server sees it close17
2An application-level freeze: a local proxy stops forwarding bytes and closes nothing9
3Real packet loss: iptables drops the flow both ways in a Linux container, nothing is closed6
stage 2Round 3's packet loss, with a client-side recovery (recovery.py)11

Results

1. A pending tool call survives a resume that works

In all 14 runs that had a pending call and a successful resume, the resumed session accepted the FunctionResponse for the call id issued before the drop. The server never re-issued the call, never sent toolCallCancellation, and returned no error. There was exactly one booking per run. The call stayed answerable for as long as the session could be resumed: 154 s after it was issued in round 2, 83.6 s in round 3.

If the client does not send that response after the resume, the model stays at "in progress". In 3 of 3 such runs it answered "Did you book it?" with an in-progress sentence ("I'm booking the 3 pm slot for you now." in run 1), 2.9 to 3.1 s after the booking had committed.

2. The immediate resume after an idle drop is refused, and works about 1.5 s later

Round 1, where the client aborts the connection and the server sees the close:

From the drop to the resumed setupComplete: 2.0 to 2.3 s after an idle drop (n=12), 0.5 to 0.6 s after a drop during speech (n=3). Each connection received exactly one sessionResumptionUpdate, right after setupComplete, and none later.

3. After a silent loss, every resume is refused while the server holds the old connection

A silent loss is what a phone sees when it drops off Wi-Fi: packets stop, and no close frame, FIN or RST reaches the server. The client detected it 1.5 to 2.0 s after the loss.

In every silent-loss run the booking committed 1 to 3 s after the loss. In the runs where the old path never came back, the user was never told.

4. Client-side recovery (stage 2)

recovery.py is a reference pattern for that case. On a detected loss it tries to resume for 4 s (an attempt at detection, then one per second) and sends a close on the old socket in parallel. If no resume is accepted, it opens a new session without a handle and restores the conversation with one send_client_content: the last turns of the client's own transcript, plus a "System note" with one status line per side effect, taken from the backend, not from the model. The old call id is never answered. A call the model issues again is deduplicated by business key and answered from the existing job.

Two panels on a time axis since the loss. Without recovery: five runs with real packet loss, every resume refused with 1011 up to 15 minutes, the booking committed within seconds, the user never told. With recovery: four runs, loss detected, four resumes refused in a 4 s window, a new session, first model audio 6.1 to 7.1 s after the loss, and the model says the slot is booked.
The lockout (B1, B2) and the recovery (BR1). From results/figures/lockout.png, drawn by make_figure.py from the JSONL.

Other findings

The clip

MP4, 52.5 s, 1.6 MB. Two sessions recorded on 2026-10-03 for the clip in the round 3 container, not part of the 43. The user's lines are spoken by af_heart, a voice of the open-weight Kokoro-82M text-to-speech model, sent as real input. Without recovery (B1 timing, the resume loop stopped 60 s after the loss): 6 of 6 resumes refused with 1011, one booking, the model silent after its tool call. Shown in real time to 9.3 s, then compressed ×4, marked on screen. With recovery (BR1 timing): 4 refusals, a new session 3.75 s after detection, first model audio 6.57 s after the loss, one booking, no re-issued call, a correct answer. From results/clip/lockout-recovery-v2.mp4, drawn by make_clip.py.

One-file reproduction

repro_resume_lockout.py (179 lines, google-genai 2.25.0 and python-dotenv only) shows finding 3 on a laptop, without Docker or root. It embeds a small HTTP CONNECT proxy that can freeze its tunnel, routes only the first connection through it, freezes the tunnel after one text turn, and tries to resume every 5 s. As a control it then unfreezes the tunnel, closes the first connection cleanly, waits 2 s, and tries once more. Two runs on 2026-10-03, macOS: 12 of 12 attempts refused with 1011 while frozen in each run (+5 to +60 s, 510 to 676 ms and 497 to 665 ms per refusal); after the clean close, resume setupComplete after 582 ms and 543 ms. This is the reproduction in the report to Google.

What a client must do

These follow from the measurements. They are not guarantees from the API.

  1. Keep your own ledger of tool calls by call id, across connections. After a resume, send the result for the old id, even if the side effect finished during the gap.
  2. Retry a refused resume. A 1011 right after an idle drop is not a dead handle: wait about 1.5 s or retry after 1 s. Keep the handle you have; only one arrives per connection, and it restores the latest state.
  3. Detect the loss yourself with WebSocket pings. An idle healthy connection carries no server frames for many seconds, and the server never pinged the client.
  4. Get a close to the server when you can. If the path comes back, close the old connection before resuming; that is what unlocks the session.
  5. Put a short cut-off on resuming and fall back to a new session. After a silent loss on a path that does not come back, the session stayed locked for the full 15 minutes tried.
  6. In the new session, restore the status of each side effect from your backend, not only the transcript. Without it the model re-issued the call and announced a booking that was already done.
  7. Deduplicate side effects by business key, in the new session too. A re-issued call has a new id, and only the business key ties it to the old one.
  8. Do not count on transparent mode or lastConsumedClientMessageIndex on the Gemini Developer API. Buffer anything you need to replay yourself.
  9. Do not wait for goAway.
  10. Expect speech cut by a drop to be partly repeated after a resume.

Caveats

Source and reproduce

CommitDateContents
687b1af2026-10-03The harness, the 43 sessions, FINDINGS.md, the one-file reproduction
a2bc6782026-10-03The lockout figure, the first clip (macOS voice), one recorded recovery session
85a9bf82026-10-03The clip re-recorded with the af_heart user voice (the clip above)

The detailed record, with the method, every table, verbatim transcripts and the limits of each round, is FINDINGS.md. results/summary.md is generated from the JSONL by make_summary.py. No resumption handle is stored anywhere in the results.

cp .env.example .env    # then set GEMINI_API_KEY in .env
uv run repro_resume_lockout.py --seconds 60
uv run resume_test.py --scenario R1 -n 3 --results-dir out/r1

The first two commands rerun the lockout on a laptop. The third reruns one round 1 scenario; the README lists the commands for every round. Rounds 3 and stage 2 need more than three commands: a Linux container with NET_ADMIN, under Colima started with --network-address --network-preferred-route, and each B1 and B2 run lasts 15 minutes. Offline checks, no key: uv run test_freeze_proxy.py; uv run test_recovery.py.