Frontier on Cloud › Tests › gemini-live-stop-test
What "stop" does to an in-flight tool call on Gemini 3.8 Live
- Model
gemini-3.8-live- Endpoint
- Gemini API (Google AI Studio, API key). Not Vertex AI.
- SDK
google-genai2.25.0, Python 3.13- Runs
- 2026-09-29 (text from 08:27 CEST, speech from 09:47 CEST) and 2026-09-30 (from 22:03 CEST)
- Sessions
- 41: 21 text, 20 speech. N=3 per scenario, plus two smoke sessions.
- Repository
- frontier-on-cloud/gemini-live-stop-test (MIT), commits listed below
- Report to Google
- google-gemini/gemini-live-api-examples#59, opened 2026-10-02 with the one-file reproduction
- Part 1, text (2026-09-29) · Part 2, speech and the clip (2026-09-30)
Question
A voice assistant on the Live API starts booking a slot through a tool. One second later the user says "Actually, stop. Don't book it." The booking takes 4 s to commit. Does the API tell the app to cancel the pending call, and what does the model tell the user?
The Live API reference describes toolCallCancellation as a notice that a previously issued tool call should be cancelled, and says it occurs "only in cases where the clients interrupt server turns". The gemini-3.8-live model page says NON_BLOCKING is now the default function calling mode.
Setup
- One tool,
book_slot(slot), backed by a fake in-process booking service that commits 4.0 s after the call (7.0 s in G2). The tool response goes out as soon as the job commits. - The user says "Book me the 3pm slot tomorrow, please.", then "Actually, stop. Don't book it."
- Part 1: typed text, sent with
send_realtime_input(text=...). Part 2: synthetic speech (macOSsay, voice Samantha, 16 kHz 16-bit mono PCM), streamed in 100 ms chunks in real time, with server voice activity detection at its defaults. - Output: AUDIO with output transcription. The server rejected a TEXT-only session with close code 1007.
- Every server message is logged with a timestamp in ms since session start.
| Scenario | Tool behavior | When the stop is sent |
|---|---|---|
| A, B | unset (model default, NON_BLOCKING). A and B differ only in whether the harness would act on a cancellation. | 1.0 s after the tool call |
| C | BLOCKING | 1.0 s after the tool call |
| D | unset | 5.5 s after the tool call, after the commit |
| E | unset, tool response with scheduling: SILENT | 1.0 s after the tool call |
| F | unset | 0.3 s after the request, before any tool call (speech: 0.3 s after the end of the request clip) |
| G2 | unset, 7.0 s service, a follow-up question 0.5 s after the call so the model speaks while the call is pending | 1.0 s into that speech, call still pending (speech only) |
Results
Part 1: text input, 2026-09-29
toolCallCancellation: 0 of 18 sessions (A, B, C, E, the D re-run and F). The first D pass, 3 sessions, ended early because of a bug in the harness stop condition. It is kept in the data, marked, and left out of this count. It had no cancellation either.- Default mode (A, B, E, 9 runs). The model's turn held only the
book_slotcall and no speech, so the stop interrupted nothing. 0.5 to 0.9 s after the stop the model said, in some form, that the booking was already made. At that moment the fake service was still about 2 s from committing, and no tool response had been sent. 9 of 9. - BLOCKING (C).
interruptedarrived 16 to 18 ms after the stop, with no cancellation. 0.6 to 0.8 s later the model issued a secondbook_slotcall with a new id and the same slot. Both bookings committed, and the model said the slot was already booked. 3 of 3. - Stop after the commit (D re-run).
interrupted15 to 16 ms after the stop, no cancellation. 0.5 to 0.7 s after the stop the model said the booking was already made, which was true this time. It is the same kind of sentence as in A, B and E, where it was not true. - Stop before the call (F). No
interrupted, no cancellation.book_slotstill arrived in 2 of 3 runs, 1047 and 1115 ms after the stop; the booking committed and the model said it was already made. In run 1 no call came, and the model said "I haven't made that booking for you." - E (
scheduling: SILENT). The model did not speak after the tool response.
results/figures/stop-timeline-A1-C1.png, drawn by make_figure.py.Part 2: speech input, 2026-09-29 and 2026-09-30
Counts are runs out of sessions. Offsets are ms after the stop: for text, after the stop text was sent; for speech, after the first chunk of the stop clip, so they include the server's speech detection delay. Text A is A plus B.
| Scenario | Input | Cancellation | interrupted after the stop | Double booking | Tool call after an early stop |
|---|---|---|---|---|---|
| A | text | 0/6 | 0/6 | 0/6 | n/a |
| A | speech | 0/3 | 0/3 | 0/3 | n/a |
| C | text | 0/3 | 3/3, +16 to +18 ms | 3/3 | n/a |
| C | speech | 0/3 | 3/3, +149 to +152 ms | 1/3 | n/a |
| D | text | 0/3 | 3/3, +15 to +16 ms | 0/3 | n/a |
| D | speech | 0/3 | 3/3, +139 to +144 ms | 0/3 | n/a |
| F | text | 0/3 | 0/3 | 0/3 | 2/3, +1047 and +1115 ms |
| F | speech | 0/3 | 0/3 | 0/3 | 3/3, +3465 to +3475 ms |
voiceActivityACTIVITY_STARTfor the stop arrived 139 to 154 ms after the first chunk of the stop clip (A, C, D). In C and D,interruptedarrived in the same millisecond or 1 ms later.- BLOCKING with speech (C). Run 2 booked twice. In runs 1 and 3 one booking committed and the tool response said
status: booked, then the model said "I have not booked the slot. It has been canceled as requested." and "I've stopped the process, and the booking booking was not made." - Default with speech (A). The "already booked" reply started 1056 to 1130 ms after the commit. With speech and a 4 s service, the text-mode finding (the claim before the commit) was not reproduced.
- Early stop with speech (F). The server transcribed the request and the stop as one turn and sent
book_slotafter it, 3 of 3. - Barge-in while the model speaks (G2, a smoke session plus N=3, 2026-09-30). After the follow-up question the model spoke while the call was pending, for example "I am booking the 3 PM slot for you now. Please bring your ID and any necessary paperwork to the appointment." The stop interrupted that speech:
interruptedarrived in the same millisecond asACTIVITY_START, 144 to 238 ms after the stop clip started. NotoolCallCancellation, no second call, and the pending call was never cancelled. The booking committed 0.9 to 1.0 s afterinterrupted. - BLOCKING re-run with the model's audio saved (C, 2026-09-30). Run 2 booked twice again. In runs 1 and 3, 1256 and 1172 ms after the commit, the model said "The booking was not made, so nothing has been scheduled." and "The booking was not made as you requested."
Across both days: 41 sessions (21 text, 20 speech), none with a toolCallCancellation.
The clip
interrupted arrives at 5330 ms (+147 ms). The fake service commits BK-1001 at 8181 ms and the tool response status: booked goes out at 8186 ms. From 9437 ms the model says "The booking was not made, so nothing has been scheduled." No toolCallCancellation. From results/clip/stop-test-C1.mp4, drawn by make_clip.py.One-file reproduction, 2026-10-02
repro_blocking_reissue.py (119 lines, google-genai and python-dotenv only) reproduces scenario C with typed text. In 3 BLOCKING runs, interrupted arrived 14, 21 and 14 ms after the stop; a second toolCall with a new id came 713 and 469 ms after the stop in runs 1 and 2; no run had a toolCallCancellation. In run 1 the re-issued call had the slot tomorrow at 3pm where the first had tomorrow 3pm. Raw output is in results/repro/. This is the reproduction in the report to Google.
What it means for an app
On this model the Live API does not hand the app a cancellation for a pending tool call. The app holds the call id and runs the job, so "stop" has to be handled on the client: hold the side effect for a short grace window, deduplicate repeated calls for the same booking, and take the booking status from the backend, not from the model's sentence. The follow-up test, a client-side commit guard, measures that.
Caveats
- The booking service is a fake in-process
asynciojob. Cancelling it isasyncio.Task.cancel(). - Part 1 used typed text. Part 2 used synthetic speech: one macOS voice, one wording per utterance, no background noise, no room acoustics, no echo of the model's own audio.
- Gemini API with an API key only. Not Vertex AI, not other models, not session resumption.
- N=3 per scenario, plus one smoke session each for G1 and G2, over two days. It shows what can happen, not how often.
- The system prompt asked the model to say plainly whether the booking was made. A prompt that forbids status claims without a tool result may change the wording. It does not produce the missing cancellation.
- G needed a follow-up question to make the model speak while the call was pending. With the G1 wording, one combined request, it did not.
- "What the model said after the stop" counts everything received after the stop was sent, including output generated just before it that arrived late.
Source and reproduce
| Commit | Date | Contents |
|---|---|---|
fc6ab7a | 2026-09-29 | Text runs (A to F, the D re-run) and the 12 speech sessions |
687cc72 | 2026-09-30 | Scenario G and the BLOCKING re-run with saved model audio |
e8b7c99 | 2026-09-30 | Clip script and clip |
0d8ab9b | 2026-10-02 | One-file reproduction and its output |
The numbers above come from results/summary.md and the README findings, built from the JSONL timelines committed next to them.
cp .env.example .env # then set GEMINI_API_KEY in .env
uv sync
./run_all.sh # text, scenarios A to F, N=3
Speech: ./run_audio.sh. Scenario G2 and the BLOCKING re-run: ./run_g.sh. One-file reproduction: uv run repro_blocking_reissue.py. The key is read only from GEMINI_API_KEY and never logged.