Frontier on Cloud

Frontier on Cloud › Tests › gemini-live-stop-test

What "stop" does to an in-flight tool call on Gemini 3.8 Live

Model
gemini-3.8-live
Endpoint
Gemini API (Google AI Studio, API key). Not Vertex AI.
SDK
google-genai 2.25.0, Python 3.13
Runs
2026-09-29 (text from 08:27 CEST, speech from 09:47 CEST) and 2026-09-30 (from 22:03 CEST)
Sessions
41: 21 text, 20 speech. N=3 per scenario, plus two smoke sessions.
Repository
frontier-on-cloud/gemini-live-stop-test (MIT), commits listed below
Report to Google
google-gemini/gemini-live-api-examples#59, opened 2026-10-02 with the one-file reproduction
Reddit
Part 1, text (2026-09-29) · Part 2, speech and the clip (2026-09-30)

Question

A voice assistant on the Live API starts booking a slot through a tool. One second later the user says "Actually, stop. Don't book it." The booking takes 4 s to commit. Does the API tell the app to cancel the pending call, and what does the model tell the user?

The Live API reference describes toolCallCancellation as a notice that a previously issued tool call should be cancelled, and says it occurs "only in cases where the clients interrupt server turns". The gemini-3.8-live model page says NON_BLOCKING is now the default function calling mode.

Setup

ScenarioTool behaviorWhen the stop is sent
A, Bunset (model default, NON_BLOCKING). A and B differ only in whether the harness would act on a cancellation.1.0 s after the tool call
CBLOCKING1.0 s after the tool call
Dunset5.5 s after the tool call, after the commit
Eunset, tool response with scheduling: SILENT1.0 s after the tool call
Funset0.3 s after the request, before any tool call (speech: 0.3 s after the end of the request clip)
G2unset, 7.0 s service, a follow-up question 0.5 s after the call so the model speaks while the call is pending1.0 s into that speech, call still pending (speech only)

Results

Part 1: text input, 2026-09-29

Timeline of run 1 of the default scenario and run 1 of the BLOCKING scenario, text input. Default: tool call at 0.808 s, stop at 1.809 s, the model says the booking was already made at 2.639 s, the booking commits at 4.811 s. BLOCKING: tool call at 0.891 s, stop at 1.892 s, interrupted at 1.910 s, second tool call at 2.714 s, bookings commit at 4.893 s and 6.716 s.
Run 1 of A (default) and run 1 of C (BLOCKING), text input, on one time axis. From results/figures/stop-timeline-A1-C1.png, drawn by make_figure.py.

Part 2: speech input, 2026-09-29 and 2026-09-30

Counts are runs out of sessions. Offsets are ms after the stop: for text, after the stop text was sent; for speech, after the first chunk of the stop clip, so they include the server's speech detection delay. Text A is A plus B.

ScenarioInputCancellationinterrupted after the stopDouble bookingTool call after an early stop
Atext0/60/60/6n/a
Aspeech0/30/30/3n/a
Ctext0/33/3, +16 to +18 ms3/3n/a
Cspeech0/33/3, +149 to +152 ms1/3n/a
Dtext0/33/3, +15 to +16 ms0/3n/a
Dspeech0/33/3, +139 to +144 ms0/3n/a
Ftext0/30/30/32/3, +1047 and +1115 ms
Fspeech0/30/30/33/3, +3465 to +3475 ms

Across both days: 41 sessions (21 text, 20 speech), none with a toolCallCancellation.

The clip

Run 1 of the BLOCKING re-run with speech, real time, with the model's own audio. MP4, 15 s, 373 KB. The stop clip starts at 5183 ms. interrupted arrives at 5330 ms (+147 ms). The fake service commits BK-1001 at 8181 ms and the tool response status: booked goes out at 8186 ms. From 9437 ms the model says "The booking was not made, so nothing has been scheduled." No toolCallCancellation. From results/clip/stop-test-C1.mp4, drawn by make_clip.py.

One-file reproduction, 2026-10-02

repro_blocking_reissue.py (119 lines, google-genai and python-dotenv only) reproduces scenario C with typed text. In 3 BLOCKING runs, interrupted arrived 14, 21 and 14 ms after the stop; a second toolCall with a new id came 713 and 469 ms after the stop in runs 1 and 2; no run had a toolCallCancellation. In run 1 the re-issued call had the slot tomorrow at 3pm where the first had tomorrow 3pm. Raw output is in results/repro/. This is the reproduction in the report to Google.

What it means for an app

On this model the Live API does not hand the app a cancellation for a pending tool call. The app holds the call id and runs the job, so "stop" has to be handled on the client: hold the side effect for a short grace window, deduplicate repeated calls for the same booking, and take the booking status from the backend, not from the model's sentence. The follow-up test, a client-side commit guard, measures that.

Caveats

Source and reproduce

CommitDateContents
fc6ab7a2026-09-29Text runs (A to F, the D re-run) and the 12 speech sessions
687cc722026-09-30Scenario G and the BLOCKING re-run with saved model audio
e8b7c992026-09-30Clip script and clip
0d8ab9b2026-10-02One-file reproduction and its output

The numbers above come from results/summary.md and the README findings, built from the JSONL timelines committed next to them.

cp .env.example .env    # then set GEMINI_API_KEY in .env
uv sync
./run_all.sh            # text, scenarios A to F, N=3

Speech: ./run_audio.sh. Scenario G2 and the BLOCKING re-run: ./run_g.sh. One-file reproduction: uv run repro_blocking_reissue.py. The key is read only from GEMINI_API_KEY and never logged.