Frontier on Cloud

Frontier on Cloud › Tests › gemini-live-commit-guard

A client-side commit guard for in-flight tool calls, before and after

Model
gemini-3.8-live
Endpoint
Gemini API (Google AI Studio, API key). Not Vertex AI.
SDK
google-genai 2.25.0, Python 3.13.7
Runs
2026-10-01, 08:35 to 08:45 CEST, ablations 08:52 to 08:56 CEST
Sessions
36, speech input: 24 with the guard off and on (4 scenarios, N=3 each), 12 ablations
Repository
frontier-on-cloud/gemini-live-commit-guard (MIT), numbers at 20330de
Follows
What "stop" does to an in-flight tool call
Reddit
Part 3, the commit guard and the before/after clip (2026-10-02)

Question

The stop test found that on gemini-3.8-live the client never receives toolCallCancellation, that in the default NON_BLOCKING mode a later "stop" interrupts nothing, that in BLOCKING mode the server drops the pending call and the model may re-issue it with a new id, and that the model's account of the booking can contradict the backend. Can the client handle all of this itself, and which parts of the handling matter?

Setup

  1. Hold on speech onset. After prepare(), a grace window of 1.5 s. If the user starts speaking (ACTIVITY_START) while a job is pending, the commit is held until that utterance's input transcript arrives, and a keyword check decides: stop, cancel, don't, do not, wait, never mind, hold on. The utterance that triggered the call is checked too. With no transcript 3 s after the later of the speech onset and the moment the commit became due, the job is cancelled.
  2. Deduplicate by business key. The key is the tool name plus the canonical arguments (the slot reduced to day plus 24 h time, tomorrow 15:00). A call whose key matches a job from the last 30 s is not executed; it is answered with that job's result.
  3. Status note. After a cancel decision, or a commit that happened despite a stop, the guard sends the model a user turn with send_client_content(..., turn_complete=True), for example "System note: the booking for tomorrow 3pm was cancelled before it was committed. Nothing is scheduled." The FunctionResponse still carries the real status.
  4. Abandoned BLOCKING calls. If interrupted arrives while a BLOCKING call has no response yet, the guard marks the call abandoned, holds the commit and does not answer the dead id. A re-issued call for the same key gets the existing job's status on its live id.
ScenarioWhat happens
ADefault (NON_BLOCKING); stop clip 1.0 s after the call; 4 s prepare
CBLOCKING; stop clip 1.0 s after the call; 4 s prepare
G2Follow-up question 0.5 s after the call, so the model speaks while the call is pending; stop clip 1 s into that speech; 7 s prepare
FStop clip 0.3 s after the end of the booking clip; 4 s prepare

Unwanted commit: a booking committed after the stop clip had started. Statement consistent: the model's turns after the stop clip, classified from the transcript as "claims booked", "claims not booked" or neither, compared with the fake service; neither counts as inconsistent. The last claim and the first claim are scored separately. The rules were checked by hand on 51 post-stop statements.

Results

Counts are runs out of sessions. "Off, stop test" rescores the stop test's speech sessions (no guard, same prompt, clips, latency and timing) with the same metric code; for C it includes the stop test's re-run.

ScenarioGuardSessionsUnwanted commitDouble bookingStatement consistent (last claim)First claim consistentStop to decision, ms (median, range)Note cut a started reply
Aoff, stop test33/30/33/33/3--
Aoff33/30/33/33/3--
Aon30/30/33/33/33481 (3479 to 3503)1/3
Coff, stop test66/62/62/62/6--
Coff33/32/32/32/3--
Con30/30/33/33/33477 (3473 to 3507)1/3
G2off, stop test33/30/33/33/3--
G2off33/30/33/33/3--
G2on30/30/33/33/33466 (3463 to 3466)1/3
Foff, stop test33/30/33/33/3--
Foff33/30/33/33/3--
Fon30/30/33/33/33494 (3469 to 3827)0/3

With the guard off the model usually said "already made", which was true, so most off runs score as consistent, including the C runs that booked twice. The exceptions are C off run 3 ("The booking was not made." after a commit) and 4 of the 6 stop-test C runs.

Dot chart. Booked after the stop: A, C, G2 and F each 3 of 3 with the guard off and 0 of 3 with it on. First reply matched the booking state, no status note against full guard: A 0 of 3 against 3 of 3, C 3 of 3 against 3 of 3, G2 0 of 3 against 3 of 3, F not run. Double booking in C: 2 of 3 off, 0 of 3 on.
Each dot is one run. From results/figures/before-after.png, drawn by make_figure.py.

Which behaviors mattered

Ablations

Guard on with one behavior switched off, N=3 each. Compare with the off and on rows above.

ConfigScenarioUnwanted commitDouble bookingStatement consistent (last claim)First claim consistent
no status noteA0/30/33/30/3
no status noteC0/30/33/33/3
no status noteG20/30/33/30/3
no hold, no abandonC3/30/33/33/3

The note can cut a reply in half

A note sent with turn_complete=True interrupts a reply already in progress. In 8 of the 12 full-guard runs, interrupted arrived 12 to 78 ms after the note. In 3 runs the model had already emitted text. In C run 3 the reply was split "I have not" / "booked that slot for you." A client that drops its queued audio on interrupted plays only the second half, which says the opposite of the truth. Two options were not tested: send the note only while the model is idle, or carry the status in a FunctionResponse with a scheduling mode.

The clip

MP4, 33.7 s, 802 KB. First half: the stop test's BLOCKING run without the guard (the clip on the stop test page). Second half, real time with the model's audio: one extra C session with the full guard, recorded on 2026-10-02 for the clip and not counted in the tables (rule: cancelled before commit, note sent, a clear correct reply the note did not cut; up to 3 tries allowed; run 1 met it). The guard cancels the job at 8720 ms on the stop transcript, 0.50 s after prepare finished. The model re-issues book_slot at 8721 ms and dedupe answers cancelled at 8724 ms. The note goes out at 8726 ms, and from 10063 ms the model says "The booking was cancelled and nothing is scheduled." No toolCallCancellation. From results/clip/before-after.mp4, drawn by make_clip.py.

Caveats

Source and reproduce

CommitDateContents
20330de2026-10-01The guard, the harness, the 36 sessions and results/summary.md
844258b2026-10-01Before/after figure and its script
229d9442026-10-02Audio capture, the one recorded session for the clip (not counted), the clip
3dc74912026-10-03README correction: the abandon count is 6 of 6 runs, re-issue in 4 of 6 (it read 9 of 9 and 4 of 9)

Per-run tables, metric definitions and the rescored stop-test sessions: results/summary.md. Every event and guard decision is in results/<scenario>_<mode>.jsonl.

cp .env.example .env    # then set GEMINI_API_KEY in .env
uv sync
./run_guard.sh          # A, C, G2, F with the guard on, then off, N=3: 24 sessions

Offline, no key and no network: uv run python -m unittest -v (18 tests that replay recorded timelines). Ablations: ./run_guard.sh A_noinject C_noinject G2_noinject C_nohold_noabandon. Rebuild the summary: uv run aggregate.py.