Frontier on Cloud › Tests › livekit-agents-js-2615-check
LiveKit agents-js#2615 on gemini-3.8-live: reproduction and patch check
- Model
gemini-3.8-live(base model)- Endpoint
- Gemini API (Google AI Studio key). No LiveKit server.
- Framework
@livekit/agents1.9.1,@livekit/agents-plugin-google1.9.1,@livekit/rtc-node1.1.0,@google/genai2.27.0 (plugin dependency),zod3.25.76; Node v26.9.0, macOS- Run
- 2026-10-02
- Runs
- 10: 3 + 3 with the "call first" prompt, 2 + 2 with "speak first", unpatched then patched
- Repository
- frontier-on-cloud/livekit-agents-js-2615-check (MIT) at
1559574 - Issue checked
- livekit/agents-js#2615
Question
Issue #2615 reports that in the agents-js Google plugin, with toolBehavior unset, a content-free generationComplete after a toolCall interrupts the tool's own turn, and the tool result and FunctionToolsExecuted are lost. Does that happen on the base model gemini-3.8-live, and does the patch to isNewGeneration() proposed in the issue fix it?
Setup
AgentSession.start({ agent })with no room and no job context. Input:FileAudioInput, 20 ms PCM16 frames paced in real time: 1 s of silence, then "Book me the 3pm slot tomorrow, please." (16 kHz mono, synthetic speech, 2.69 s), then silence. Gemini's default server-side voice detection handles turn-taking. Output:SimAudioOutput, a simulated real-time playout with the same contract asParticipantAudioOutput.google.beta.realtime.RealtimeModel({ model: 'gemini-3.8-live', connOptions: { maxRetry: 0 } }), withtoolBehaviorandtoolResponseSchedulingunset.- One tool,
book_slot(slot): waits 3 s and ignores the abort signal, like a real side effect. - Two prompts: "call first", and "speak first", which asks the model to say one short sentence and then call the tool in the same turn.
- Evidence: one JSONL per run with agents-js and plugin debug logs, harness marks, and a wrapper on
Session.prototype.sendToolResponseof@google/genai, so "function response sent" is checked where the SDK sends it.
Results
| Prompt | Plugin | Runs | Phantom generation after toolCall | Owning speech interrupted | FunctionToolsExecuted | Function response sent | Line spoken before the call |
|---|---|---|---|---|---|---|---|
| call first | 1.9.1 | 3 | 3/3, +3 to +4 ms | 3/3 (not aborted) | 3/3 | 3/3 | none |
| call first | patched | 3 | 0/3 | 0/3 | 3/3 | 3/3 | none |
| speak first | 1.9.1 | 2 | 2/2, +2 to +7 ms | 2/2 (aborted +2, +8 ms) | 0/2 | 0/2 | cut after 442 and 643 ms |
| speak first | patched | 2 | 0/2 | 0/2 | 2/2 | 2/2 | played in full |
- Unpatched, agents-js 1.9.1 opened a phantom agent-initiated generation 2 to 7 ms after every
toolCall(5 of 5). The trigger is the content-freegenerationCompletethat follows the call. The phantom'sinput_speech_startedinterrupts the speech that owns the call. - Call first (3 runs). The owning speech had already finished its generation tasks and was waiting on the tool. It was marked interrupted, but the code after the tool does not check that flag again, so
FunctionToolsExecutedwas emitted and the function response was sent: delivered 3 of 3. - Speak first (2 runs, the issue's case). The owning speech was still playing out, so it took the interrupted branch. The spoken line was cut after 0.4 to 0.6 s,
execute()still completed 3 s later,FunctionToolsExecutedwas never emitted, no function response was sent, and the model stayed silent until the session closed 12 s later: delivered 0 of 2. - Patched: no phantom generation and no
input_speech_startedafter the call (0 of 5), the result delivered 5 of 5, the line spoken before the call played in full, and the model spoke the post-tool confirmation.
On the base model the drop is a race: any audio before the call loses it. The result is dropped in one place, in agent_activity.ts (_realtimeGenerationTaskImpl): the if (speechHandle.interrupted) check right after waitIfNotInterrupted.
Caveats
- Small N: 3 + 3 runs with "call first", 2 + 2 with "speak first".
- The speak-first variant is an addition. The issue's own steps do not ask the model to speak first.
- Simulated audio output, no WebRTC track.
- Not tested:
gemini-3.8-live-extended-thinking, NON_BLOCKING and PR #2594, a fast-failing tool, Vertex AI. - Each run also opens a short setup-only websocket (about 0.3 s, no audio) when the session is created: 21 connections for 11 run attempts. One attempt stopped at startup and is kept as
runs/aborted-0.jsonl. - The absolute path of the WAV file in the logs was rewritten to
assets/book.wavfor publication. Nothing else in the logs was changed. - No clip for this check.
Source and reproduce
Everything is in one commit, 1559574 (2026-10-02): the harness app/run.mjs, the runs in app/runs/, the analysis app/analyze.py, and the patch as patches/dist.diff and patches/src.diff. The unmodified plugin files kept under patches/orig/ are LiveKit's, Apache-2.0, not covered by the repository's MIT license.
This check takes more than three commands: it runs the published plugin, patches it, then runs it again.
cd app
npm ci
cp ../.env.example .env # put GEMINI_API_KEY=... in it
LK_GOOGLE_DEBUG=1 node --env-file=.env run.mjs baseline-1 > runs/baseline-1.jsonl
VARIANT=speak_first LK_GOOGLE_DEBUG=1 node --env-file=.env run.mjs baseline-speakfirst-1 > runs/baseline-speakfirst-1.jsonl
python3 ../patches/apply_dist_patch.py . # exact-text edit of dist/realtime/realtime_api.js and .cjs
LK_GOOGLE_DEBUG=1 node --env-file=.env run.mjs patched-1 > runs/patched-1.jsonl
VARIANT=speak_first LK_GOOGLE_DEBUG=1 node --env-file=.env run.mjs patched-speakfirst-1 > runs/patched-speakfirst-1.jsonl
python3 analyze.py runs/*.jsonl