Frontier on Cloud

Measured tests of cloud AI products

Frontier on Cloud is a set of communities, one per cloud, that test cloud AI products at the frontier and publish the measurements. Every number on this site comes from a public repository at a named commit, with the raw data committed next to the summary. Each README has a Reproduce section that reruns the test in three commands: copy the example environment file, install, run. Where a setup needs more than that, the test page says so. The rules are on the Standard page.

Communities

CloudCommunityStatus
Google Cloudr/FrontierOnGCPActive. All four tests below.
AWSr/FrontierOnAWSComing
Azurer/FrontierOnAzureComing

Tests

Newest first. Each page has the setup, the exact numbers, the caveats and the commit the numbers come from.

  1. A pending tool call across a lost connection and session resumption Google Cloud · Gemini Live API, gemini-3.8-live · 2026-10-02 to 2026-10-03 · 43 sessions · repo

    Measured: what happens to a pending booking when the Live API WebSocket is lost, and what session resumption recovers. Result: after a silent loss every resume was refused with close code 1011 while the server held the old connection (450 of 450 attempts over 15 minutes with real packet loss); a client-side fallback to a new session answered correctly in 9 of 9 runs.

  2. LiveKit agents-js#2615 on gemini-3.8-live: reproduction and patch check Google Cloud · Gemini Live API through LiveKit agents-js 1.9.1 · 2026-10-02 · 10 runs · repo

    Measured: whether a tool result is lost after a content-free generationComplete, as the issue reports, and whether the patch proposed in the issue fixes it. Result: unpatched, the result was lost in 2 of 2 runs where the model spoke before calling the tool and delivered in 3 of 3 where it called first; patched, delivered in 5 of 5.

  3. A client-side commit guard for in-flight tool calls, before and after Google Cloud · Gemini Live API, gemini-3.8-live · 2026-10-01 · 36 sessions · repo

    Measured: a reference guard (hold the commit on speech onset, dedupe by business key, send the model a status note, handle abandoned BLOCKING calls) on the stop test's speech scenarios. Result: a booking committed after the user's stop in 12 of 12 runs with the guard off and in 0 of 12 with it on.

  4. What "stop" does to an in-flight tool call on Gemini 3.8 Live Google Cloud · Gemini Live API, gemini-3.8-live · 2026-09-29 to 2026-09-30 · 41 sessions · repo

    Measured: whether the Live API cancels a pending book_slot call when the user says "stop", with text and speech input. Result: no toolCallCancellation in any of the 41 sessions; with behavior: BLOCKING the model re-issued the call and the slot was booked twice in 3 of 3 text runs.