Frontier on Cloud

The standard

A test is published here only when someone else can rerun it and check every number without trusting the person who measured it. These are the rules.

What each repository contains

Each test has a public repository in frontier-on-cloud. At the commit named on the test page, it contains:

  1. A README that states what is measured, the exact model id, the SDK name and version, the endpoint (for example the Gemini API or Vertex AI), the region when it matters, and the run date. Its Reproduce section takes at most three commands: cp .env.example .env, one install command, one run script. When a setup cannot fit in three, such as a container that drops packets or a dependency patched between two runs, the README lists every step and the test page says so.
  2. A pinned environment: pyproject.toml with uv.lock and .python-version, or the equivalent for another language, such as package-lock.json.
  3. Scripts that regenerate everything from scratch: raw logs, summary tables, figures and clips. No number comes from a manual step.
  4. The raw data, committed next to the summary: JSONL timelines, audio, sidecar files. A reader can check a claim without rerunning anything.
  5. A section on what the test does not establish: fake components, synthetic inputs, N, the endpoint, anything else that limits the claim.
  6. No secret anywhere. .env is ignored, .env.example is present, and outputs are scanned before a push.

How numbers are reported

When a number looks wrong

Open an issue in the test's repository with the file and the line, or post a rerun with your commit, date, endpoint and N. See Discuss.