Skip to content
Self-Driving DB LabCOMP90050 · G40

Repeated runs, with uncertainty

Advisor benchmark

The arena shows one run. Here every advisor runs R times on seeded workloads, and each number comes with a 95% interval: mean workload time, paired comparisons against greedy what-if search, the bandit's regret against a hindsight optimum, and how all of it moves as the workload drifts. It runs on both datasets, in your browser.

Results

Ready

One run is an anecdote

The benchmark repeats the arena R times, each on its own seeded workload, and reports every advisor's mean cumulative workload time with a 95% bootstrap interval. Advisors replay the same workloads, in a seeded random order after one discarded warm-up replicate, so the comparison against the greedy what-if advisor (AutoAdmin) is paired: the difference, the change, the effect size d_z, an exact sign test and the share of replicates won.

The intervals are conditional on this browser session: they describe how results vary from workload to workload, not from run to run. Running the same seeds again moves measured times by more than that, so compare two runs before trusting a small difference.

The bandit's regret is measured against a hindsight reference: the configuration CoPhy's branch and bound finds best, under the what-if model, for the whole workload, and the page says whether it proved it optimal. The drift sweep repeats everything at five levels between a static and a shifting workload. How each number is computed is on the methods page.

Optional · bring your own key

Evaluating an LLM as an index advisor

Ask your own model for a configuration, review it, then tick “LLM advisor” in the settings and run the benchmark: the same seeded workloads, the same paired statistics, against the greedy what-if advisor and the bandit. Compare on “build + run” time, since a provider's response time says nothing about the advice. Each measurement is linked to its call in the audit log.

LLM index advisor (your key)

Optional. Your Anthropic (Claude) key proposes indexes from the schema, a summary of round 1 and SQLite's current plans. The lab validates them and builds only valid indexes, at the start of round 2. See the AI use statement.

No key in this browser. Everything else works without one.

Invalid-proposal rate

Per model, prompt version and dataset, over the LLM advisor calls logged in this browser. 95% intervals.

No answered proposals in this browser yet.

Calls are the independent unit, so the call-level rate has a Wilson interval. Indexes from one reply share its mistakes, so the index-level interval resamples whole calls. Provider, network and cancelled calls (0) are left out: they say nothing about the model.

Every model output is shown with an AI-generated label.