docs / cli / benchmarks

python -m benchmarks

Run the capability benchmark suite that builds the capability map.

Synopsis

python -m benchmarks [options]

Options

FlagEffect
--quickRun 20% of each dataset, ~4 minutes per model
--modelBenchmark a single model
--categoryBenchmark a single category
--reportShow last results without re-running
--gapsRun gap analysis: where is capability missing?
--historyShow previous runs
--reset-staleClear partial/inconsistent entries from the map