python -m benchmarks
Run the capability benchmark suite that builds the capability map.
Synopsis
python -m benchmarks [options]
Options
| Flag | Effect |
|---|---|
--quick | Run 20% of each dataset, ~4 minutes per model |
--model | Benchmark a single model |
--category | Benchmark a single category |
--report | Show last results without re-running |
--gaps | Run gap analysis: where is capability missing? |
--history | Show previous runs |
--reset-stale | Clear partial/inconsistent entries from the map |