Status
Status & illustrative leaderboard
The headline signal is skill in bits over the own-routine baseline R2, with calibration and permutation specificity as gates. The "overall" column is an illustrative roll-up for display only, not a canonical composite.
⚠
No empirical leaderboard results exist for v1.0. Every row below is a synthetic illustrative mock baseline on benchmark protocol v1.0 — not a real submission, not an official result, and not evidence about any system or product. Higher is better for every metric except calibration error and long-horizon degradation (lower is better). A large prediction count from few targets is not a large independent sample.
Loading leaderboard…
Metric definitions
- Overall score — illustrative pilot roll-up for display only.
- Target adaptation gain — Skill in bits over the own-routine baseline R2 (headline).
- Calibration error — top-label ECE; lower is better.
- Evidence attribution — whether cited evidence supports the prediction.
- Counterfactual consistency — coherence of action-conditioned predictions.
- Temporal-order sensitivity / wrong-target penalty — skill lost under shuffled history / the permutation gate (higher = more target-specific).
- Long-horizon degradation — skill decay at long horizon; lower is better.
- Modality contribution — marginal skill from added evidence streams.