World Model Evaluation Methodology
Not the 21st benchmark: independently re-run public scores (L1), probe uncertainty calibration and failure boundaries (L2), then rate by use-domain as level × evidence strength (L3). Seven dimensions stay separate; run conditions ship on every card. Full v0.1 will be published.