若水研究院

Docs

Beginner's guide

One diagram to see what MEI rates, how we rate it, and how to read a conclusion. World models only. No cross-domain leaderboard.

World-model rating system at a glanceFlow from rated object to dual-axis conclusion, seven dimensions, three method layers, and run conditions.01What we rateModel version × use-domain02Dual-axis callLevel × evidence strength03Seven dimsShown separatelyThree method layersL1Replication audit

Independently re-run public benchmarks; publish values and conditions.

L2Independent probes

Fill gaps on uncertainty calibration and failure boundaries.

L3Rating judgment

Level × strength by use-domain, with a reasoning chain.

04Run conditions are evidenceN · precision · GPU · subset / offload — on the same card
Read left to right, then down: version × domain → level × strength → seven dimensions → L1–L3 sources → run conditions as the trust boundary.

The rated object is always world-model version × use-domain. Domains lack a reliable common order, so we never publish a total ranking.

Conclusions use two axes: level (1 / 2 / 3 / NR) for maturity, evidence strength (high / medium / low) for how strong the evidence is. They must appear as a pair.

Seven dimensions stay separate—never one total score. Run conditions (N, precision, GPU, subset, …) ship on the same card. Beyond “does it look good,” we ask “dare you trust it.”