Replication audit
Independently re-run public benchmarks; publish our measured values, run conditions, and variance—we invent no metrics, we run and record honestly.
What we do
Three layers stack: replication audits earn trust, deployment risk is the moat, longitudinal tracking compounds over time.
Independently re-run public benchmarks; publish our measured values, run conditions, and variance—we invent no metrics, we run and record honestly.
Failure boundaries, uncertainty calibration, long-horizon drift. Buyers care less about “does it look good” than “when will it deceive me.”
A version × date evidence archive. After updates, prior conclusions are labeled expired—an asset others cannot catch up to a year later.
Value
Buyers put world models into cars, robots, and factories. Others score how good generation looks; we ask whether you dare trust it.
Selection is not chasing a board. It is a defensible adoption decision under safety constraints—where failure boundaries, uncertainty, and run conditions must share one screen.
World-model evaluation is already crowded. The contribution is not another board—it is making scores independently verifiable, citable, and open to further pressure from follow-on work.
How to read
A rating is not a place on a board. Start with one diagram, learn the model card and methodology—then every issue can stand.
Now
Latest issue · Model of the Week
#1 · FeatureNVIDIA Cosmos Predict-1·Prediction: driving & robotics
Tech (architecture & capability) → application (what works, where it breaks) → business. Includes a prediction-domain rating card with run conditions and evidence strength on-screen.
Free to read until September 11, 2026
2026-08-12
Each card is an absolute rating for one world-model version × use-domain, paired with evidence strength and on-card run conditions—from different domains, not a cross-domain leaderboard.
Loading ratings…
Trust
Credibility rests on open, reproducible, time-bound ratings—run conditions are evidence.
Methodology and rating conclusions stay permanently free. N, precision, GPU, and subset status ship on the same card for anyone to verify.
Read the methodologyRatings ship as model × use-domain. We never publish a cross-domain total ranking—industry trade-offs and academic comparisons both need the dimensions kept.
Browse the ratings libraryEvery rating binds to a model version and test date; after updates, prior conclusions are labeled expired so stale evidence is never treated as present truth.
Read the rating constitutionVision
We do not build world models, and we do not compete with those we rate. We aim to be neutral evaluation infrastructure for the world-model era—so deployers in cars, robots, and factories hold evidence that stays verifiable, time-bound, and correctable, instead of unauditable vendor self-reports.
Replication pipelines, evidence archives, and use-domain protocols become reusable assets—so each world-model rating settles into a citable public record.
Interest order is constitutional: public interest first; information walls between ratings and commerce; when principle conflicts with revenue, drop the revenue.
Methodology and rating conclusions stay permanently free; corrections have deadlines; transparency reports ship even at zero revenue—authority comes from verifiability, not volume.
Model Evidence Institute (若水研究院) is public—an independent rating body for world models only.
Founder
Project Assistant Professor at The University of Tokyo, working on driving perception, world models, and safety-critical vision–language systems. Founder & CTO of AquaAge. She started Model Evidence Institute.
Personal siteJoin
Follow Model of the Week: one world model per issue—tech → application → business, ending with a rating card. Biweekly features plus mid-cycle replication briefs; membership unlocks the archive, deep reviews, and specials.
¥500 / month (tax included). Stripe checkout lands in a later milestone.