Finished benchmark shape.
Dossier / live reference cards
Every card is a fact.
Hourly source check is green.
AI Resource Hub
v2 / terminal# route boot: benchmarks / meta / mistral watchlist
airh > since-last-visit
now Recomputed benchmark-weighted quality scores
now Synced Chatbot Arena benchmark track
now Updated speed measurements
now Pulled latest OpenRouter price index
airh > open /benchmarks/meta/mistral-watchlist/
Mistral
Mistral top model
Track only if a new Mistral release lands near the current frontier.
confidence: low
No current comparable operating row.
Watchlist only
No current comparable research-trust evidence.
Current signal
Watchlist only
Evidence completeness
15%
1 score dimensions represented.
DSWE
Not reported
Keep watch for a DeepSWE row.
Calculation
Score components
Correct, grounded answers on ordinary verifiable tasks.
Follows exact user constraints and output formats.
Avoids unsupported confidence and remains stable across runs.
Handles non-flashy reasoning without brittle failures.
Completes tool, coding, and agent-style tasks when evidence exists.
Cost, speed, token burn, verbosity, and friction.
Evidence
Sources used
- watchlist
Caveat
What to watch
Not ranked in the first table because there is no recent frontier-level Mistral signal comparable to the current leaders.
Meaning
How to read the score
Each score is a 0-100 saturation scale. A score of 100 means that benchmark lane is complete for this model. Floor Capability is the finished benchmark shape; Operating Envelope, Frontier Reality, and Research Trust are provisional tabs that need prompt-pack backfill. Confidence is evidence quality, not another model grade. Evidence % is source coverage, not capability.
Citations
Benchmark sources