Native model calls
Serial, native rendering + inference + readout on the actual suite. Model loading is excluded; precision and kernels are in the receipts. Heterogeneous input lengths. This is not an HTTP service SLA.

SYSTEM ONE · REPRODUCIBLE EVALUATION · SEPTEMBER 2026
Nagi-ENORMOUS, Nagi-HUGE, Jev, OpenJev and Laya on the same public decision tasks. Accuracy is only part of the story: we also measure what each model can read, how stable its choices are, and how certain the comparison is.
Four Nagi tiers → Full-suite accuracy: SMOL 38.85% · BIG 76.14% · HUGE 77.92% · ENORMOUS 83.07%. ENORMOUS−Jev: −1.20 pp [−2.70, +0.31], statistically tied.
01 / THE COMPARISON
These are public-dataset subsets adapted to typed choice, not official leaderboard scores. Nagi's developers conducted this evaluation. The data and protocol are available for independent inspection.
Loading frozen results…
| System | Accuracy | Valid outputs | Untruncated | Order flips |
|---|
Accuracy: mean of eight source accuracies, each weighted 1/8. “Same complete evidence” uses the 4,518 core examples all local systems can process without truncation. Jev receives the same full semantic payload; its internal truncation is not observable. Full operational scores retain native truncated answers and count unsupported, invalid or missing answers as incorrect.
Results are loaded from the frozen analysis artifact.
02 / UNDER THE AVERAGE
Select a source to inspect its accuracy. A strong aggregate can hide a weak task; a small dataset can have substantial uncertainty.
03 / SPEED, WITH BOUNDARIES
Serial, native rendering + inference + readout on the actual suite. Model loading is excluded; precision and kernels are in the receipts. Heterogeneous input lengths. This is not an HTTP service SLA.
Client-observed requests include network and service time, with four requests in flight. Jev hardware is unknown. These numbers do not establish a matched-hardware speed ranking.
Separate Nagi-HUGE deployment probe: warm H100 loopback HTTP P95 145 ms at 1,024 input tokens, and 664 ms at 4,096 tokens; 30 observations/tier. Length filler measures compute cost, not long-context reasoning quality.
04 / EXPERIMENTAL DESIGN
Identical state, question and option semantics. Each system uses its pinned native prompt and readout. No test-time demonstrations, generated reasoning, prompt search or test calibration.
Native tokenizers audit all inputs before predictions. Laya truncates 153 core examples; the common-evidence track excludes those upfront. The full suite reports the real product behavior.
ANLI, XNLI, BoolQ, CB, COPA, RTE, WiC and MMLU-Pro have human or expert labels. Exact local train/dev overlap was checked. Model pretraining exposure remains unknown.
The preceding synthetic release gate did not establish a Jev win: Nagi 85.83%, Jev 88.33%; paired difference −2.50 points [−5.83, +0.67]. HUGE is released for research with this failed gate preserved.
05 / THE SYSTEMS
Qwen3.8-27B + rank-8 LoRA. Flagship line; first in Arena Live, tied with Jev here.
Nagi-HUGE ↗12BGemma 4 + rank-8 LoRA. Dev-selected step125. Third Nagi tier, alongside SMOL and BIG.
Laya ↗421M / 322MOfficial Router v0.3.20. English and multilingual checkpoints, default routing without language hints.
OpenJev / SemIf ↗4BOfficial native Qwen3.5-4B direct reference. Not the MiniCPM browser default or the unrelated 27B HF project.
Jev ↗1.13.0Actual pinned native API requests. No third-party published score substituted for measurement.
Qwen3.8 base ↗ · Gemma4 base ↗ · Qwen3.5 base ↗ · OpenJev website ↗ · Laya SDK ↗ · Nagi SDK ↗
LIGHTWEIGHT FOLLOW-UP / CONTEXT
SDK v0.4.1 raises Big’s default prompt limit to 4,096 tokens. On eight simple lookup/rule cases repeated across lengths and evidence positions, long-input correctness increased from 36/48 to 48/48; short-input probabilities were identical. This is a small engineering probe, not a general long-context guarantee.
Smol’s experimental 2,048-token mode scored 24/48, so its default remains 512. Dynamic state padding is available with the opt-in mode. Big’s observed P95 at roughly 3k state tokens was 276–285 ms: short-input latency does not apply to every length.
The benchmark tables retain SDK v0.4.0 settings (Big 768, Smol 512). To reproduce Big’s archived run with the new SDK, set max_input_tokens=768.
ADDITIONAL EXPERIMENT / NAGI TIERS
SMOL, BIG v3, HUGE and ENORMOUS on the same frozen tasks, using their public SDK defaults. SMOL/BIG/HUGE come from the three-tier analysis; ENORMOUS was scored later on the same rows.
| Model | Full-suite accuracy | All three untruncated | Truncated / unsupported | H100 P50 / P95 |
|---|
Loading tier extension…
06 / THE AUDIT TRAIL
Recompute the scores. Inspect every returned probability, failure, source ID and hash. Distinguish invoices from reserved budgets.
Data source attribution and source-specific terms are in the manifests. Canonical third-party text is reconstructed from pinned sources; raw text is not redistributed here. No implied endorsement by dataset or comparator authors.