Nagi: Typed decisions. Visible probabilities.
New · Arena Live · 25 September 2026Four AIs. Three real-time games. Watch every decision.Nagi-ENORMOUS 83/90 points · 23 of 30 rounds won vs Jev, OpenJev and Laya.Watch the replays →

SYSTEM ONE · REPRODUCIBLE EVALUATION · SEPTEMBER 2026

Fast decisions.
Open evidence.

Nagi-ENORMOUS, Nagi-HUGE, Jev, OpenJev and Laya on the same public decision tasks. Accuracy is only part of the story: we also measure what each model can read, how stable its choices are, and how certain the comparison is.

Four Nagi tiers → Full-suite accuracy: SMOL 38.85% · BIG 76.14% · HUGE 77.92% · ENORMOUS 83.07%. ENORMOUS−Jev: −1.20 pp [−2.70, +0.31], statistically tied.

4,671public examples
8equally weighted sources
320option-order probes
4languages in XNLI

01 / THE COMPARISON

One table. Visible trade-offs.

These are public-dataset subsets adapted to typed choice, not official leaderboard scores. Nagi's developers conducted this evaluation. The data and protocol are available for independent inspection.

Loading frozen results…

Model comparison on the frozen public evaluation
SystemAccuracyValid outputsUntruncatedOrder flips

Accuracy: mean of eight source accuracies, each weighted 1/8. “Same complete evidence” uses the 4,518 core examples all local systems can process without truncation. Jev receives the same full semantic payload; its internal truncation is not observable. Full operational scores retain native truncated answers and count unsupported, invalid or missing answers as incorrect.

WHAT THE EVIDENCE SAYS

Results are loaded from the frozen analysis artifact.

02 / UNDER THE AVERAGE

Generalization has a shape.

Select a source to inspect its accuracy. A strong aggregate can hide a weak task; a small dataset can have substantial uncertainty.

03 / SPEED, WITH BOUNDARIES

Hardware is part of the result.

LOCAL · NVIDIA H100

Native model calls

Serial, native rendering + inference + readout on the actual suite. Model loading is excluded; precision and kernels are in the receipts. Heterogeneous input lengths. This is not an HTTP service SLA.

REMOTE · JEV API

Network + service

Client-observed requests include network and service time, with four requests in flight. Jev hardware is unknown. These numbers do not establish a matched-hardware speed ranking.

Separate Nagi-HUGE deployment probe: warm H100 loopback HTTP P95 145 ms at 1,024 input tokens, and 664 ms at 4,096 tokens; 30 observations/tier. Length filler measures compute cost, not long-context reasoning quality.

04 / EXPERIMENTAL DESIGN

Designed for
independent inspection.

01

Same decision, native interfaces

Identical state, question and option semantics. Each system uses its pinned native prompt and readout. No test-time demonstrations, generated reasoning, prompt search or test calibration.

02

Coverage before correctness

Native tokenizers audit all inputs before predictions. Laya truncates 153 core examples; the common-evidence track excludes those upfront. The full suite reports the real product behavior.

03

Public does not mean unseen

ANLI, XNLI, BoolQ, CB, COPA, RTE, WiC and MMLU-Pro have human or expert labels. Exact local train/dev overlap was checked. Model pretraining exposure remains unknown.

04

Failed hypotheses stay visible

The preceding synthetic release gate did not establish a Jev win: Nagi 85.83%, Jev 88.33%; paired difference −2.50 points [−5.83, +0.67]. HUGE is released for research with this failed gate preserved.

05 / THE SYSTEMS

Follow the models. Inspect the pins.

Qwen3.8 base ↗ · Gemma4 base ↗ · Qwen3.5 base ↗ · OpenJev website ↗ · Laya SDK ↗ · Nagi SDK ↗

LIGHTWEIGHT FOLLOW-UP / CONTEXT

A larger limit, tested.

SDK v0.4.1 raises Big’s default prompt limit to 4,096 tokens. On eight simple lookup/rule cases repeated across lengths and evidence positions, long-input correctness increased from 36/48 to 48/48; short-input probabilities were identical. This is a small engineering probe, not a general long-context guarantee.

Smol’s experimental 2,048-token mode scored 24/48, so its default remains 512. Dynamic state padding is available with the opt-in mode. Big’s observed P95 at roughly 3k state tokens was 276–285 ms: short-input latency does not apply to every length.

The benchmark tables retain SDK v0.4.0 settings (Big 768, Smol 512). To reproduce Big’s archived run with the new SDK, set max_input_tokens=768.

Usage and limits ↗

ADDITIONAL EXPERIMENT / NAGI TIERS

What does a larger Nagi buy?

SMOL, BIG v3, HUGE and ENORMOUS on the same frozen tasks, using their public SDK defaults. SMOL/BIG/HUGE come from the three-tier analysis; ENORMOUS was scored later on the same rows.

Native Nagi tier comparison
ModelFull-suite accuracyAll three untruncatedTruncated / unsupportedH100 P50 / P95

Loading tier extension…

06 / THE AUDIT TRAIL

Don't take the table on faith.

Recompute the scores. Inspect every returned probability, failure, source ID and hash. Distinguish invoices from reserved budgets.

Data source attribution and source-specific terms are in the manifests. Canonical third-party text is reconstructed from pinned sources; raw text is not redistributed here. No implied endorsement by dataset or comparator authors.