Expanded evaluation in progress. This page currently reports the completed pilot. Full BoolQ, RTE and WiC validation benchmarks, larger synthetic sets and repeated timing/cache measurements are being added.
Model evaluation · OpenRouterExploratory study / v2
Empirical technical report

Decision quality, latency and cost:
GPT-6 Luna and Jev

A paired evaluation with matched probability outputs and a controlled prompt-caching experiment.

Abstract

Objective. Compare two deployed decision services on answer quality, complete-response latency and billed cost, while constraining GPT-6 Luna to Jev’s native probability-of-yes answer schema.

Methods. We evaluated 180 paired cases: 60 synthetic policy decisions, 60 BoolQ reading questions and 60 problems with exact event probabilities. A separate 24-case experiment compared Jev, uncached Luna and cached Luna on a shared 1,903-word rulebook. Requests were sequential, with randomized model order within each case. Reported probabilities, raw payloads, timings and billed usage were retained.

Results. Both models answered 60/60 simple policy cases correctly. Agreement with published BoolQ labels was 57/60 for Jev and 49/60 for Luna. Probability mean absolute error was 6.44 and 3.28 percentage points, respectively. Across the short-input tasks, median latency was 0.238 s for Jev and 1.143 s for Luna. On the long-prefix workload, Luna cache reads were observed in 24/24 calls; cached Luna cost 76.2% less than uncached Luna after priming, and 34.9% less than Jev after allocating the extra prime across 24 decisions.

Interpretation. Jev was faster in this run. Quality varied by task, and prefix reuse changed the cost ordering. Small samples, one measurement window, public-data exposure and identified BoolQ label limitations constrain generalization.

Keywords: model evaluation · binary decisions · probability estimation · latency · billed inference cost · prompt caching

1Research questions

The practical question is whether a general language model constrained to compact decision outputs can match a dedicated decision model on quality, speed and price. Output length alone does not control the comparison: reasoning settings, request encoding, provider overhead and input reuse also matter.

  1. Quality: How do the models compare on deterministic rules, passage-based answers and exact probability calculations?
  2. Service performance: What complete-response latency and billed cost are observed for the same cases?
  3. Input reuse: Does a verified cached prefix make Luna cheaper, and how does the extra priming call change the comparison?

This is an exploratory evaluation of two service configurations. No single composite quality score is used because the tasks have different targets and error meanings.

2Methods

2.1 Models, output contract and controls

Jev was requested as typesafe/jev-1.13 and resolved to typesafe/jev-1.13-20260917 through TypeSafe. Luna was requested and returned as openai/gpt-6-luna, pinned to the OpenAI provider with fallbacks disabled. Jev used the OpenRouter Decisions endpoint; Luna used Chat Completions. Model IDs and providers were checked on every response. The API envelopes differ; the answer object is identical.1–3

Matched answer object
{"type":"noul","noul":0.73}
The number is p = P(yes); P(no) = 1 − p. A binary answer is “yes” when p ≥ 0.5. The value 0.73 above illustrates the schema.

Luna used strict JSON Schema, a maximum of 128 output tokens, reasoning.effort = none and non-streaming responses. Temperature was left unspecified. All Luna responses recorded zero reasoning tokens. Jev supplied its native Noul probabilities; Luna generated numerical probabilities. Neither adapter used next-token log probabilities.

Both models received the same state, question and yes/no criteria. Luna additionally received a format instruction and schema. Case IDs, target answers and source metadata were excluded from model input. The exact instruction, schema and example requests appear in Appendix A; every recorded request is inspectable in Appendix C.

2.2 Evaluation data

Table 1. Frozen evaluation sets and reference targets.
ExperimentCasesCompositionReference target
Simple policies60Refund, delivery and laboratory access; 20 each; 30 yes / 30 no overallExact Boolean rules evaluated in Python
BoolQ reading60Uniform sample from 3,270 validation rows; 36 yes / 24 noPublished human label
Exact probabilities60Conditional population tables, draws without replacement and weighted mixtures; 20 eachExact rational event probability
Long policy / cache24One shared dispatch rulebook and varied records; 12 yes / 12 noDeterministic rule oracle with independently specified expected labels

The base seed was 20261002; the probability generator used the derived seed 20261003. Request order used the base seed. Simple policy cases were balanced within family and label, then sampled uniformly within each label. This does not exhaust all boundaries or policy branches. Exact probabilities were computed from counts with rational arithmetic and independently checked by enumeration in tests.

BoolQ rows were sampled without replacement and without filtering by labels or model outputs. The original Google Storage download returned HTTP 403, so the complete Google-owned Hugging Face validation mirror was retrieved and normalized to JSONL. The manifest records the mirror, row indices and normalized-file hash. The 60 selected passages and labels retain BoolQ’s CC BY-SA 3.0 attribution.5–7

2.3 Request schedule and timing

The main run shuffled cases and randomized model order within each pair. Calls were sequential over a shared persistent Python HTTPS connection to OpenRouter. Two warmup cases per model were excluded from the measured sample. Four initial setup probes were also excluded. There were no automatic retries or answer repairs.

Latency was measured with a monotonic clock immediately before connection setup/request, through receipt, JSON parsing and schema validation of the complete response. It includes network and service overhead. Payload serialization occurs before the timer. The reported value is complete-response latency; first-token timing and concurrent throughput were not measured.

Main Luna calls used explicit cache mode with no marked breakpoint. Every main Luna response reported zero cache reads and zero writes. The main window was 08:42:23–08:46:47 UTC; the cache window was 08:49:58–08:50:56 UTC, on 2 October 2026.

2.4 Metrics and statistical analysis

Binary reference labels, y ∈ {0,1}

Accuracy = mean[1(p ≥ 0.5) = y]
Brier = mean[(p − y)²]
Log loss = −mean[y ln p + (1 − y) ln(1 − p)]

Brier and log loss are lower when predictions better match labels. Probabilities are clipped to [10⁻¹⁵, 1 − 10⁻¹⁵] only for log loss.

Exact event probabilities, q ∈ [0,1]

MAE = mean[|p − q|]
RMSE = √mean[(p − q)²]
Expected Brier = mean[(p − q)² + q(1 − q)]

Excess Brier is mean[(p − q)²]. These targets are known probabilities, so no randomly sampled outcome replaces q.

Latency is summarized by the median, mean and linearly interpolated 95th percentile. Cost is the reported usage.cost in USD. Cost per 1,000 decisions equals measured total cost / measured requests × 1,000. It is a workload projection. Input token counts can differ because the providers use different tokenization and API overhead.

Main contrasts use a paired percentile bootstrap over case identities with 2,000 draws. Quality differences are Luna minus Jev; time contrasts are the ratio of Luna’s median to Jev’s median. Seeds are deterministic per task. Intervals are exploratory, with no multiple-comparison correction. They describe variation across sampled cases in this run, not repeated days or service conditions. Invalid calls would remain in success-rate and cost accounting; every measured response here was valid.

3Main results

3.1 Quality depended on the task

@@QUALITY@@

Both models were correct on all 60 simple policy cases. Luna’s probabilities were exactly 0 or 1 and matched the targets; Jev’s less extreme probabilities produced a nonzero Brier score despite identical binary accuracy. The all-correct result supports this sampled task set, not perfect policy reasoning generally.

Jev agreed with 95.0% of published BoolQ labels, compared with Luna’s 81.7%. The paired difference was −13.33 percentage points for Luna (95% bootstrap interval −21.67 to −5.00). This is published-label agreement: the passage audit in Section 5 found a clear label contradiction and several ambiguous items.

Luna had lower absolute error on exact probability calculations: 3.277 percentage points versus Jev’s 6.443. The paired MAE difference was −3.166 points (95% interval −5.212 to −1.114). The excess-Brier difference interval includes zero, so the evidence for a difference depends on which error metric is used.

@@FIG2@@
Additional probability scores, 60 cases per model. Lower is better.
ModelExcess BrierExpected Brier
Jev0.0099770.156336
Luna0.0055100.151869

3.2 Jev completed requests sooner at lower short-input cost

@@FIG1@@@@LATENCY@@

Pooling timing and cost across the 180 measured requests per model, Jev’s median was 0.238 s and Luna’s was 1.143 s: a 4.81× ratio of medians. The measured workload cost $0.019806 per 1,000 Jev decisions and $0.043876 per 1,000 Luna decisions. Quality is reported separately by task.

3.3 Paired uncertainty

@@INTERVALS@@

The zero-width policy accuracy interval is a consequence of resampling an all-correct sample. It does not establish zero population error. Supplementary Wilson intervals and an independent bootstrap check are included in Appendix D.

4Cached-input experiment

4.1 Controlled prefix reuse

A second experiment used one 1,903-word fictional equipment-dispatch rulebook and 24 varied records. The three arms were Jev, uncached Luna and cached Luna. Both Luna arms used exactly the same text and message segmentation; the cached arm marked the first text block with an explicit cache breakpoint. Both used explicit cache mode, a 30-minute TTL and the same experiment-specific cache key.4

One distinct case primed the cache before the 24 measured cases. Arm order was randomized within each case. The prime reported 2,533 cache-write tokens and zero read tokens. All 24 cached calls then reported exactly 2,533 read tokens and zero writes. Every uncached control reported zero reads and writes. Jev did not expose comparable cache counters, so its cache state remains unknown.

The cached prefix accounted for 60,792 of 69,426 reported prompt tokens across measured cached calls (87.56%). Each request still generated a fresh answer. Cache hits are established from returned counters, rather than inferred from lower latency.

@@CACHE@@@@FIG3@@

4.2 Billing with and without the cold prime

The measured cached Luna calls cost $0.001711320 in total, versus $0.007180200 for uncached Luna and $0.003184104 for Jev. After priming, cached Luna was 76.17% cheaper than uncached Luna and 46.25% cheaper than Jev on this workload.

Amortization over the observed 24 decisions
$0.001711320 measured + $0.000362525 extra prime
= $0.002073845 total
÷ 24 × 1,000 = $0.086410 per 1,000 decisions

With this allocation, cached Luna cost 34.87% less than Jev. The prime itself took 1.204 s; that one-off latency is excluded from the measured steady-state latency distribution.

The observed prime-plus-call model predicts cost parity with Jev after about six reused decisions: $0.000362525 + N × $0.000071305 versus N × $0.000132671. This is an extrapolation using mean observed costs and a fully successful cache, not a measured result for every reuse count. It counts an extra priming request, as this experiment did.

4.3 Speed and quality under caching

Cached Luna’s median was 0.918 s, uncached Luna’s 1.060 s and Jev’s 0.254 s. The ratio of these medians is 0.866 for cached versus uncached Luna and 3.607 for cached Luna versus Jev. The original cache summary separately reports the median of per-case ratios (0.913 and 3.618); those are different statistics.

Both Luna arms answered 21/24 correctly, and Jev answered 20/24 correctly. Cached and uncached Luna disagreed on two individual cases despite equal aggregate accuracy. Jev’s Brier score was lower in this set because its probabilities were less extreme on errors. With one answer per arm per case, the study does not isolate whether caching changed answer quality or whether generation variability explains the differences.

5Reference-label quality

A passage-only audit reviewed all 60 selected BoolQ question–passage–label triples without inspecting model predictions or scores. It began after the scored run started, following identification of a possible source-label error. It is therefore a post-hoc audit, not a preregistered filter.

38 / 60Published labels supported by the passage
21 / 6015 ambiguous or insufficient; 6 time/scope flags
1 / 60Clear passage–label contradiction
Case boolq-dev-1842. The question asks whether carbon is a metal. The published label is “yes”; the supplied passage describes carbon as nonmetallic. Both models answered “no”. The primary scores retain the published label and therefore count both answers as incorrect.

A transparent post-hoc sensitivity check excluding only this contradictory row yields Jev 57/59 (96.61%) and Luna 49/59 (83.05%). No other flagged rows are excluded. Ambiguity flags do not establish that the opposite label is correct.

Cases boolq-dev-0659 and boolq-dev-0383 share the same World Cup host-qualification passage with paraphrased questions. Both remain in the frozen sample. Sampling rows does not guarantee independent topics; row-level intervals may understate uncertainty from topic dependence.

Complete passage audit — all 60 cases
@@AUDIT@@

Read the original audit · Inspect each passage, label and model answer

6Discussion and limits

The results support a task-dependent comparison. Jev combined lower latency with stronger agreement on this BoolQ sample. Luna was more accurate numerically on the sampled probability calculations. Reusing a large prompt prefix reduced Luna’s measured input cost enough to reverse the long-policy price comparison.

For repeated decisions against a stable reference document, the relevant cost includes the prefix length, cache hit rate, number of reuses and any cache-writing charge. For short, mostly unique inputs, the uncached comparison is more informative. Both scenarios also require a task-specific quality target; a faster or cheaper answer is useful only if its error pattern is acceptable for that task.

  1. Configuration scope. Luna’s reasoning was disabled. Higher effort, different prompting, alternative providers or different model versions could change both quality and latency. Matching the output object does not equalize model architecture or API overhead.
  2. Small, narrow samples. There are only 60 cases per main task and 24 cache cases. Synthetic rule families and exact probability puzzles do not establish broad operational reasoning, prospective forecasting skill or real-world probability calibration.
  3. One service window. Timing includes this machine, network, router and provider load. No repeated-day study, concurrency test or throughput benchmark was performed. Random order reduces systematic order effects but cannot remove all transient load effects.
  4. Public benchmark and imperfect labels. BoolQ may overlap with training data. Its labels contain the documented limitations, and two selected rows share a passage. Primary results should be read as agreement with published annotations.
  5. Uncertainty scope. Bootstrap intervals resample cases, not latent task families or independent model runs. All-correct samples can produce zero-width bootstrap intervals. Comparisons are exploratory and unadjusted for multiple testing.
  6. Probability semantics. A generated numeric probability and Jev’s native probability are different mechanisms. Brier scores reflect both probability reliability and discrimination; this pilot cannot establish general calibration.
  7. Cache generalization. All measured cached requests hit within a short run. Results do not estimate expiration, cold traffic, low reuse, cache eviction or concurrent-request behavior. Quality differences between cache arms are not attributable to caching from this design alone.
  8. Billing scope. Costs are reported per-response charges; no independent account invoice was reconciled. Future rates and tokenization may differ. Cost per 1,000 scales this sample rather than promising a market price.

6.1 Conclusion

For the recorded settings, Jev offered the lower latency in every tested task. Quality had no universal winner: Jev led on published-label reading agreement, while Luna led on mean absolute probability error. Verified prefix caching made Luna cheaper on the long-rulebook workload, including the observed prime at 24 reuses. A deployment choice should preserve this separation between task quality, service latency and workload-specific cost.

7Reproducibility and accounting

The archive contains frozen cases, source attribution, every request and response, timing and token fields, reported charges, request IDs, model IDs, code, tests and verification results. All 441 recorded calls returned schema-valid probabilities; 432 belong to measured experiments. The total reported API charge was $@@TOTAL@@ (about 2.41 US cents).

@@COSTLEDGER@@

Thirty-four offline tests passed. Completed-run verification checked frozen hashes, input and target identity, expected request counts, served models, reported billing, zero Luna reasoning tokens, and both cache controls. These checks establish record consistency; they do not independently validate upstream model identity or human source labels.

All completed-run verification checks@@CHECKS@@

7.1 Reproduce the analysis

From the extracted benchmark directory, the following commands inspect the saved evidence or rebuild the report without calling a model:

PYTHONDONTWRITEBYTECODE=1 python3 -m unittest discover -p 'test_*.py' -v
python3 verify_run.py --help
python3 report.py --help
PYTHONDONTWRITEBYTECODE=1 python3 build_paper.py

The figure script requires Matplotlib; the benchmark runner and report builder use Python’s standard library. The frozen case files are sufficient to repeat the benchmark. A fresh model run uses the secure environment or stdin key mechanism documented in the README and incurs new charges.

Rerun instructions · Offline test output · Source manifest · Complete archive

References

  1. OpenRouter. Jev tutorial and typed decisions; Decisions API reference. Used to implement the native Noul answer contract.
  2. OpenRouter. Jev 1.13 model page; GPT-6 Luna model page. Pricing context captured on 2 October 2026; primary cost evidence comes from saved response usage.
  3. OpenAI. GPT-6 Luna model documentation. Configuration context for the evaluated model.
  4. OpenRouter. Prompt caching. Explicit cache controls and reported usage fields. See also OpenAI prompt caching documentation.
  5. Clark C, Lee K, Chang M-W, Kwiatkowski T, Collins M, Toutanova K. BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions. NAACL, 2019.
  6. Google Research. BoolQ dataset repository; Google-owned Hugging Face mirror. Sample drawn from the named validation split, 3,270 rows.
  7. Creative Commons. CC BY-SA 3.0. License retained for selected BoolQ source passages and annotations. Per-case source fields record normalization and reformulation.

AExact prompts and schema

The following are the actual system instruction and schema used by the Luna adapter. Each case supplies its own state, question and criteria. The recorded request bodies retain all settings and text.

Luna system instruction

@@SYSTEM@@
Strict JSON Schema
@@SCHEMA@@
@@PAYLOADS@@
Cache settings and explicit breakpoint
@@CACHECONTROL@@

The uncached arm omits the prefix breakpoint. Other settings, text, segmentation and the experiment key are shared.

Complete 1,903-word dispatch rulebook
@@RULEBOOK@@

BEvery measured test case

Inspect all 204 cases, including the complete input, reference answer, provenance and model outputs. Filters affect this appendix only; the paper’s tables remain the full frozen study.

Open a case to inspect its full evidence.Download all case data

CEvery recorded API call

All 441 calls are retained: measured requests, warmups, cache priming and setup probes. Expand a row for the exact request, API response, tokens, bill and timing. The saved records omit authorization headers.

Export all 441 request summaries (CSV) · Export full request and response records (JSON)

DSupporting evidence and downloads

Token counts by task and model@@TOKENS@@
List-price context at evaluation time
Table 6. Captured base rates per million tokens in USD, 2 October 2026. All tested inputs are below the catalog’s long-context pricing threshold. Actual recorded bills control all reported costs.
ModelInputOutputCached input readCache write
Jev$0.042$0.000UnspecifiedUnspecified
Luna$0.100$0.500$0.010$0.125

No search calls or tools were used. The captured catalog record also preserves long-context rate overrides, which were not applicable to these inputs.

Source manifest, sampling indices and hashes
@@MANIFEST@@
@@AUXSTATS@@
Complete artifact index@@ASSETS@@
Download complete study (.zip)All call summaries (.csv)Complete evidence (.json)