Decision quality, latency and cost:
GPT-6 Luna and Jev
Full validation benchmarks, exact-oracle tasks, repeated timing and independently primed cache blocks.
Abstract
Objective. Compare a general language model constrained to compact probability outputs with a dedicated decision model on quality, latency and billed inference cost, including the effect of reusing a cached input prefix.
Methods. We evaluated the complete labeled validation splits of BoolQ (3,270), RTE (277) and WiC (638), plus 300 synthetic policy and 300 exact-probability cases. The quality protocol used 9,570 requests. A sequential timing protocol repeated 60 selected cases three times per model. A four-block cache protocol evaluated 96 rulebook cases, with one cold prime per block, cached and uncached Luna arms, and Jev. The primary datasets and analysis specification were fixed before primary outcomes were inspected.
Results. @@ABSTRACTRESULTS@@
Interpretation. @@ABSTRACTINTERPRETATION@@ Public validation data, one measurement day, task adaptation and shared synthetic templates still limit generalization. The study provides a reproducible service comparison, without claiming an official SuperGLUE leaderboard result or an overall model ranking.
1Study design
The study asks three practical questions: whether the models make equally good compact decisions, how long those decisions take, and what they cost when inputs are unique or repeatedly shared. Both models answer with a probability of “yes”, from which the same fixed binary decision threshold is applied.
One study combines three complete public validation benchmarks, exact-oracle tasks, repeated response timing and controlled prefix caching. Each protocol answers a different part of the same quality, speed and cost comparison. Supplementary measurements and setup checks are retained in the evidence appendix and complete call ledger.
| Protocol | Cases and repeats | Measured requests | Purpose |
|---|---|---|---|
| Bulk quality | 4,785 cases; one answer/model/case | 9,570 | Quality and actual cost. Up to 16 case pairs processed concurrently. |
| Isolated timing | 60 cases; 3 repeats/model | 360 | Sequential latency and within-case variation. 12 warmups excluded. |
| Long-prefix caching | 96 cases; 3 arms; 4 blocks | 288 | Cache control, quality, latency and amortized cost. 4 extra primes accounted separately. |
Concurrency in the large quality run keeps evaluation time practical. Its timings are retained as diagnostics and excluded from claims about relative response speed. Timing and cache runs execute after bulk traffic finishes.
2Methods
2.1 Models and matched outputs
Jev was requested as typesafe/jev-1.13 through OpenRouter’s Decisions endpoint. Luna was requested as openai/gpt-6-luna through Chat Completions, pinned to OpenAI with provider fallbacks disabled. @@MODELIDENTITY@@
{"type":"noul","noul":0.73}Here p = P(yes), and P(no) = 1 − p. “Yes” is predicted when p ≥ 0.5. The value 0.73 illustrates the format.Luna used strict JSON Schema, at most 128 output tokens, reasoning.effort = none, no tools and non-streaming output. Temperature was unspecified. Both models received the same state, question and yes/no criteria. Luna also received a format instruction and schema; Jev used its native typed answer. The API envelopes and tokenization differ. No reference answer, source label, case ID or provenance field was sent.
Jev’s native Noul probabilities and Luna’s generated numeric probabilities are compared as returned. Neither adapter extracts token log probabilities. Exact system text, payloads and schema are included in Appendix A.
@@OUTPUTAUDIT@@2.2 Established benchmark datasets
@@DATASETTABLE@@BoolQ. The complete 3,270-row development split was read from the normalized Google-owned Hugging Face mirror used for this evaluation. The original Google Storage URL had returned HTTP 403. Source labels are retained exactly, including the documented contradiction in case 1842. This is a full-split result rather than a selected sample.5
RTE. Recognizing Textual Entailment asks whether a premise supports a hypothesis. Non-entailment includes both contradictions and unsupported hypotheses. We used all 277 validation rows from the accessible aps/super_glue Hugging Face dataset. The published class ordering is 0 = entailment, 1 = not entailment; the probability-of-yes target is therefore 1 − label.6
WiC. Word-in-Context asks whether two occurrences of a word share the same sense. All 638 validation rows were used. The published target maps directly to no = 0 and yes = 1. The target spans were checked against source offsets and marked with [TARGET]…[/TARGET]; inflected surface forms were retained.7
These are binary adaptations of established tasks using the shared probability-output protocol. They are not the complete SuperGLUE suite, hidden test splits or official leaderboard submissions. No training or fine-tuning was performed. Public validation sets may have appeared in either model’s training data.
2.3 Synthetic targets and test coverage
The 300 simple-policy cases contain 100 each for refunds, delivery and laboratory access, balanced to 50 yes and 50 no per family. Each family includes six paired condition contrasts before seeded stratified sampling fills the set. Exact Boolean rules determine targets. Thirteen policy inputs share decision-relevant facts with supplementary cases; the templates are shared.
The 300 probability cases contain 100 conditional-table problems, 100 without-replacement draws and 100 weighted mixtures. Every target is an exact rational event probability, independently checked by outcome enumeration. Probability inputs identical to supplementary cases were excluded. The cache protocol includes 48 counterfactual pairs: an approved case and a case with exactly one changed fact that violates a named rule. Expected labels are checked by two independently structured rule implementations.
The dataset manifest records actual derived seeds, row mappings, input hashes and overlap flags. The base seed is 20261002; timing case selection uses 20261004 and timing order uses 20261005. The primary dataset, timing sample and analysis specification were fixed before primary outcomes were inspected. The chronology and supplementary measurements are documented in Section 8.
2.4 Request schedule and measurement
The quality run shuffles cases and model order within each pair. Sixteen workers each execute pairs sequentially over their own persistent HTTPS connections. There are no automatic retries or answer repairs. Failed calls remain in the logs, costs and success accounting; a circuit breaker stops sustained failures or unknown billing.
The isolated timing sample contains 12 cases from each of the five tasks, chosen before primary outcomes were inspected. The same 60 cases are evaluated in three sequential randomized blocks. Two warmup cases per model per block are excluded from timing estimates but retained in cost accounting. The cache protocol uses a separate sequential connection and independently marked prefixes.
A monotonic timer starts immediately before connection setup/request and ends after the complete body is received, parsed and validated. Payload serialization occurs before the timer. Latency includes the local network, OpenRouter, provider processing and client parsing. It is not time to first token. No concurrent throughput or multi-day reliability claim is made.
Ordinary Luna requests use explicit cache mode without a breakpoint. Returned cache-read, cache-write and reasoning-token counters are checked. API bills use the recorded usage.cost field. The study budget was $1. Before each subsequent protocol, known costs plus a $0.01 allowance for the unreported charge were checked against that budget. Each runner also uses a local cost guard; this is not an account-level spending cap.
2.5 Metrics and uncertainty
Binary targets y ∈ {0,1}
Brier = mean[(p − y)²]
Log loss = −mean[y ln p + (1 − y) ln(1 − p)]
Invalid responses count as wrong for primary accuracy. Balanced accuracy averages recall for yes and no. Brier and log loss use valid responses, with denominators shown. Log-loss clipping is 10⁻¹⁵.
Exact probabilities q ∈ [0,1]
RMSE = √mean[(p − q)²]
Expected Brier = mean[(p − q)² + q(1 − q)]
Excess Brier is the first squared-error term. A known event probability is not replaced with a sampled binary outcome. Errors in percentage points equal the probability-scale error × 100.
Quality contrasts are Luna minus Jev. Paired percentile intervals use 5,000 resamples. Exact duplicate BoolQ passages are resampled together as clusters; other main tasks are resampled by case. Wilson intervals describe model accuracy. All intervals are nominal, exploratory and unadjusted for multiple comparisons. Public fixed-split scores are exact descriptions of this run; these intervals rely on a hypothetical broader case population and do not measure model-training uncertainty.
For timing, each case/model is first reduced to its median across three repeats. We then compare medians across cases and bootstrap paired case identities. Repeat-block summaries are also shown. Cache inference groups the two counterfactual cases by scenario; four blocks still represent one session, not independent days. Cost per 1,000 is observed billed cost / attempted requests × 1,000. Missing bills are never treated as zero.
3Benchmark results
3.1 Decision quality
@@QUALITYTABLE@@@@QUALITYFIG@@@@QUALITYTEXT@@3.2 Exact probability estimation
@@PROBTABLE@@@@PROBFIG@@@@PROBTEXT@@3.3 Paired contrasts and uncertainty
@@INTERVALTABLE@@Accuracy contrasts retain invalid answers as wrong. Proper-score contrasts use pairs valid for both models; excluded pairs are reported in the analysis files. Positive accuracy differences favour Luna; negative error-score differences favour Luna.
3.4 Billed cost by workload
@@COSTTABLE@@These costs cover fresh quality calls, with Luna caching disabled; Jev’s cache state is unobserved. The bulk-run latency is available in the raw records but is not used for speed comparisons. Model-native token counts and format overhead differ; the bill measures the service cost of obtaining one answer.
4Repeated, isolated timing
The following results come from the separate serial run: 60 cases, three repeats per model, 360 measured responses. The primary summary first takes a within-case median, then the median across cases. No quality-run requests run concurrently with this protocol.
@@TIMINGTABLE@@@@TIMINGFIG@@@@TIMINGTEXT@@All three timing blocks
@@TIMINGBLOCKS@@Within-case prediction consistency across repeats
@@CONSISTENCY@@This is descriptive run-to-run answer variation on the 60 timing cases. It is not an independent larger quality dataset and is not pooled into the main quality score.
5Four-block prompt-caching experiment
5.1 Cold writes, cache hits and controls
The same 1,903-word fictional dispatch rulebook is used in four sequential blocks, each with a fresh administrative identifier inside the prefix. The identifier changes no rule or fact and appears in every arm. Each block contains 24 cases and one extra Luna cache-prime call. Random arm order compares Jev, Luna uncached and Luna cached on every case.
Both Luna arms have identical text, segmentation, model settings, explicit cache mode, 30-minute TTL and a shared block key. Only the cached arm marks the prefix breakpoint. Distinct prefix text, rather than a new key alone, isolates each cold prime. Returned counters establish cache treatment.4
@@CACHEVERIFICATION@@@@CACHEBLOCKTABLE@@5.2 Quality, latency and cost
@@CACHETABLE@@@@CACHEFIG@@@@CACHETEXT@@Paired cache uncertainty and counterfactual consistency
@@CACHESTATS@@Cache savings are conditional on the long common prefix and the observed hit rate. The experiment does not simulate long idle periods, expiry, eviction, low reuse or concurrent cache traffic. Same aggregate accuracy does not imply identical answers, and answer differences alone cannot be attributed to caching.
6Data quality, overlap and sensitivities
The published benchmark labels are preserved. Data checks verify source fidelity and transformations; they cannot establish that every human annotation is correct. BoolQ, RTE and WiC scores should be interpreted as agreement with their respective published labels.
A passage-only BoolQ audit covered 60 cases: 38 supported, 15 ambiguous or insufficient, six time/scope limitations and one clear contradiction. The audit began after supplementary measurements started and did not inspect model predictions. Those findings do not establish a label-error rate for the full 3,270-row split.
Sixty BoolQ rows also occur in the supplementary measurements. All 3,270 primary benchmark predictions were obtained from new calls. The subgroup table shows the 60 overlapping rows and the remaining 3,210 rows to make this overlap transparent; no rows were excluded based on performance.
@@OVERLAPTABLE@@Duplicate passages within BoolQ are grouped in the paired bootstrap. The exact cluster count is @@CLUSTERS@@. Synthetic cases still share a small number of rule families and probability templates; larger counts do not create new domains. RTE and WiC were not subject to a complete independent human label audit.
The complete 60-row label audit
@@AUDIT@@7Discussion and limits
@@DISCUSSION@@7.1 Design strengths
- Full labeled public splits reduce sensitivity to selecting a small, favourable sample.
- Frozen target mappings, exact-oracle tests, request hashes and saved outputs make the comparison inspectable and repeatable.
- Isolated repeated timing avoids using concurrent bulk-request times as a latency benchmark.
- Fresh cache writes in four blocks test whether the discount is reproducible within the session.
- Counterfactual policy pairs test whether a model changes its decision when one controlling fact changes.
- Duplicate-passage clustering, paired comparisons, validity counts and complete cost accounting make uncertainty and failures visible.
7.2 Remaining limits
- Model configuration. This compares Luna with reasoning disabled and the recorded Jev service. Other reasoning budgets, prompts, providers or versions can alter the trade-off. Matching output schema does not equalize internal processing.
- Public validation data. Training contamination is unknown. Published labels are fallible. Results are not official hidden-test or leaderboard scores.
- One day and environment. Repeated cases and cache blocks improve within-session evidence but do not measure variation across days, regions, accounts or deployments. Network and service load are part of observed latency.
- Limited task diversity. Five task types and one fictional long rulebook do not represent every application. Synthetic rows share templates; probability arithmetic is not real-world forecasting calibration.
- Inference assumptions. Confidence intervals are exploratory, unadjusted and conditional on the chosen resampling units. Clustering exact BoolQ passages does not resolve all topical dependence. Narrow intervals do not remove dataset or model-selection bias.
- Cache conditions. Savings depend on actual hit rate, prefix length, write charges, reuse and expiry. Fresh answers can vary despite identical text. Four blocks do not establish long-run cache availability.
- Billing. Reported charges are reconciled against saved usage, not an independent invoice. A per-1,000 projection describes these inputs, not a universal future price.
7.3 Conclusion
@@CONCLUSION@@8Reproducibility and complete accounting
@@ACCOUNTING@@The archive retains the frozen datasets and analysis specification, every API payload and response, generation IDs, measured times, token usage, billing, checks and plotting/report code. Supplementary measurements were available before the primary protocols were finalized. They remain in the evidence appendix and call ledger; the primary result tables use only their stated protocol denominators. Authorization headers and API keys are excluded.
@@LEDGER@@Completed-run verification and offline tests
@@VERIFICATION@@8.1 Rebuild without model calls
python3 expanded/analyze.py --help python3 expanded/build_report.py python3 -m unittest discover -s expanded -p 'test_*.py' -v
The analysis/report code uses saved evidence. Figure generation requires the plotting packages documented in the README. A fresh model evaluation is a separate paid run using an environment or stdin credential, never a key embedded in source or command arguments.
Study protocol and reproduction · Frozen analysis specification · Dataset manifest
References
- OpenRouter. Jev typed-decision tutorial and Decisions API reference. Native Noul schema and request adapter.
- OpenRouter. Jev 1.13 and GPT-6 Luna model pages. Contextual rates captured on 2 October 2026; measured charges use response usage.
- OpenAI. GPT-6 Luna documentation. Model setting context.
- OpenRouter. Prompt caching controls. Explicit breakpoints and usage counters.
- Clark C, Lee K, Chang M-W, Kwiatkowski T, Collins M, Toutanova K. BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions. NAACL, 2019. Repository; Google-owned dataset mirror. CC BY-SA 3.0.
- SuperGLUE. Accessible Hugging Face dataset and RTE task documentation. Source class mappings, fields and validation counts. Underlying dataset terms apply; no unified permissive RTE license was verified.
- Pilehvar MT, Camacho-Collados J. WiC: the Word-in-Context dataset. NAACL, 2019. Task definition and author-stated CC BY-NC 4.0 license; source spans from the SuperGLUE mirror.
This report is served locally. Source attribution and individual dataset terms remain attached to the data; the bundle does not impose a new license over third-party benchmark material.
AExact prompts and schema
Luna system instruction
@@SYSTEM@@
Strict output JSON Schema
@@SCHEMA@@
Complete shared dispatch rulebook
@@RULEBOOK@@
Frozen dataset manifest
@@MANIFEST@@
DSupporting evidence
Captured pricing and token accounting
@@TOKENSTABLE@@@@PRICING@@
Complete analysis results (JSON)
@@SUMMARY@@