Reader activity semantic contrast
The curation replay recovered mechanically valid quotations, but independent review still accepted a disputed Japanese format-only passage. The selector also returned NPC-only actions and narrator directives. This experiment separates that semantic issue from quote provenance and budget reliability.
Frozen contract and inputs
The clarification requires a concrete fictional action with consequence, or a specific interaction involving the reader or a clearly established player-controlled role. A cast list, co-presence, group-chat format, distinctive voices or a generic promise to react to input does not establish this narrower activity aspect. Fictional choice/consequence rules remain eligible when written as instructions; markup or imperative wording alone is not a reason to reject them. Premise and format information remain available in the unchanged baseline.
The authored fixtures contain eight paired scenarios in EN, ZH-Hans, ZH-Hant, JA and ES: reader training, matched NPC training, format-only co-presence, concrete group rehearsal, a generic narrator directive, an instruction-shaped fictional rule, a denied action and uncertain participant identity. Their intended decisions are three supported, four unsupported and one ambiguous per language. Independent agent review found a translation drift from unlocking a gate to opening it; that was corrected before freezing. This was agent review, not independent bilingual human adjudication.
All forty quotations are complete single structural units below the unchanged quote limits. Each language's selector sees all eight scenarios in one packet. A separate reviewer judges all eight original complete quotations in two groups (six plus two), regardless of selector output. Expected labels, case names and model selections are not supplied to that reviewer. Selector coverage and reviewer agreement therefore have separate denominators. Fixture file SHA-256: bcd00c97c8a824f9cbee6a8e515d4fe326141c09a6b26685cb855273265c152b.
The protocol freezes both old source hashes and the exact predicted source hashes after adding one predefined clarification to the selector and reviewer rubrics. It changes no model, transport, schema, source selection, segmentation, quote budget or scoring rule. Each arm permits five Gemini selectors and ten Claude fixed-quote reviews, fifteen sequential requests total, with no retries, a twenty-minute timeout and USD 3 observed-cost next-request stop. Unknown cost stops subsequent spend. A private durable journal records actual usage before validation; no raw provider responses or credentials are retained. Protocol SHA-256: f36eac06f12bd0546755be3bfd7745be72a9dc2560cb73b681d37afe294693d8.
Before the clarified arm, helper review found that the planned edit targeted a rubric shared with the unused literal proposer. That would also change a third prompt. The amendment instead inserts the identical clarification into the reviewer-specific prompt, leaving the failed literal proposer intact. V1 plan, helpers and completed old observations remain preserved. Amended v2 protocol SHA-256: f50db2647e59fd0d5c8bcff52a0cdba90047236d7155ebaab9b267eeb70a1053. Case text, intended labels, models, limits and clarification text are unchanged; the amendment binds the old plan and result hashes. No clarified-arm calls occurred before this correction.
The advancement check requires all three intended positives per language to be selected and supported, and no intended negative or uncertain case to be selected or supported. Exact ambiguous-versus-unsupported disagreement is reported separately from safe abstention. Failures and omissions remain explicit; an incomplete run cannot pass.
Existing-prompt result
The old arm completed all fifteen requests, cost USD 0.1517515, and recorded no unknown costs with complete journals. Its selector retained all fifteen intended positive cases but additionally selected NPC-only training in ZH-Hant and the explicitly denied action in Japanese. Fixed-quote review supported all fifteen intended positives and supported none of the twenty-five negative/uncertain cases. Three uncertain-role cases received unsupported instead of the authored ambiguous label; both verdicts abstain. Exact verdict agreement was 37/40; translated variants are not independent observations. Result SHA-256: 152c2be73d7f57b8179315ede10ddddf8f058a22e6dc9c600f8bbc93bff06d2e.
The existing reviewer therefore already passes the supported-versus-abstained boundary on these simplified cases. The observed old-arm problem is two extra selector choices; no claim of improved reviewer accuracy can follow merely from matching this result.
A separately frozen natural-development followup uses the exact prior EN, ES and JA packets: an English positive control, the Spanish selector's unsupported outputs, and the disputed Japanese format result. It permits at most six requests, ten minutes, no retries and USD 3 observed-cost/unknown-cost stopping. It preserves empty/unsupported outcomes and requires inspection of any supported passage. The original followup protocol was frozen before the old contrast completed; its amended v2 only updates the prompt/helper hash chain, with the same chosen packets and rules. V2 SHA-256: 5341f57df24d639f8deb43504cafa14137ab9a5de7b7bc85428e9115d0cbe778.
Clarified-prompt result
The clarified arm completed all fifteen requests, cost USD 0.1557675, and recorded no unknown costs with complete journals. It met the predeclared development boundary in every language: all three intended activities selected and supported, with no negative/uncertain case selected or supported. Both arms together cost USD 0.307519. Clarified result SHA-256: 32ae62440823632b8c0ff5c6cb4096530b6b9ad219898995e9a31135ffdccfc2.
| Language | Intended activities selected, old/new | Extra cases selected, old/new |
|---|---|---|
| EN | 3 / 3 | 0 / 0 |
| ZH-Hans | 3 / 3 | 0 / 0 |
| ZH-Hant | 3 / 3 | 1 / 0 |
| JA | 3 / 3 | 1 / 0 |
| ES | 3 / 3 | 0 / 0 |
The reviewer supported the same fifteen intended positive cases and none of the twenty-five negative/uncertain cases in both arms. Four uncertain-role cases received unsupported in the clarified arm, versus three in the old arm; exact three-way agreement therefore changed from 37/40 to 36/40. Both labels abstain, as predeclared. This result supports the narrower selector contract on these authored cases; it does not establish improved reviewer accuracy, statistical significance, natural-world extraction recall or ranking quality.
After the prompt change, the 49 focused model/curation/billing tests and the complete 522-test selected regression suite pass. Full workspace typecheck passes eight tasks and build five. The diff only adds the predefined wording to the unit selector and reviewer-specific prompts; the old literal-proposer prompt is unchanged. Final-head CI follows the commit.
Natural-development followup
Five requests completed for USD 0.153851, with no unknown costs and complete journals. EN proposed six quotes: two were supported and four unsupported. ES proposed four: two supported and two unsupported. JA proposed none, so no review request or activity aspect was produced; its baseline remains intact. Result SHA-256: f4125d78dba870c64915e9d28680c5b6c6f341be0fde76e01d7f2224f9d15b1e.
The supported English quotes describe a coach directing the reader to decide whether to hold or rotate, and a separate opening assigning the reader the shotcaller role. The supported Spanish quotes explicitly connect defeating a ranked expert in public to reputation, and acquiring a top-ranked weapon to the ability to challenge stronger opponents. Main inspected all four passages in their complete original fields. The English openings remain separate bundles; their simultaneous availability does not mean they occur together. The earlier Japanese format-only judgment did not recur because no quote was selected. That is not proof that the world has no relevant activity or that the reviewer would always reject such text.
The same inspection raises a false-negative concern: the English reviewer rejected “bring the analyst a new pattern” and its draft-influence consequence because it was hypothetical. The rubric explicitly permits conditional fictional mechanics. This is an apparent over-rejection, preserved as a development disagreement; it was not silently changed into supported evidence. The supported passages are plausible developer-assessed evidence, not human-certified labels. The natural comparison is one stochastic run on already-used packets, not a causal or statistically reliable quality estimate.
Helper review also found a reproducibility gap after this natural run had completed: the runner and artifact-builder code hashes were recorded and checked for changes during execution, but were not pinned in its frozen plan before startup. Their values were printed by preflight before calls and recorded in the result; that is weaker than an enforced pre-start manifest. The completed observation remains explicitly qualified and was not rerun or retroactively sealed. A prospective v2 runner now requires a separate execution manifest before credentials or network access; its dry check failed without that manifest and passed with it. Manifest SHA-256: 8f0634e45ecd517c757d4ed9c76d8546c53fd65acd5391afe8068581f5e2ac83. No v2 provider run occurred. Future paid jobs must freeze every executable/artifact dependency before dispatch.
The next useful test is complete twelve-world development extraction across all 46 packets, with explicit whole-world review/selection bounds and frozen executable hashes. That must measure omissions, duplicate/alternative-opening treatment and supported-evidence coverage before a fresh ranking comparison. The current four-packet and three-packet probes cannot establish catalog-wide coverage. Neither this clarification nor the passing authored contrast enables production semantic ranking.
Interpretation boundary
These are authored development propositions, not natural-data accuracy labels, human ground truth or ranking relevance judgments. Five translations of each scenario are correlated; there are eight scenario groups, not forty independent examples. The clarification is designed around the observed development failure, so passing these cases cannot establish generalization. A separately declared natural-development check, complete cross-packet selection, fresh ranking evaluation and measured reader outcomes remain necessary before promotion. All earlier failed or disputed results stay preserved.
