Numbered source-unit selection experiment
The literal-quote development trial failed before independent review: models selected partial structural units or changed source spelling. This experiment presents complete numbered units and asks for existing range references. Code derives the exact original quotation, then applies the same strict provenance, boundary, uniqueness and quote-budget checks. The failed literal adapter and its results remain preserved.
Representation and validation
activityEvidenceRanges exposes the exact previous segmentation and conservative merging rules as a generator. It keeps conditional newlines and abbreviated names within their existing complete units. activityEvidenceBoundaries now uses this generator; no source normalization or boundary-policy change occurs.
buildActivityUnitPresentation retains every source field and its provenance. Unit text concatenates to the exact original field. The presentation includes the original packet hash, exact baseline context and runtime ICU version, with a canonical hash over its content. It permits at most 2,048 units and 160 KiB including framing and hash. Exceeding either rejects the whole presentation. Units exceeding the 8,192-byte quotation ceiling remain visible whole, with an unknown token count; smaller units carry their exact embedding-model token count, including counts above 256. Provider input tokens are a different measurement.
selectActivityEvidenceByUnits accepts at most six references with inclusive endpoints within one source. It validates the complete selection and derives contiguous quotations without rewriting or shortening them. It validates the concatenation's actual tokens rather than summing individual hints. Duplicate, ambiguous, unknown, blank, oversized or malformed selections reject the response. The existing independent reviewer receives the original full packet and canonical quotations, without proposer reasoning.
The adapter uses the same mandatory operator journal and platform evidence-proposal cost policy. Static ownership checks permit its transport call and only the proposal version constant/type import from the older adapter. Neither experimental adapter has a serving or worker caller. This change adds no provider, dependency, subscription, database write or production flag.
Same-world development coverage
All 46 packets from the same twelve development worlds fit the numbered presentation: 9,670 complete units, with at most 855 units and 126,773 bytes per packet. Total serialized size is 1,813,947 bytes. The probe reconstructs every source exactly and preserves three prior v1 artifact hashes and the prior late-source v2 artifact hash.
Of those units, 9,569 are nonblank and fit the individual 256-token quotation budget; 100 exceed that token budget and one is blank. No unit in this sample exceeds 8,192 bytes. Seventy-seven unit texts repeat within their source; the existing uniqueness check still applies to the selected complete range. These counts describe mechanical eligibility, not semantic usefulness or model accuracy.
The preserved development probe SHA-256 is 1c3f7c951e104b5d099e5e55cfdb53570e201f3a5cf65086286a8164c61d4ea8. It took about 1.2 seconds for presentation processing on this local run; this is not a production capacity benchmark. No private reader data or evaluation cohort content was used.
Checks and next test
The selected integrated suite passes 509 tests with no failures or skips. The two range tests first failed against a single-whole-field stub; the presentation's initial binding test and the model selection tests were also developed red-to-green. Tests include exact byte/count boundaries, complete reconstruction, multibyte and conditional text, canonical quote derivation, independent-review handoff, journaling and cancellation. Full workspace typecheck passes all eight tasks and build all five. Independent review found no remaining issues and additionally checked boundary parity against commit ca2803a67 across 169 cases, plus selection of unit 2,047. Final-commit CI is separate; both workflows for the preceding ca2803a67 commit passed.
A frozen development protocol used the same four deliberately chosen EN/ZH/ES/JA packets as the failed literal trial, preserving their packet and numbered-presentation hashes. It permitted eight requests, no retries, fifteen minutes and the same observed-cost/unknown-cost stop rules. The plan SHA-256 is 11661c4fcbe269165404b68b021a2b3113b03a94244c08c88eebd5582ddf2870.
Real-model results and next reliability change
The trial was mixed and does not support promotion. English returned one valid proposal and an independently supported practice/trust/seat consequence. Spanish returned six mechanically valid proposals; the reviewer rejected all six as NPC/faction profiles or configuration, leaving no aspect. Chinese and Japanese again failed selection validation, so neither reached review. Six actual requests cost USD 0.1682935, with complete journals and no unknown costs. Result SHA-256: ce5d09d7116572802047306b9f493df296702262d1391cd623a98f0e25da06e2.
Two separately bounded diagnostic proposals isolated the remaining validation failures. Chinese produced four valid ranges plus two of 300 and 377 tokens; Japanese produced five valid ranges plus one of 379 tokens. All references, envelopes and presentation hashes were valid. Because the current adapter rejects an entire response when any quote exceeds 256 tokens, it discarded nine mechanically valid ranges along with three oversized ones. These new diagnostic calls cost USD 0.0306315, with complete journals and no unknown costs; result SHA-256 831e91762f8f7138a510a0e5611c2a4ad3c2eb48d25f438a6f0286a68b659d6f. Only fixed-label diagnostics and bounded references into known public sources were persisted, never raw provider strings.
The next proposed change is explicit per-selection abstention for quotation-budget overflow: discard an oversized selection whole, retain other strictly validated selections for independent review, and record the original accepted/rejected indices and reasons. Do not shorten a quote, relax the budget or accept malformed references. This would address the measured all-or-nothing loss without treating mechanically valid text as semantically supported. Spanish's entirely unsupported selections remain a separate semantic-selector concern.
Even a successful trial would establish only bounded pipeline behavior on selected development examples. Semantic support remains fallible model judgment; a fresh preregistered ranking comparison and measured reader outcomes remain required before promoting the experimental ranking.
