Skip to content

Story representation diagnosis and next candidate

Retain published-greetings-v1. Develop one separate activity-and-consequence aspect only after the parent's performance validation. This report diagnoses information loss and scoring behavior; it establishes no quality uplift.

The consumed sixty-family diagnostic reported macro NDCG@5 of 0.6426 for the published-greetings baseline, 0.5385 for rich content, and 0.6183 for rich content plus reviewed experience with equal RRF-60. The fusion improved EN/JA relative to baseline but regressed ZH/ES. Neither weights nor promotion decisions may be tuned against that consumed sample. The protocol remains the record of that experiment, including its AI-label and correlated-translation limitations.

Evidence boundaries

  • Code-proven: deterministic selection, bounds, serialization, and ranking arithmetic described below.
  • Development-observed: replay of content-selection logic over all twelve family-distinct public development inputs, three per language, with input hashes checked and the existing blind references/reviews inspected.
  • Unverified: whether these mechanisms caused the sixty-family score gap, whether another representation improves retrieval, and any reader benefit.

Development evidence is in the ignored semantic-development-corpus.json, semantic-development-reference.json, semantic-development-reference-v2-blind.json, semantic-development-adjudication.json, and original/curated results under .local-artifacts/. These are bounded inputs and AI assessments, not human ground truth. The saved corpus lacks raw schemas and exact baseline inputs; it cannot establish all first-pass omissions or reconstruct an exact development baseline comparison. The frozen twenty-four-family corpus was not read or used. Inspection made no provider/network calls or production changes. Existing artifacts remain unchanged; this report is the only new file.

Construction and fidelity

Code-proven. The baseline builder reserves up to 3,000 UTF-8 bytes for description and 3,000 bytes/2,000 UTF-16 units for opening. The rich builder first bounds source input to 32,000 bytes, using separate reservoirs and equal serialized-byte shares among selected modern entries. It then compresses again in buildStoryEmbeddingTexts: at most 800 bytes per section, with total metadata/greeting/character/lore/mechanics budgets of 1,200/1,700/1,600/1,900/1,500 bytes. This second pass is sequential; later sections can disappear despite their reserved first-pass shares.

Unmarked system/custom entries can enter lore; marked presets enter mechanics. The second pass merges instructions, variables, rules, and reactions into one mechanics allowance. Section headings and markup consume text budgets. Validated player_role, setting, and tone facets do not protect those facts in the content embedding: the descriptor affects only the separate experience text.

Development-observed. Examples below compare saved bounded source text with the rich embedding text, not with an independently rebuilt baseline.

Development storyRetained versus omitted evidence
Ninevolt: Spring Bootcamp (EN)Description 965→799 bytes: practice-sub role survives; earning trust and potentially losing the seat disappear. Third opening retains only 64 bytes.
无敌!童话王! (ZH)Opening 1,575→798 bytes: scenery survives; the block puzzle, three-heart quest, stakes, and initial choice disappear. Both reaction sections are omitted.
KATSEYE シミュレーター (JA)Eight opening excerpts become three intact excerpts, seven bytes of the fourth, and no remaining excerpts. Description truncation removes mutual-understanding context.
El Ocaso del Vasallo Leal (ES)Player identity and family background survive; the later repeated-training→rapid-growth consequence disappears.
Nueve Provincias: Longyuan (ES)Martial-level lore, weapon, and gold sections disappear; reputation is cut through its action/consequence explanation.

Only 37/77 supported reference facets retain at least one complete cited quote in rich content. Of 121 supporting quotes, 45 survive completely, 11 are cut, and 65 begin outside retained text. This is literal quote retention, not semantic recall: other retained text might convey the same proposition. Conversely, none of the 40 accepted development experience facets exceeded the 450-byte serialization allowance. Experience-quote clipping is a prospective contract risk, not an observed explanation here.

Why can more data reduce fidelity? More source fields do not imply more retained task-relevant information. Under these budgets, extra biographies, instructions, and location lists compete with coherent premise, role, opening, and consequence passages; truncation can leave descriptions without their qualifiers or outcomes. Combining heterogeneous passages into one vector can also dilute their relevance, but that embedding effect requires measurement. Even eliminating truncation would not establish that one full-content vector is better. Full source fidelity, coherent evidence selection, and retrieval quality are distinct properties. Enlarging one byte limit alone does not resolve them.

Language and taxonomy uncertainty

An 800-byte allowance fits roughly 266 three-byte CJK characters versus 800 ASCII characters; neither character count nor token count measures equal meaning. Several translated/localized stories retain English production presets and Chinese tags. These are concrete budget and mixed-script concerns. The twelve families are not paired translations, so their differences cannot identify a causal language effect or explain the ZH/ES regressions.

The evidence rubric groups investigation and survival under mystery. Development disagreements include generic control promises versus consequential agency, rosters or NPC-generation instructions versus enacted ensemble interaction, replenishment versus progression, and real-person idol fandom versus fictional franchises. The latest complete blind grid contains 77 supported, 71 unsupported, and 20 ambiguous decisions. Of 70 semantic predictions, it supports 53, rejects seven, and marks ten ambiguous. These disagreements expose taxonomy uncertainty, not objective error rates. Literal provenance and independent AI calls do not establish semantic correctness; greater facet coverage is not retrieval uplift.

Missingness changes the ranking experiment

Code-proven. In .local-artifacts/assemble-retrieval.py, equal RRF adds 1 / (60 + rank) for content and, when available, experience. With fifteen candidates, every story with both vectors has score at least

2 / 75 ≈ 0.02667 > 1 / 61 ≈ 0.01639,

the maximum possible score for a story missing experience. Thus every dual-vector story must outrank every missing-experience story in its language, regardless of content rank. Keeping failed extractions eligible does not remove this availability advantage. This proves an ordering confound, not its numeric contribution to the measured NDCG gap.

The next rule must be missingness-neutral: score = baselineScore + delta, where absent, failed, or ambiguous aspect evidence yields delta = 0. Available evidence must earn a bounded, query-dependent adjustment; mere extraction success must not add a constant bonus or guarantee dominance. Other candidates' adjustments can still change a missing story's ordinal position. Define the score scale, adjustment bound, tie behavior, and calibration on development data/synthetic fixtures, then freeze them before evaluation. This report does not select an optimal bound or authorize serving changes.

Candidate and gates

  1. Preserve baseline. Keep its exact text, hash, vector, and serving behavior. Use a separately versioned experimental artifact; preserve all prior results.
  2. Add one activity aspect: what the reader does and what follows. Select complete source passages describing concrete activities, enacted relationships, or action→consequence rules; retain role/situation needed to interpret them. Broad taxonomy labels are not substitutes for evidence. Inspect passages before the second compression. Exclude production boilerplate without blanket-banning system-shaped sections containing legitimate fictional facts.
  3. Bound and review evidence. A starting proposal is at most three evidence bundles, each at most 256 embedding-model tokens, plus an 8-KiB artifact ceiling. These limits are engineering proposals, not established optima. Preserve source IDs, offsets, hashes, route/opening identity, and verdicts. Bound complete clauses before review; never silently truncate accepted evidence. Keep incompatible openings separate, prevent cross-route conjunctions, and cap/deduplicate their story-level contribution. Record overflow and abstention.
  4. Run development regressions after performance validation. Require exact baseline parity; preservation of the observed action/consequence examples; complete reviewed quotes; deterministic provenance; and explicit fallback. Use paired synthetic EN, ZH-Hans, ZH-Hant, JA, and ES scenarios with the same intended propositions, including late evidence, long cast lists, headings, English presets, negation, and alternative openings. Test danger versus survival, secrets versus deduction, rosters versus interaction, and generic choices versus concrete consequences. These fixtures test declared semantics, not natural-data generalization. Ranking tests must expose the RRF inequality and verify neutral missingness rather than assuming eligibility suffices.
  5. Freeze a new blind diagnostic. A proposed workload is 120 new families, 30 per language, and 16 independently authored intent groups translated into four languages: 1,920 story–intent judgments and 64 rankings. This is neither an optimal cohort size nor a power calculation. A custodian supplies exclusions for pilot, development, frozen-24, and consumed-60 families without exposing the frozen-24 contents. Include premise/role and compound activity/consequence requests; check translation equivalence independently.
  6. Predeclare the decision. Freeze public snapshots, exclusions, queries, models, selection/review rules, scoring, failure handling, and metric definitions before rankings. Compare baseline against baseline plus the activity aspect with the same embedding model. If fusion also changes, preregister an old- experience/new-fusion diagnostic arm to help distinguish representation and scoring effects. Assessors receive common public sources, never representations or rankings. Preserve uncertainty and literal evidence. Independent bilingual human adjudication would strengthen validity if arranged; none is provided or assumed. An AI-only assessment must remain explicitly AI-referenced.

Primary evaluation remains macro NDCG@5, with strong-match recall@5, per-language paired differences, relevant-inventory counts, and extraction coverage. Freeze undefined-inventory handling and uncertainty treatment; never silently drop failed candidates or convert missing judgments to irrelevance. Translations are correlated observations of sixteen intent groups, not sixty-four independent queries. Paired intervals over intent groups describe the fixed sampled inventory; catalog-wide claims require broader family sampling. Predeclare a meaningful effect target and per-language noninferiority margins and assess whether the planned workload can resolve them. Wide intervals are inconclusive; unmet improvement or noninferiority gates block advancement.

No tuning or promotion follows from the consumed sixty families. A revised candidate needs another untouched evaluation after development; even a successful static diagnostic would not establish reader uplift or authorize production rollout. Candidate development, provider expenditure, and reader experiments are future work, separate from this documentation change.