Skip to content

Static retrieval diagnostic: retain the baseline

The richer content representation did not beat the published-greetings baseline in the frozen sixty-family diagnostic. Adding reviewed experience evidence recovered some of the loss, with substantial differences by language. Keep experimental personalization disabled. These results do not establish that the production feed became better or worse: this is a small current-content retrieval test with AI relevance labels, not a personalized reader experiment.

Frozen comparison

The protocol and annotation amendment were recorded before the first ranking/metric computation. Sixty family-distinct public stories (fifteen per language) exclude the earlier pilot, development and frozen twenty-four-story evaluation families. Eight reader intents translated into four languages produce 32 query/language observations. Each method ranks the same fifteen candidates per query, without removing failed experience extractions. All vectors use text-embedding-3-small.

RepresentationMacro NDCG@5Strong-match recall@5
Published title, description, tags and opening (foundation PR baseline)0.64260.6923
Rich bounded content (characters, lore and mechanics included)0.53850.6490
Rich content + reviewed experience, equal-weight RRF-600.61830.6699

NDCG uses gains 2^grade - 1 and all 32 queries. Strong-match recall uses the 26 queries with at least one grade-two candidate; six queries have no strong inventory and are explicitly undefined, not assigned zero. Translations are correlated observations of eight intents. No significance or causal claim is made.

LanguageBaseline NDCG@5Rich contentContent + experience
English0.76240.72820.8214
Chinese0.54770.38290.3786
Japanese0.60780.52940.7071
Spanish0.65260.51360.5660

Content vectors succeeded for all sixty stories. Reviewed experience vectors were available for 56; two proposals and two reviews failed their bounded contracts. Those four stories remained in every method's eligible inventory, with only their content contribution in RRF. This is a deliberately simple, fixed fusion diagnostic, not the serving policy's learned score or an optimized fusion weight. Do not tune weights on this consumed evaluation sample.

Label provenance and limitations

An independent relevance call receives full bounded public sources and the reader intents, never descriptors, vectors or rankings. The initial free-form quote interface rejected thirteen responses. Thirty-three same-rubric retries left eight invalid. The uniform replacement reference grades all sixty stories by selecting numbered literal source passages; thirty paired-story requests supplied all 480 decisions with valid references. The original grades are not mixed into the primary evaluation. The same model reviews experience evidence and judges relevance, so correlated AI bias remains possible despite separate blind calls. Quote provenance is not proof of semantic correctness or human preference.

The separate twelve-story development audit compared two blind AI references: seventy positive facets agreed, two earlier positives were rejected by the new reference, five became ambiguous, and seven additional positives were identified. The complete new grid contains 77 supported, 71 unsupported and 20 ambiguous decisions. Against its supported decisions, the cue-validated baseline matched 18/77; semantic review matched 53/77. Of seventy semantic predictions, seven were unsupported and ten ambiguous. Better facet coverage therefore does not by itself establish better retrieval. Generic choice promises, cast lists and ordinary danger should not be overinterpreted as agency, ensemble or survival experiences.

The diagnostic used 374 provider requests including failed attempts and one quote-format diagnostic. OpenRouter reported $6.323260 total; embeddings used 178,970 tokens. This is observed experiment usage, not a catalog-wide cost estimate. No private conversation text, reader personality inference or production profile write was used.

Reproduction and next decision

trainer/story_retrieval_evaluation.py validates complete judgment/ranking coverage, family exclusions, input hashes, frozen protocol binding, identical embedding-model declarations and eligible permutations before calculating metrics. It reports paired differences on matching denominators and exposes per-language, per-intent and undefined-inventory counts. Fifty-one synthetic tests cover those contracts and hand-computable arithmetic. The artifact assembler additionally checks the pinned source/protocol/amendment bytes and literal relevance evidence; the scorer alone does not certify the producer's declared provenance.

All sixty canonical story inputs and baseline hashes were revalidated. Rich content was rebuilt from source; all 56 successful experience representations were independently reconstructed from complete, hash-bound review decisions. Raw public-story text, vectors and assessor outputs remain in ignored local artifacts. Aggregate reports contain no reader identifiers.

Proceed with the foundation's continuation, exclusions and measurement work while retaining its scoring baseline. Diagnose representation dilution and multilingual source budgets on development data before freezing another candidate. A useful next candidate should preserve the coherent premise/opening representation and use additional aspects as separate evidence rather than assuming that packing more text into one vector improves matching. That is a hypothesis for a fresh comparison, not a result of this test. Promotion still requires reliable canonical exposures, matured outcomes, capacity checks and a controlled reader experiment.