Skip to content

Personalization readiness evidence

Latest follow-up: the frozen retrieval comparison did not justify semantic promotion. Production preparation has since installed source history, measurement schema and the warehouse archive; the dated checks below describe the earlier isolated readiness run.

Follow-up: semantic evidence and Redis storage measurements use a separate 12-story development sample and 18 paired Railway testing runs. Semantic reference coverage improved, but disagreements remain; compact storage reduced bytes without a consistent latency gain. Neither change is promoted. Production descriptor validation and raw storage remain the defaults.

The continuation, identity and worker paths passed isolated integration checks. Personalization is not approved for production promotion. Story-facet coverage, production capture/CDC, full origin HTTP capacity and an online outcome experiment remain open. Managed warehouse access and the isolated live migration tests are now verified. Production story data, application schema, ranking and feature flags were not changed; a restricted migration identity and its two Railway variables were added without deployment. Work belongs to draft PR #179, stacked on the foundation and measurement pull requests.

Reader experience

Discover stays a feed of new stories. A reader can optionally choose a few interests when the eligible catalog supports those choices, skip the chooser, and edit or reset choices later. Reading and saving provide stronger evidence than a click; several enduring interests and current-visit intent remain separate. Related stories mix with discovery opportunities rather than treating one click as a permanent preference. Saved or started story families leave Discover and remain accessible through Library. Continuation reaches eligible inventory beyond the initial recommendation pool; a finite catalog cannot provide infinitely many unique eligible stories.

Real infrastructure checks

Testing PostgreSQL and Redis were verified distinct from production before writes. The shared testing catalog was not modified. A uniquely named disposable PostgreSQL database received the current schema (151 tables), then the unmodified measurement, world-history and personalization SQL installers twice each. PostgreSQL 17.11 and pgvector 0.8.2 executed the tests. Credentials stayed in memory; no private messages or production user records were copied.

PathObserved resultLimit
Live story worker23 of 24 public worlds stored with real embeddings; dry run made zero provider callsOne provider refusal remains a failed first-pass example
Worker races/faultsSeven real-SQL checks passed: unchanged source, cache freshness, source edit, concurrent write, failure preservation and unpublishFaults use injected deterministic providers, separately from live extraction
pgvectorThree favorites from distinct families produced three clusters; five eligible English results respected an exclusion; pair features finiteOnly 23 stored profiles; not a recall-quality or large-index benchmark
Hono over TCP11 phases, 119 HTTP requests, all passedFour real route modules and real auth, not a full deployed app/browser
Inventory continuationAll 713 synthetic families reached without duplicate families; 717 total including eligible corpus rowsFinite seeded catalog, no claim of infinite unique content
Library and identityConfirmed save emitted one canonical outcome and verified handoff; two saved language families stayed excludedSynthetic accounts inside disposable storage
Preferences/failureSkip/select/reset, idempotence, account isolation, one-time handoff and optional preference SQL failure passedEligible English chooser options were empty; no claim that the UI offer was shown
Redis correctnessNine checks passed with v2 snapshots and preference-aware recovery; exact-key cleanup verifiedRuntime credentials explicitly selected testing Redis

The HTTP probe also checked baseline/shadow ordering equality, cursor replay, stolen-cursor rejection and canonical impression deduplication. A deliberately failed optional preference query preserved baseline IDs and HTTP 200. Its final cleanup found zero owned Redis keys and zero synthetic worlds, removed synthetic auth sessions, preserved the corpus and closed all connections.

Windows-loopback HTTP with remote testing SQL/Redis had p95 704/1,195/1,766 ms at 1/4/8 actors. These are tiny WAN samples, not stable tail estimates. The app pool reached 8 connections with 72 queued requests. SQL/queue work needs measurement inside the actual serving topology before selecting launch capacity.

Railway Redis probe

A roughly 19.6-second desktop clock offset correctly blocked the stricter local capacity harness before traffic. The same probe ran in a temporary Node 22 Railway testing service using only the existing testing Redis reference. No clock guard was weakened. The service had no public endpoint, no provider/DB credentials and no restart policy; it was removed after verification.

The final machine-readable run used 1,000 v2 cards, two Redis connections, at most two simultaneous calls per actor, a ten-second dispatch ceiling and 200 requests per run. Deliberate duplicate cursors exercise compare-and-swap contention.

ActorsSuccessful / attemptedAdmission rejectionsSuccessful p50 / p95 / p99 msSampled owned key memory
1172 / 2002812.4 / 107.2 / 136.7459,496 B
4200 / 2000176.0 / 402.7 / 531.44,196,973 B
8200 / 2000234.3 / 804.3 / 1,367.88,393,843 B

All runs reported zero unexpected request/transport errors and verified exact-key cleanup. The one-actor probe deliberately limits admission to one resident session: exhausted fresh-session retries can be rejected. These are reported separately, not counted as successful pages. Raw CAS conflicts are also expected from duplicate concurrent requests. The earlier run passed but emitted multiline logs that Railway interleaved; the final run changed only report capture to one structured record per probe.

This validates bounded session behavior, not a scale target. Maximum attempted serialized session body was about 1.25 MB. In the eight-actor run Redis read/write traffic was about 223/213 MB over 200 attempted calls, including retries. This cost is material: increase neither admission limits nor confidence in launch capacity from these samples. Profile serialization, queued snapshots, CAS retries and SQL work before a broad rollout. Memory samples exclude unsampled peaks, allocator reserve and replication; no production throughput projection is justified. The probe's generic WAN warning is overridden only as to its execution location here: it ran inside Railway, but did not include HTTP, SQL or actual candidate ranking.

Frozen extraction comparison

The benchmark contains 24 family-distinct public worlds, six each with English, Chinese, Japanese and Spanish metadata, covering idols, fandom, relationships, simulation, fantasy and other content. It excludes the earlier eight-world development pilot. Selection was fixed and stratified, not population weighted. Canonical input SHA-256: b0a5f5da97549e48c688a860ce76c1a1da086a808921fcb360becb3cfdd306f8.

Three independent source-only AI audits were completed without access to model predictions. All 163 reference interest/experience labels include verified literal source quotations. This is an independent AI reference, not human ground truth; disagreement does not establish which side is wrong. No prompt, taxonomy or successful-response validator was tuned on these 24 worlds.

Metric, failed first requests includedGemini 3 Flash PreviewClaude Sonnet 4.6
Accepted descriptors23 / 2423 / 24
Worlds with an experience facet20 / 2416 / 24
Retained / proposed facets, all kinds123 / 20887 / 224
Retained / proposed quotes150 / 32295 / 326
Interest/experience predictions6343
Labels shared with reference5538
Reference agreement among predictions87.3%88.4%
Coverage of reference labels33.7%23.3%
Reported extraction cost$0.122646$0.775653

Exact model IDs: google/gemini-3-flash-preview and anthropic/claude-sonnet-4.6. Both used story-content-v2 and story-descriptor-v3; live vectors used text-embedding-3-small with 1,536 dimensions. Costs exclude embeddings and separate diagnostic calls. They are observed small-run costs, not subscription or catalog-backfill estimates.

Gemini reference coverage by language was 39.0% English, 22.2% Chinese, 33.3% Japanese and 39.0% Spanish; Sonnet was 24.4%, 13.9%, 22.2% and 31.7%. With six worlds per language these values identify investigation targets, not reliable language-wide quality estimates.

Literal-evidence validation and conservative vocabulary discard many proposed facets. The gap may contain both correct rejections and missed semantic support; retention counters alone cannot distinguish them. More expensive extraction did not solve it. Content vectors still represent the sampled published prose independently, so missing facets do not erase all semantic matching.

Next quality work must separate source selection, model omissions, verifier false rejections and reference ambiguity on a separate development set. Build contrastive multilingual cases for explicit support, paraphrase, denial, incidental mention and model-facing instructions. Evaluate a semantic verifier against conservative cues while retaining exact source provenance and bounded inputs; do not simply remove validation. Freeze the resulting configuration before a new unseen, independently adjudicated test. Keep descriptor-driven ranking and onboarding off until this gate passes. This benchmark supplies no personality inference, calibrated confidence, reader uplift or model promotion.

Actual defect fixed

One live provider returned HTTP 200 with an embedded error code 403 and no model identity. The adapter had classified that as a model mismatch. It now recognizes error envelopes first, retains only a validated numeric status, and never logs the provider body. Permanent refusals are not retried; transient statuses remain within existing worker budgets. Caller cancellation wins over status and no successful-response provenance checks were relaxed. Focused adapter/worker regressions passed 83 tests; an independent review found no actionable issue.

Warehouse access and rollout gates

The initial access audit found that the analytics credential could not create test storage and the admin MCP connection enforced read-only mode. The user subsequently enabled the documented connector write/cleanup flags. A fresh connection through that same configured MCP transport verified readonly=0; no admin credential was copied to a direct HTTP client.

On September 20, 2026 at 05:26 UTC (September 19 Pacific), all nine managed ClickHouse checks passed without skips on version 26.2.1.641. The unchanged migration ran repeatedly in an owned database, with real SharedReplacingMergeTree storage. The run covered replacement across months, UInt64/microsecond preservation, receipt-age retention filtering, permanent-user/guest erasure and late replay, importer recovery, and exact-ID pruning reconciliation. It inserted 135 synthetic rows / 59,650 bytes, then dropped the database synchronously. Independent metadata verification confirmed the database was absent.

The first run passed all eight behavioral subtests but its cleanup failed: dropping a nonempty database required DROP TABLE in addition to DROP DATABASE. That missing grant was added only for the exact owned database. After cleaning that database, a fresh full run passed including cleanup. No test writes touched yumina_raw or yumina_analytics. Physical TTL merge completion was not asserted; the test proves expired rows are excluded from consumer reads.

The separate yumina_discovery_migrator identity is stored in the existing yumina-analytics Railway service as CLICKHOUSE_MIGRATION_USERNAME and CLICKHOUSE_MIGRATION_PASSWORD. Its durable grants are SELECT/CREATE TABLE on the two discovery archive tables and SELECT/CREATE VIEW/DROP VIEW on source_discovery_events. It has no roles or grant option. Eight live permission checks verified allowed DDL and denied unrelated-table creation, archive INSERT, production-database deletion, raw-user reads, user management and test-database creation after revocation. All exact-test-database grants were revoked.

Both Railway values were supplied through stdin, with deployment suppressed; all existing variables were verified unchanged. Secrets were generated in memory and not written to source, artifacts or command arguments. The two temporary MCP flags were removed afterward; a new normal MCP connection confirmed locked readonly=1. No user setup or new subscription is needed for this access step. The identity intentionally cannot run unrelated analytics migrations or provision another disposable database without a new scoped grant.

Explicit worker migration commands now require the migration variable pair; normal runs keep their existing importer credentials. Missing or partial pairs fail before network access, with variable names only in validation errors. Focused configuration and real-CLI tests passed seven checks; the CLI regression was also verified to fail when the worker's migration-mode selection was deliberately reverted. A separate compiled-worker loopback probe verified all three DDL requests use the migration identity, invalid configurations make zero warehouse requests, and migration exits without a Redis connection. Hosted build/typecheck and the affected analytics suite passed (96 tests, seven environment-gated skips) before the additional CLI regression was added. The managed test launcher maps the saved migration pair to its explicit ANALYTICS_TEST_CLICKHOUSE_* inputs. The migration and test source SHA-256 values were 276b979d98b5e05a5f1bf2ac594a8d36f58cbbdeac28c4f35da1088a42a7c75a and 635d11535bebd999e11bf9e4d0713fe552bea9362dbdf9a5ab1c4dbda640afe7 respectively.

The refreshed read-only production audit still found the new discovery source tables/history, interval columns, publication membership, warehouse archive/views and import cursor absent. Existing analytics readiness flags do not prove this new pipeline is ready. Install and verify the foundation/measurement capture and CDC path before enabling new measurement. Then validate feature availability and latency in shadow mode before any stable actor-level ranking experiment. The online experiment needs complete outcome windows and existing rollout guardrails; unit tests cannot establish improved qualified reading or retention.

Reproduction and remaining checks

  • Real Redis correctness: pnpm exec tsx scripts/check-discovery-testing-redis.ts --personalization.
  • Bounded Redis traffic: pnpm exec tsx scripts/check-discovery-redis-capacity.ts --actors 8 --cards 1000 --seconds 10 --requests 200.
  • Both require controller-verified testing Redis supplied through DISCOVERY_TEST_REDIS_URL and ACK_TESTING_REDIS=true; neither loads app credentials.
  • Warehouse: direct pnpm exec tsx --test src/analytics/discovery-warehouse.integration.test.ts with the explicit ANALYTICS_TEST_CLICKHOUSE_* environment. The standard isolated test launcher strips remote credentials intentionally. Managed reruns require temporary grants on the exact owned database; include DROP TABLE so nonempty-database cleanup succeeds, then revoke those grants after the test.
  • Capacity regressions use direct pnpm exec tsx --test scripts/discovery-redis-capacity.test.ts; the server test launcher accepts only source-directory tests.

Hosted typecheck and build passed; existing bundle-size warnings remain. Local capacity regressions passed 11 tests with one local-Redis availability skip. The managed warehouse test subsequently passed all nine checks with cleanup. Exact-commit CI results and remaining rollout gates are recorded in the PR. No browser verification was performed.