Personalization readiness execution plan
User authorization: complete the testing and preparation required for recommendations, using existing access. Work starts from ce3b4a515, PR #179, in the existing isolated worktree. The user subsequently enabled temporary MCP writes/cleanup for restricted ClickHouse account provisioning. That access step is complete: both migration credentials are in Railway, managed tests passed, owned storage/test grants are removed, and the connector is read-only again.
Acceptance and limits
- Production product-data observations remain bounded and read-only. The separately authorized access step may provision the restricted migration identity, its Railway variables and an isolated synthetic ClickHouse database on the managed service; it does not activate production capture, migrations or ranking. Provider calls use public published story content and declared request limits.
- Prove testing endpoints differ from production before writes. Prefer a disposable database and an exact-key Redis namespace; never overwrite the shared testing catalog or clear shared Redis.
- Disable outbound product analytics, billing, email and unrelated background work in the test application. Synthetic actors must stay inside isolated test storage.
- Report actual SQL/Redis/HTTP measurements separately from synthetic codec timings. Passing correctness tests is not measured reader uplift or a scale guarantee.
- A model-produced independent reference is an audit reference, not human ground truth. Real retention/outcome windows cannot be accelerated.
Tasks
- [x] 1. Audit current testing/source readiness; provision isolated test storage and run the unmodified measurement/history/personalization installers twice to establish compatibility/idempotence.
- [x] 2. Freeze a family-distinct, language/category-stratified public-story benchmark. Independently annotate source evidence before examining extraction predictions. Compare supported facets, omissions, empty profiles, latency and cost without tuning on the held-out slice. Benchmark completed; quality promotion remains unpassed.
- [x] 3. Exercise the real bounded worker and pgvector retrieval on the isolated database: dry run, apply, repeat, edited/unpublished source, concurrent extraction, missing/stale profiles and provider failure. Fix actual regressions with focused tests.
- [x] 4. Extend and run the existing exact-key testing Redis harness with v2 snapshots and preference-aware recovery; measure bounded concurrent traffic, memory and cleanup. Keep all pre-existing safeguards and distinguish WAN measurements from origin capacity.
- [x] 5. Run isolated HTTP and SQL checks with synthetic guests/accounts: feed continuation, library exclusions, preference lifecycle/handoff, shadow equivalence, failures/erasure and bounded concurrency. Audit warehouse capture prerequisites and install/test only in isolation. The managed warehouse suite passed all nine checks, including cleanup; production CDC remains a separate rollout gate.
- [x] 6. Review fixes, run affected verification, publish evidence to PR #179, and document the remaining online experiment duration and explicit promotion gates. Explain the user experience simply. Completing this readiness investigation does not promote the model or close the quality, production CDC, full-origin capacity and online-outcome gates.
Files and ownership
- Controller: environment scripts under ignored
.local-artifacts/readiness-*, test application/storage setup, benchmark corpus preparation, readiness report and integration fixes. - Capacity worker:
packages/server/scripts/check-discovery-testing-redis.ts, a new bounded capacity script and narrowly related script tests only. No infrastructure writes until controller supplies verified runtime credentials. - Quality auditor: ignored frozen-source annotations and rubric report only; no provider output or production credentials required.
Execution decisions
- Current Railway testing app and Redis exist, but its catalog is shared (1,998 rows). Use separate storage for tests.
- Keep production flags off until the documented engineering gates pass. An offline benchmark cannot replace the subsequent stable online experiment.
