Skip to content

Actual catalog, reader behavior, and Railway Redis

Read-only investigation, September 19, 2026, approximately 21:15–21:30 UTC. This updates the generic category examples in the personalization proposal. No production data, settings, services, subscriptions, or deployed code changed.

Finding: Yumina's current catalog and audience call for recommendation models that understand fandom, character relationships, and kinds of interactive play. A conventional fiction-genre picker misses much of what distinguishes these worlds. Sci-fi and horror are small tagged catalog segments; K-pop is also small in supply but has substantial concentrated engagement. Neither inventory counts nor overall popularity should determine every reader's feed.

Evidence and definitions

  • Production Neon DATABASE_READ_URL: confirmed transaction_read_only=on and pg_is_in_recovery()=true. SQL ran with a 30-second statement timeout.
  • Current public catalog: is_published=true, status='published', visibility='public'. First-party game cards are excluded from story analysis. Translations are grouped by language_group_id, falling back to world ID. Counts are current, not a reconstructed historical catalog.
  • ClickHouse: CDC source views / raw tables with FINAL and deletion filtering. The empty destination play_intervals and world_dimensions tables are not evidence of absent data; the deployed source views read the populated CDC tables.
  • Main comparison window: September 14 00:00 through September 19 00:00 UTC, five complete days. Foreground collection began September 13, so an earlier full week of this metric does not exist.
  • A qualified reader/story pair has at least 60 unioned foreground seconds over the window and at least one foreground-engaged-60s event. Intervals are clipped to the window, split at UTC midnight, and unioned per person/story family/day. Intervals longer than 90 seconds are excluded. Current creators reading their own worlds and known admin/internal accounts are excluded.
  • Repeat-day readers have qualifying engagement on two or more UTC dates for that family in the window. This is descriptive repeat use, not D1/D7 retention or a recommendation-attributed return rate. Published titles identify groups; the reader may have used a translation.
  • Historical cross-check: August 22–September 18 UTC, counting readers with at least three send usage records per story family. Regenerations, memory jobs, Studio calls, and empty-send endpoints are not counted as new story turns. Usage records remain a behavioral proxy, not audited successful completions.
  • PostHog: yumina.io, surface='recommended', same five-day window, aggregated hub_impression / hub_click events. Event-time $is_identified separates anonymous from identified SDK activity. This is an approximation of login state, not an independent authentication check. All 3,041 world/cohort rows were paginated and reconciled to the independent overall totals.
  • Inspected public descriptions across prominent stories and eight selected worlds' published greeting entries. This is qualitative sampling, not a claim to have read every story or validated generated semantic labels.
  • No private player chat text, account names, or emails were needed. Actor-level intermediate aggregates remained in process memory; retained results are catalog metadata and aggregate statistics.

Remaining limitations: old interval/activity records lack an explicit historical consumer-eligibility field. Excluding current creators/internal accounts cannot perfectly remove historical collaborator playtests or identify deleted sessions. Requiring engagement reduces idle-time noise but does not make every foreground second meaningful reading. Summed story-hours can overlap across different stories on concurrent tabs. Exposure, acquisition, language, world age, existing Library use, and current-catalog survivorship all influence these descriptive results.

What is available

There are 1,650 public published rows, including four first-party game cards; 1,646 story rows, 1,634 primary story variants, and 1,017 distinct story families. These are wider than any individual's eligible Discover inventory.

LanguagePrimary public story variantsRated all-ages
English840541
Chinese (zh)360235
Japanese215156
Spanish213159

Six other primary variants use French, Portuguese, German, or zh-CN. Across all languages, 1,096 primary variants are rated all-ages and 538 sensitive. Language, content settings, audience filters, Library membership and seen-history reduce supply further. A large global catalog does not guarantee endless unique content in every language and eligibility combination.

Tags below are overlapping evidence labels, not exclusive genres. Romance combines 恋爱, 言情, and romance-labelled subtypes; fantasy combines fantasy labels with 玄幻/仙侠/修仙; horror includes 恐怖 and horror subtypes; sci-fi includes 科幻, 科幻奇幻 and English sci-fi spellings. Existing shared tag canonicalization is already present; the missing piece is richer semantic coverage, not simply adding another alias table.

Tagged categoryDistinct public story familiesQualified readers, five days
Fandom3662,162
Simulators2282,030
Romance159990
Anime139622
Fantasy / cultivation115449
Mystery / investigation57239
School life42620
Everyday life39153
Horror30191
Science fiction3067
K-pop29810

The main qualified sample contains 3,587 readers across 644 story families, with 6,615 reader/story pairs. Counts across categories must not be added. Sparse tags also understate real themes: a story can contain science-fiction worldbuilding or relationship development without using those labels.

K-pop-tagged families account for 22.6% of sampled readers and approximately 31.6% of attributed story-hours. Horror and sci-fi are each tagged on about 3% of the catalog, but horror has a substantial individual hit. Do not interpret small catalog supply as proof of unwanted content or fill the global feed with K-pop solely because some readers spend many hours there.

Worlds that people actually use

Story familyQualified readersRepeat-day readersAttributed foreground hours
SEVENTEEN Simulator4373351,987.2
Sakura Season27835117.8
Liriel16835229.7
Poison in the Bottle · Battle Royale13236199.8
Marvel RPG11834258.1
BTS Simulator10348277.8
Jujutsu Kaisen931777.9
ATEEZ Simulator8472478.5
Stray Kids Simulator7945235.7
Naruto: The Will of Fire7340281.0

This is a selected comparison, not an exhaustive ranking. In particular, these repeat-use counts are not fair head-to-head quality scores: audiences, exposure, publication age, and pre-existing reader relationships differ.

The 28-day cross-check also places Sakura Season (996 readers with at least three send records), SEVENTEEN (791), Liriel (651), Poison in the Bottle (445), Marvel RPG (353), and Jujutsu Kaisen (318) near the top. This makes the broad audience patterns more credible than a single day's spike, without establishing causality.

The published content explains why a single genre vector is insufficient:

  • SEVENTEEN: an ensemble relationship simulator with group-chat interaction, everyday companionship, romance routes, and multiple settings. Even its unit variants emphasize different experiences: atmosphere/exploration, group chemistry, or gentle emotional interaction. Fandom membership alone is too coarse.
  • Sakura Season: school-life reunion and character relationships. Sharing a school setting with another card does not imply the same desired experience.
  • Poison in the Bottle: school setting combined with survival, alliances, betrayal, and consequential choices. Its resemblance to Sakura Season on a school tag should not outweigh this tonal and mechanical difference.
  • Marvel RPG / Jujutsu Kaisen: self-insert participation in a known universe, character/power creation, missions and progression. This differs from merely chatting with a favorite canon character.
  • Love and Deepspace: relationship development and long-form companionship inside a sci-fi mystery setting. A one-label science fiction classification would miss the likely attraction, while its existing tags omit sci-fi altogether.
  • RPG Engine & World Sim.: explicitly promises persistent time, inventory, quests, NPC relationships and open-ended decisions, yet has zero tags. It still attracted 59 qualified readers. The catalog has 80 untagged primary story families, 45 of which appear in this qualified activity sample.
  • 萬靈天下: a mythic world with multiple life paths, species and social settings. Its opening offers career, military, commerce, cultivation and exploration choices; these interaction affordances matter alongside fantasy.

These observations support separate descriptor axes for franchise/characters, relationship dynamics, player role, interaction mechanics, tone/intensity, setting, and pace, with multilingual labels, evidence spans and source/version tracking. Descriptions and opening text are evidence, not instructions to the classifier. Keep creator declarations distinct from model interpretations and actual behavior.

Reader differences and discovery implications

Among qualified readers, 1,993 read an English variant and 1,517 a Chinese variant; 113 read Spanish and six Japanese. These groups overlap. Language-tagged inventory does not imply a comparably sized active audience.

  • Of the Chinese-variant readers, 787/1,517 (51.9%) read a K-pop-tagged family. Of English-variant readers, 790/1,993 (39.6%) read a 恋爱-tagged family, and 585 read an anime-tagged family. Use language as context and an eligibility signal; these aggregate patterns are not rules for every person using that language.
  • 2,389 readers (66.6%) qualified on one family; 1,198 (33.4%) qualified on more than one during the five days. This does not measure an enduring personality or prove that one-story readers dislike exploration.
  • The top ten families by attributed time account for 42.3% of time; the top twenty, 54.3%. Raw playtime optimization would give existing long-session worlds a large advantage. Discover's objective must remain valuable new-story starts, with old-story continuity handled through Library.
  • There are useful behavioral connections beyond tags: 10 qualified readers shared Mushoku Tensei and Reincarnated as a Slime; 16 shared Aspen and Liriel. These are candidate collaborative signals, not sufficient sample sizes to declare a universal relationship. Language, exposure, support thresholds, shrinkage, and a later validation period must qualify those edges.

PostHog supplies the guest perspective missing from authenticated play intervals:

Event-time SDK identityRecommended impressionsClick eventsRaw clicks / impressions
Anonymous139,1254,8173.46%
Identified380,89616,1304.23%

Anonymous impressions are 26.8% of this telemetry. Click/impression ratios are descriptive event ratios, not deduplicated conversion rates or evidence that one policy wins. Browser collection misses opt-outs/blockers and reflects the existing ranking. Person-level historical identity can merge after login, hence the explicit event-time split (PostHog identity documentation).

Anonymous clicks are led by Liriel (619), Poison in the Bottle family (286), Sakura Season family (282), Marvel RPG (247), and Aspen (237). K-pop-tagged variants receive 2,397 anonymous impressions / 36 clicks, versus 22,088 / 231 identified impressions/clicks. This is not evidence that guests inherently dislike K-pop: current exposure, language/acquisition mix, Library exclusions and returning fandom traffic differ. Separate cold discovery from established reading.

Signal correction: POST /sessions automatically inserts a published story into user_library. The seven-day Library-insertion ranking therefore mixes manual saves with starts and is not a preference/vote ranking. The new measurement contract should preserve manual-save intent, successful insertion and session-start provenance separately, and avoid counting one action twice as independent evidence. The automatic insertion was verified in the source of the live production commit, 631752cd9cfaa4d037690bfdd6cff4456b168528, as well as the discovery worktree.

Revised optional-interest proposal

Use a short, skippable experience picker, then learn from actual use. Candidate choices grounded in this audit are:

  1. Romance and character relationships.
  2. Anime and game universes.
  3. Idols and K-pop.
  4. Open-world adventures and simulators.
  5. Fantasy and cultivation.
  6. Mystery and survival.

These are proposed labels, not final implemented taxonomy or a forced six-way classification. Experiences overlap. Show only choices supported by sufficient eligible content in the selected language, and keep niche interests reachable through search/refinement. Sci-fi and horror can remain valid interests without being prominent universal onboarding defaults. Tone can be a later optional refinement; choices must not bypass content settings or infer the reader's gender.

Before shipping: map and review representative stories for each choice, inspect false matches and missing tags, test coverage after access/Library exclusions, and run the already proposed randomized offer-vs-no-offer experiment. A choice seeds a soft, separate interest; skipping or leaving another choice unselected does not imply dislike. Do not interrupt direct links to a selected story.

Existing Railway Redis

Confirmed against live Railway service variables, without retaining credentials:

CheckResult
Production Railway Redis serviceRunning, Redis 8.6.2
Production app REDIS_URL points to that serviceYes
Production analytics worker shares that serviceYes
Testing Redis is a separate endpointYes
Production DISCOVERY_REDIS_URL overrideAbsent
Memory at inspection183.97 MiB; historical process peak 485.47 MiB
Connected clients / operations at inspection6 / approximately 95 per second
Redis maxmemory / policy0 / noeviction
Evicted keys / rejected connections since process start0 / 0

Redis is the same technology already used for fast caches, rate limits and cross-instance messages. The separate subscriber client in code is another connection to the same service, not a second Redis deployment.

The proposed dedicated recommendation Redis means another instance of the same technology, potentially in the same Railway project, so feed state can have its own capacity limits and failures. It does not require opening a Redis Cloud or Upstash account. A separate key prefix or Redis database number would not provide that memory/failure isolation.

maxmemory=0 means Redis itself has no dataset memory ceiling; Railway/container limits still apply. This snapshot is not a load test or evidence of imminent failure. Before large feed-state rollout, establish an explicit memory/admission budget and measure concurrent feed sessions, keys and payload sizes. An isolated instance is a reasonable scale boundary, not a purchase justified solely by 184 MiB of current use. See Redis memory and eviction documentation.

User action now: no new account or subscription is required for this work. Existing data access was sufficient. Prioritize the measured serving/attribution foundations, evidence-backed story descriptors, and language-aware multi-interest profiles before selecting additional storage or training vendors.

Reproduction artifacts

Credential-free local audit scripts and aggregate outputs are retained under .local-artifacts/ in this worktree and are intentionally untracked:

  • catalog-audit.mjs: read-replica catalog, Library-insertion aggregates, Redis inventory and metrics, ClickHouse source coverage.
  • catalog-demand.mjs: interval union, qualifying activity, family/tag/language metrics, and aggregated co-reading edges.
  • catalog-history.mjs: 28-day send-record cross-check and published-content samples.
  • catalog-audit-{supply,demand,categories,history,connections,redis,coverage}.json.
  • catalog-audit-posthog.json: complete paginated world/cohort aggregates, with totals reconciled to 520,021 impressions and 20,947 clicks.

No application tests were rerun for this research-only change. Verification was query/result reconciliation, source-schema/code inspection, public catalog eligibility checks, and credential-free output review.