gitoriaLog in with ident

tracker

All repositories: gitoria

ReadmeCodePull requestsReleasesTicketsSettings
Main branchmain5cb85d75deploy.sh: a backup taken while a background job writes (tar exit 1) is a warning; archive checked with gzip -tmremain/lib/search.hl

14.9 KB

  1. // lib/search.hl — THE SEARCH (tracker.worldapi.org#14): our own database first, then "Fetch from web" (TMDB) for titles
  2. // we don't have yet, and the import of the one picked. Used by components/search.hl (`/search/<text>`).
  3. //
  4. // dbSearch(text) our shows, movies and people — from an IN-MEMORY INDEX built once at boot (no table scan per
  5. // keystroke): every title/name split into words; each word's first 2 and first 3 characters are
  6. // a bucket key → the entries holding such a word. A query looks at ONE bucket (its longest word's
  7. // first 3 characters, or 2 for a 2-letter word) and keeps the entries where EVERY query word is
  8. // the start of one of the entry's words. A title's year is one of its words ("doctor who 2005").
  9. // webSearch(text) GET <TMDB_BASE_URL>/search/multi?query=… (one request; TV + movies only, people dropped) —
  10. // the titles we already have (same TMDB id and kind) are left out
  11. // importTitle(kind, id) a new show record from GET /tv/<id>?append_to_response=external_ids,aggregate_credits (a movie:
  12. // /movie/<id>?…,credits), then sync.hl `syncShowWith` fills it from that same answer exactly like
  13. // the daily sync does (seasons, episodes, poster, external ids); added to the index. The answer goes
  14. // back as `details`: the caller (components/search.hl) stores its cast + crew (credits.hl applyCredits,
  15. // tracker#28 — details.hl imports this file, so it cannot be called from here)
  16. //
  17. // ADULT TITLES (mission 054): our database lists only titles with `adult == false` (shows.hl isPublicTitle; unknown is
  18. // hidden) — `titleOk` holds that per show id, set by indexShow and by `searchTitleChanged` (the adult backfill and the daily
  19. // sync call it per title). The web list asks TMDB with include_adult=false AND keeps only results with `adult` false.
  20. //
  21. // ONE PROCESS PER STORAGE (hybriel#21): the index lives in the app's process and is kept up to date by importTitle; the
  22. // daily sync never changes a title. New rows get self-assigned 16-hex ids (hybriel#113), like sync.hl.
  23. import { now } from 'hl:time'
  24. import { textOr, newId } from './util.hl'
  25. import { showsTable, personsTable, genresTable, isPublicTitle, kindOfTitle, titlePath, liveShowOf } from './shows.hl'
  26. import { tmdbGet, tmdbImageBase, adultOf } from './tmdb.hl'
  27. import { syncShowWith, persistSync } from './sync.hl'
  28. import { refreshCatalogShow } from './catalog.hl'
  29. import { claimShortId } from './shortids.hl'
  30. import { wordsOf, bucketKeys, tmdbKeyOf, yearWordOf, scoreOf, matchesAll, insertTop, yearOfDate, slugBaseOf } from './search-helpers.hl'
  31. // ---- the index --------------------------------------------------------------------------------------
  32. // entries[i] = { kind ('title' | 'person'), id, words, text (the words joined by one space, without the year) }
  33. static entries = []
  34. static buckets = {}
  35. // what we have, by TMDB id: 'series:<id>' / 'movie:<id>' → our show id (the web list leaves these out)
  36. static haveTmdb = {}
  37. // tracker#18 (the daily sync's TVmaze change list, dailysync.hl): TVmaze id → our show id (kept by indexShow and
  38. // searchTitleChanged — the sync fills TVmaze ids)
  39. static haveTvmz = {}
  40. // every show slug in use (an imported title gets a free one)
  41. static slugTaken = {}
  42. // show id → true when the title may be listed (not adult, mission 054)
  43. static titleOk = {}
  44. // tracker#18 (a title renamed by the daily sync): show id → its entry's index; an entry replaced by a newer one is `stale`
  45. // (skipped by dbSearch — the buckets only ever grow)
  46. static titleEntry = {}
  47. static staleEntry = {}
  48. static indexStats = { built = false titles = 0 people = 0 keys = 0 ms = 0 }
  49. // `tie` orders equally good matches: lower first (a title: series before movies, then the newer one)
  50. static addEntry = (kind, id, &name, &extra, tie) => {
  51. words = wordsOf(name)
  52. if (words.length == 0) { return false }
  53. i = entries.length
  54. all = []
  55. for (w of words) { all.push(w) }
  56. if (extra != null && extra != '') { all.push(extra) }
  57. entries.push({ kind = kind id = id words = all text = words.join(' ') tie = tie })
  58. if (kind == 'title') { titleEntry[id] = i }
  59. for (k of bucketKeys(all)) {
  60. if (buckets[k] == null) { buckets[k] = [i] indexStats.keys = indexStats.keys + 1 } else { buckets[k].push(i) }
  61. }
  62. return true
  63. }
  64. static indexShow = (&s) => {
  65. if (s.urlSegment != null) { slugTaken[s.urlSegment] = true }
  66. // tracker#31: a merged tombstone keeps its slug taken, nothing else (its keeper is the title)
  67. if (s.mergedInto != null) {
  68. titleOk[s.id] = false
  69. return null
  70. }
  71. if (s.tmdbId != null) { haveTmdb[tmdbKeyOf(s.type, s.tmdbId)] = s.id }
  72. if (s.tvmzId != null) { haveTvmz['' + s.tvmzId] = s.id }
  73. titleOk[s.id] = isPublicTitle(s)
  74. y = toNumber(yearWordOf(s))
  75. if (addEntry('title', s.id, s.title, yearWordOf(s), (s.type == 'movie' ? 10000 : 0) + (y != null ? 3000 - y : 2999))) { indexStats.titles = indexStats.titles + 1 }
  76. return null
  77. }
  78. // tracker#26 (details.hl): a person added from a title's cast — findable by the search at once
  79. static indexPerson = (&p) => {
  80. ensureIndex()
  81. if (addEntry('person', p.id, p.name, '', 0)) { indexStats.people = indexStats.people + 1 }
  82. return null
  83. }
  84. // a title's adult flag changed (the backfill / the daily sync): listed from now on, or not
  85. static searchTitleChanged = (&showId) => {
  86. s = showsTable.fetch(showId)
  87. if (s != null) { titleOk[s.id] = isPublicTitle(s) }
  88. if (s != null && s.tvmzId != null) { haveTvmz['' + s.tvmzId] = s.id }
  89. return null
  90. }
  91. // tracker#31 (merge.hl): a duplicate became a tombstone of `keeperId` — its entry goes stale, its TMDB id (and TVmaze id) now
  92. // answer the keeper, so no import or filmography adds the title again
  93. static searchTitleMerged = (&tombId, &keeperId, &type, &tmdbId) => {
  94. ensureIndex()
  95. old = titleEntry[tombId]
  96. if (old != null) { staleEntry['' + old] = true }
  97. titleOk[tombId] = false
  98. if (tmdbId != null) { haveTmdb[tmdbKeyOf(type, tmdbId)] = keeperId }
  99. k = showsTable.fetch(keeperId)
  100. if (k != null) {
  101. if (k.urlSegment != null) { slugTaken[k.urlSegment] = true }
  102. if (k.tvmzId != null) { haveTvmz['' + k.tvmzId] = k.id }
  103. titleOk[k.id] = isPublicTitle(k)
  104. }
  105. return null
  106. }
  107. // a title's name changed (the daily sync, tracker#18): its old entry goes stale, a new one is added — found by the new name at once
  108. static searchTitleRenamed = (&showId) => {
  109. if (!indexStats.built) { return null }
  110. s = showsTable.fetch(showId)
  111. if (s == null) { return null }
  112. old = titleEntry[s.id]
  113. if (old != null && entries[old].text == wordsOf(s.title).join(' ')) { return null }
  114. if (old != null) { staleEntry['' + old] = true }
  115. indexShow(s)
  116. return null
  117. }
  118. // built once: at boot (project.hl calls it), else by the first search
  119. static ensureIndex = () => {
  120. if (indexStats.built) { return true }
  121. indexStats.built = true
  122. t0 = now()
  123. for (s of showsTable.find(null, null)) { indexShow(s) }
  124. for (p of personsTable.find(null, null)) {
  125. if (addEntry('person', p.id, p.name, '', 0)) { indexStats.people = indexStats.people + 1 }
  126. }
  127. indexStats.ms = now() - t0
  128. console.log('search index: ' + indexStats.titles + ' titles, ' + indexStats.people + ' people, ' + indexStats.keys + ' keys, ' + indexStats.ms + ' ms')
  129. return true
  130. }
  131. // tracker.worldapi.org#16 (people.hl): our show id for a TMDB title — `type` 'series' | 'movie' — or null
  132. static titleIdByTmdb = (&type, &tmdbId) => {
  133. ensureIndex()
  134. return haveTmdb[tmdbKeyOf(type, tmdbId)]
  135. }
  136. // tracker#18: our show id for a TVmaze id, or null
  137. static titleIdByTvmaze = (&tvmzId) => {
  138. ensureIndex()
  139. return haveTvmz['' + tvmzId]
  140. }
  141. // ---- our database ------------------------------------------------------------------------------------
  142. static titleLimit = 30
  143. static peopleLimit = 10
  144. // { text, words, titles: [{ id }], people: [{ id }], titleCount, peopleCount, tooShort }
  145. static dbSearch = (text) => {
  146. ensureIndex()
  147. out = { text = text titles = [] people = [] titleCount = 0 peopleCount = 0 tooShort = false ms = 0 }
  148. t0 = now()
  149. qw = wordsOf(text)
  150. let longest = ''
  151. for (w of qw) { if (w.length > longest.length) { longest = w } }
  152. if (longest.length < 2) { out.tooShort = true return out }
  153. key = longest.slice(0, longest.length >= 3 ? 3 : 2)
  154. cand = buckets[key]
  155. if (cand == null) { return out }
  156. qtext = qw.join(' ')
  157. let titles = []
  158. let people = []
  159. for (i of cand) {
  160. let e = entries[i]
  161. if ((e.kind != 'title' || (titleOk[e.id] == true && staleEntry['' + i] != true)) && matchesAll(e, qw)) {
  162. let item = { score = scoreOf(e, qw, qtext) len = e.text.length tie = e.tie i = i }
  163. if (e.kind == 'title') {
  164. out.titleCount = out.titleCount + 1
  165. titles = insertTop(titles, item, titleLimit)
  166. } else {
  167. out.peopleCount = out.peopleCount + 1
  168. people = insertTop(people, item, peopleLimit)
  169. }
  170. }
  171. }
  172. for (t of titles) { out.titles.push({ id = entries[t.i].id }) }
  173. for (p of people) { out.people.push({ id = entries[p.i].id }) }
  174. out.ms = now() - t0
  175. return out
  176. }
  177. // ---- the web (TMDB) -----------------------------------------------------------------------------------
  178. // TMDB's search: [{ kind ('tv' | 'movie'), tmdbId, title, year, overview, posterUrl }] of what we DON'T have, or
  179. // { failed }. One request (search/multi, page 1 = up to 20 results); people and titles we have are left out.
  180. static webSearch = (&text) => {
  181. ensureIndex()
  182. qw = wordsOf(text)
  183. if (text == null || text.trim().length < 2 || qw.length == 0) { return { rows = [] left = 0 failed = null } }
  184. j = tmdbGet('/search/multi?query=' + encodeURIComponent(text.trim()) + '&include_adult=false&page=1')
  185. if (j.failed != null) { return { rows = [] left = 0 failed = j.failed } }
  186. rows = []
  187. let left = 0
  188. if (j.results != null) {
  189. for (r of j.results) {
  190. let kind = r.media_type
  191. // mission 054: only what TMDB says is not adult (include_adult=false already, and the result's own flag)
  192. if ((kind == 'tv' || kind == 'movie') && r.id != null && hlTypeName(r.id) == 'Number' && adultOf(r) == false) {
  193. if (haveTmdb[tmdbKeyOf(kind == 'tv' ? 'series' : 'movie', r.id)] != null) {
  194. left = left + 1
  195. } else {
  196. let title = kind == 'tv' ? textOr(r.name) : textOr(r.title)
  197. let year = yearOfDate(kind == 'tv' ? r.first_air_date : r.release_date)
  198. if (title != null) {
  199. rows.push({ kind = kind tmdbId = r.id title = title year = year
  200. overview = textOr(r.overview) != null ? r.overview : ''
  201. posterUrl = textOr(r.poster_path) != null ? tmdbImageBase + '/w92' + r.poster_path : '/posters/none' })
  202. }
  203. }
  204. }
  205. }
  206. }
  207. return { rows = rows left = left failed = null }
  208. }
  209. // ---- the import ------------------------------------------------------------------------------------
  210. // "Doctor Who" → "Doctor-Who" (the migrated urlSegment form: letters and digits, everything else one '-'); a slug in use
  211. // gets "-2", "-3", … like the migrated ones ("Raised-by-Wolves-2"); a title with no Latin letter at all → "title-<tmdbId>"
  212. static slugOf = (&title, &tmdbId) => {
  213. out = slugBaseOf(title, 'title-' + tmdbId)
  214. if (slugTaken[out] != true) { return out }
  215. let n = 2
  216. while (slugTaken[out + '-' + n] == true) { n = n + 1 }
  217. return out + '-' + n
  218. }
  219. // our genres by lower-case name (TMDB's "Drama" → ours); TMDB-only names ("Sci-Fi & Fantasy") are left out
  220. static genreIdsOf = (&tmdbGenres) => {
  221. out = []
  222. if (tmdbGenres == null) { return out }
  223. byName = {}
  224. for (g of genresTable.find(null, null)) { if (g.name != null) { byName[('' + g.name).toLowerCase()] = g.id } }
  225. for (tg of tmdbGenres) {
  226. let id = tg.name != null ? byName[('' + tg.name).toLowerCase()] : null
  227. if (id != null && !out.includes(id)) { out.push(id) }
  228. }
  229. return out
  230. }
  231. // `kind` 'tv' | 'movie', `tmdbId` a number → { slug, path (its page, shows.hl titlePath), show, created, sync, details } or { error }. A title we already have (same
  232. // TMDB id and kind) is not added again: its slug is answered.
  233. static importTitle = (&kind, tmdbId) => {
  234. if (kind != 'tv' && kind != 'movie') { return { error = 'unknown kind' } }
  235. // `tid`, not `id`: inside the record literal below an entry named `id` would shadow it for every later entry
  236. tid = toNumber('' + tmdbId)
  237. if (tid == null || tid <= 0 || tid % 1 != 0) { return { error = 'bad TMDB id' } }
  238. ensureIndex()
  239. type = kind == 'tv' ? 'series' : 'movie'
  240. have = haveTmdb[tmdbKeyOf(type, tid)]
  241. if (have != null) {
  242. // tracker#31: a merged tombstone answers its keeper
  243. let s = liveShowOf(showsTable.fetch(have))
  244. if (s != null) { return { slug = s.urlSegment path = titlePath(s) show = s.id created = false } }
  245. }
  246. d = tmdbGet('/' + kind + '/' + tid + '?append_to_response=external_ids,' + (kind == 'tv' ? 'aggregate_credits' : 'credits'))
  247. if (d.failed != null) { return { error = 'TMDB: ' + d.failed } }
  248. let title = kind == 'tv' ? textOr(d.name) : textOr(d.title)
  249. if (title == null) { return { error = 'TMDB has no title for ' + kind + ' ' + tid } }
  250. release = textOr(kind == 'tv' ? d.first_air_date : d.release_date)
  251. year = release != null && release.length >= 4 ? toNumber(release.slice(0, 4)) : null
  252. // tracker#18: `summary` is the creator's own text — TMDB's goes to tmdbSummary only
  253. record = {
  254. id = newId() title = title summary = null tmdbSummary = textOr(d.overview) tmdbId = tid
  255. language = textOr(d.original_language) type = type release = release year = year image = null
  256. urlSegment = slugOf(title, tid) episodesCount = null homepage = textOr(d.homepage) imdbId = null seasonsCount = null
  257. status = textOr(d.status) tagline = textOr(d.tagline) tvdbId = null tvmzId = null genres = genreIdsOf(d.genres)
  258. cast = [] seasons = [] oldId = 'tmdb-' + kind + '-' + tid imported = now() adult = adultOf(d)
  259. }
  260. // tracker.worldapi.org#21: a TV title's TMDB type and kind (set after the literal: an entry named `kind` would shadow the
  261. // parameter for the entries after it)
  262. if (kind == 'tv') {
  263. record.tmdbType = textOr(d.type)
  264. record.kind = kindOfTitle(record.tmdbType, record.genres)
  265. }
  266. // tracker#31: looked up again right before the put — the TMDB request above may have taken a while
  267. again = titleIdByTmdb(type, tid)
  268. if (again != null) {
  269. let s = liveShowOf(showsTable.fetch(again))
  270. if (s != null) { return { slug = s.urlSegment path = titlePath(s) show = s.id created = false } }
  271. }
  272. // mission 060: its short id (shortids.hl — unique across shows and persons)
  273. record.shortId = claimShortId('show', record.id, record.oldId)
  274. if (showsTable.put(record) == null) { return { error = 'could not store the title: ' + showsTable.lastError() } }
  275. indexShow(record)
  276. r = syncShowWith(record.id, d)
  277. persistSync()
  278. // the public lists (catalog.hl, #13) get it at once, like a show the daily sync touched (merge of #13 + #14)
  279. refreshCatalogShow(record.id)
  280. console.log('search import: ' + kind + ' ' + tid + ' "' + title + '" → ' + titlePath(record) + ' (newSeasons=' + r.newSeasons + ' newEpisodes=' + r.newEpisodes + ' posters=' + r.posters + ' ids=' + r.ids + ' requests=' + r.requests + ' errors=' + r.errors + (r.error != '' ? ' ' + r.error : '') + ')')
  281. return { slug = record.urlSegment path = titlePath(record) show = record.id created = true sync = r details = d }
  282. }

Branches

Latest commits

  • 5cb85d75deploy.sh: a backup taken while a background job writes (tar exit 1) is a warning; archive checked with gzip -tmre
  • 9abda75dtracker: report 035mre
  • 12595e47mission 035: theme re-vendored from layouts.worldapi.org 0222f67 (two corner radii: radiusSmall 5px, radiusLarge 10px); the tracker's 18 own radii -> radiusSmall/radiusLarge (--layout-radius is gone); check-theme 0; gates 379/0, 32/0, 53/0, 229/0, 26/0; real copy: every computed radius in {0, 5px, 10px, 50%}mre
  • e7305014tracker: report 032 (art + photos)mre
  • fbb903cctracker#33/#37 (mission 032): a title without a TMDB poster gets its backdrop (w780, posterFromBackdrop, shown 2:3 centre-cropped); movies store runtime, the page shows release date + runtime; art backfill (public titles without poster file / movies without runtime) and person photo backfill (tmdbProfile / photoCheck) as the last start jobs (TRACKER_ART, TRACKER_PHOTOS; off in every gate start); gates 379/0, 32/0, 53/0, 229/0, 26/0, check-theme 0; live copy: art 4977 titles in 42 min (115 posters, 15 backdrops, 4609 runtimes), The Remaining shows its backdrop + 7 minmre
  • 4ecc67b2tracker: STATUS/LOG for the t38 + t40 merge (gates 374/0, 32/0, 53/0, 229/0, 26/0, check-theme 0; live copy checks)mre
  • e5945d2fMerge t40 (tracker#40 curated franchises, Franchises menu, superseded collections) into main: deploy.sh lists all six gates (browser, kinds, franchises, pager, franchiseseed, check-theme); search.hl keeps #37's personPhotoOf + #40's importWithCredits; pager gate runs with TRACKER_FRANCHISE_SEED=0; gates 374/0, 32/0, 53/0, 229/0, 26/0, check-theme 0mre
  • b638e99dMerge t38 (tracker#38 pagination, #35 Returning/Airing label) into main: LOG/STATUS keep both sides; pager gate follows #37's /people (everyone, last updated first: seed updatedAt); gates pager 229/0, browser 374/0mre
  • f8ffa4dbtracker: report 032 + The Remaining + #39mre
  • 7565a863tracker: LOG timemre
  • b10f00c8tracker#39: double episodes — migrated episodes whose TMDB id TMDB replaced are adopted by their number in the sync (old id -> migratedTmdbId); merge.hl step 3 merges each season's doubles at start (keeper: most watches > synced > first; watches moved/parked; tombstones into mergedEpisodes, nothing deleted); tools/count-duplicate-episodes.hl; gate fixture + paths-m039; live copy 850 -> 0 in 64 s; gates 373/0, 32/0, 52/0mre
  • 8f1d4542tracker#40: superseded collections — a TMDB collection timeline whose titles are all in one curated timeline is hidden (supersededBy; kept: own page + editor finder), set by the collection seed when it makes one and by the curated build (lifted when the cover is gone); partly covered ones join that franchise; no second widget (First Contact: only Star Trek — Prime); gate franchiseseed 26/0, browser 365/0, kinds 32/0, franchises 53/0, check-theme 0; real copy 24 supersededmre
  • 9448d643tracker#35 follow-up (mission 033): TVmaze 'Running' is labelled 'Returning', or 'Airing' while a non-special episode of the two newest seasons is released within today +-7 days (data unchanged); gate fixtures Running/Airing/Aired + a special; browser 366/0, kinds 32/0, franchises 52/0, pager 229/0, check-theme 0mre
  • 7d4b293dtracker#40: "Franchises" in the main menu (desktop header after People, phone sidebar) → /franchises, marked on franchise and timeline pages; gates 365/0, 32/0, 53/0, 24/0, check-theme 0mre
  • 8751adb8tracker: report 032mre
  • 9bce1f65tracker mission 032: STATUS gate files + the hour-boundary flakemre
  • 718bfb89tracker#37 (mission 032): /people = everyone, last updated first (updatedAt stamped by the person fill; view built at boot, touched people first at once), photo + name tiles (person colour) with the /movies pagination, /people/<letter> removed; photo = our file, tmdbProfile, a cast/crew entry's profile (in-memory map at boot), else the new 'no photo' placeholder; new cast/crew/created_by people keep tmdbProfile; search people rows with the photo; /settings = the heading only; util.hl sortDesc starts from sorted runs (same result, 105k: 1.6 s -> 0.15 s); gates 369/0, 32/0, 52/0, check-theme 0; README/STATUS/LOGmre
  • 93dfb0batracker#40 (mission 034): the curated franchises — data/franchises.json (17 franchises, 31 timelines, 285 TMDB titles, movies + series, in-universe/release order, 12 TMDB collections attached); lib/franchiseseed.hl + jobs.hl franchiseSeedTick (last start job, imports missing titles via details.hl importWithCredits = the search's Add, one per step paced, then one build; franchiseseed.db: editor changes win, the creator's same-name franchise adopted / timeline left alone, 404 remembered, resumable, idempotent); timeline heads 'N titles · in-universe order' (orderKind) and wrap on a phone; series pages show the widget; new gate tests/franchiseseed.mjs (5th in deploy.sh), the others run with TRACKER_FRANCHISE_SEED=0; gates 365/0, 32/0, 52/0, 24/0, check-theme 0; real copy 196 imported, 0 failed, 7 min, restart unchanged=31mre
  • 03ec792ftracker#38 (mission 033): pagination goes exactly to the clicked page — tilelist read the clicked button's text after pagination.hl's own handler had rebuilt the buttons (real clicks only); now li.current, else the button's own text; new gate tests/pager.mjs (5 lists x 11 pages, 390/1280, real + script clicks, Back/Forward) 229/0; browser 365/0, kinds 32/0, franchises 52/0, check-theme 0mre
  • 96ba683adeploy.sh: a gate without a 'passed,' line (check-theme) no longer ends the scriptmre