gitoriaLog in with ident

tracker

All repositories: gitoria

ReadmeCodePull requestsReleasesTicketsSettings
Commit2b7fdd6c2b7fdd6cmission 056: signed-out header one row on phones ("Log in", nowrap), backfill skips adult titles' posters, gate checksmre2b7fdd6c/search.hl

17.3 KB

  1. // search.hl — THE SEARCH (tracker.worldapi.org#14): our own database first, then "Fetch from web" (TMDB) for titles
  2. // we don't have yet, and the import of the one picked. Used by components/search.hl (`/search/<text>`).
  3. //
  4. // dbSearch(text) our shows, movies and people — from an IN-MEMORY INDEX built once at boot (no table scan per
  5. // keystroke): every title/name split into words; each word's first 2 and first 3 characters are
  6. // a bucket key → the entries holding such a word. A query looks at ONE bucket (its longest word's
  7. // first 3 characters, or 2 for a 2-letter word) and keeps the entries where EVERY query word is
  8. // the start of one of the entry's words. A title's year is one of its words ("doctor who 2005").
  9. // webSearch(text) GET <TMDB_BASE_URL>/search/multi?query=… (one request; TV + movies only, people dropped) —
  10. // the titles we already have (same TMDB id and kind) are left out
  11. // importTitle(kind, id) a new show record from GET /tv/<id> (or /movie/<id>), then tmdbsync.hl `syncShow` fills it
  12. // exactly like the daily sync does (seasons, episodes, poster, external ids); added to the index
  13. //
  14. // ADULT TITLES (mission 054): our database lists only titles with `adult == false` (shows.hl isPublicTitle; unknown is
  15. // hidden) — `titleOk` holds that per show id, set by indexShow and by `searchTitleChanged` (the adult backfill and the daily
  16. // sync call it per title). The web list asks TMDB with include_adult=false AND keeps only results with `adult` false.
  17. //
  18. // ONE PROCESS PER STORAGE (hybriel#21): the index lives in the app's process and is kept up to date by importTitle; the
  19. // daily sync never changes a title. New rows get self-assigned 16-hex ids (hybriel#113), like tmdbsync.hl.
  20. import { now } from 'hl:time'
  21. import { randomBytes } from 'hl:crypto'
  22. import { showsTable, personsTable, genresTable, isPublicTitle } from './shows.hl'
  23. import { tmdbGet, tmdbImageBase, syncShow, persistSync, textOr, adultOf } from './tmdbsync.hl'
  24. import { refreshCatalogShow } from './catalog.hl'
  25. // ---- text → words ---------------------------------------------------------------------------------
  26. // lower case, apostrophes dropped ("Grey's" → "greys"), every other punctuation mark a word break, the common Latin
  27. // accents folded ("Amélie" → "amelie") — the same for the titles and the query
  28. static apostrophes = ["'", '’', '`', '´']
  29. static breaks = ['.', ',', ':', ';', '!', '?', '(', ')', '[', ']', '{', '}', '"', '“', '”', '„', '-', '–', '—', '/', '&', '+', '*', '#', '_', '|', '…', '·', '@', '=', '~', '<', '>']
  30. static folds = [['á', 'a'], ['à', 'a'], ['â', 'a'], ['ä', 'a'], ['ã', 'a'], ['å', 'a'], ['æ', 'ae'], ['ç', 'c'], ['é', 'e'], ['è', 'e'], ['ê', 'e'], ['ë', 'e'],
  31. ['í', 'i'], ['ì', 'i'], ['î', 'i'], ['ï', 'i'], ['ñ', 'n'], ['ó', 'o'], ['ò', 'o'], ['ô', 'o'], ['ö', 'o'], ['õ', 'o'], ['ø', 'o'], ['œ', 'oe'],
  32. ['ú', 'u'], ['ù', 'u'], ['û', 'u'], ['ü', 'u'], ['ý', 'y'], ['ÿ', 'y'], ['ß', 'ss'], ['š', 's'], ['ž', 'z'], ['č', 'c'], ['ł', 'l']]
  33. static normalize = (s) => {
  34. if (s == null) { return '' }
  35. let t = ('' + s).toLowerCase()
  36. for (p of apostrophes) { if (t.includes(p)) { t = t.replaceAll(p, '') } }
  37. for (p of breaks) { if (t.includes(p)) { t = t.replaceAll(p, ' ') } }
  38. for (f of folds) { if (t.includes(f[0])) { t = t.replaceAll(f[0], f[1]) } }
  39. return t
  40. }
  41. static wordsOf = (s) => {
  42. let out = []
  43. for (w of normalize(s).split(' ')) { if (w != '') { out.push(w) } }
  44. return out
  45. }
  46. // ---- the index --------------------------------------------------------------------------------------
  47. // entries[i] = { kind ('title' | 'person'), id, words, text (the words joined by one space, without the year) }
  48. static entries = []
  49. static buckets = {}
  50. // what we have, by TMDB id: 'series:<id>' / 'movie:<id>' → our show id (the web list leaves these out)
  51. static haveTmdb = {}
  52. // every show slug in use (an imported title gets a free one)
  53. static slugTaken = {}
  54. // show id → true when the title may be listed (not adult, mission 054)
  55. static titleOk = {}
  56. static indexStats = { built = false titles = 0 people = 0 keys = 0 ms = 0 }
  57. static bucketKeys = (words) => {
  58. let seen = {}
  59. let out = []
  60. for (w of words) {
  61. let n = 2
  62. while (n <= 3) {
  63. if (w.length >= n) {
  64. let k = w.slice(0, n)
  65. if (seen[k] != true) { seen[k] = true out.push(k) }
  66. }
  67. n = n + 1
  68. }
  69. }
  70. return out
  71. }
  72. // `tie` orders equally good matches: lower first (a title: series before movies, then the newer one)
  73. static addEntry = (kind, id, name, extra, tie) => {
  74. let words = wordsOf(name)
  75. if (words.length == 0) { return false }
  76. let i = entries.length
  77. let all = []
  78. for (w of words) { all.push(w) }
  79. if (extra != null && extra != '') { all.push(extra) }
  80. entries.push({ kind = kind id = id words = all text = words.join(' ') tie = tie })
  81. for (k of bucketKeys(all)) {
  82. if (buckets[k] == null) { buckets[k] = [i] indexStats.keys = indexStats.keys + 1 } else { buckets[k].push(i) }
  83. }
  84. return true
  85. }
  86. static tmdbKeyOf = (type, tmdbId) => { return (type == 'movie' ? 'movie:' : 'series:') + tmdbId }
  87. static yearWordOf = (show) => {
  88. if (show.year != null) { return '' + show.year }
  89. if (show.release != null && hlTypeName(show.release) == 'String' && show.release.length >= 4) { return show.release.slice(0, 4) }
  90. return ''
  91. }
  92. static indexShow = (s) => {
  93. if (s.urlSegment != null) { slugTaken[s.urlSegment] = true }
  94. if (s.tmdbId != null) { haveTmdb[tmdbKeyOf(s.type, s.tmdbId)] = s.id }
  95. titleOk[s.id] = isPublicTitle(s)
  96. let y = toNumber(yearWordOf(s))
  97. if (addEntry('title', s.id, s.title, yearWordOf(s), (s.type == 'movie' ? 10000 : 0) + (y != null ? 3000 - y : 2999))) { indexStats.titles = indexStats.titles + 1 }
  98. return null
  99. }
  100. // a title's adult flag changed (the backfill / the daily sync): listed from now on, or not
  101. static searchTitleChanged = (showId) => {
  102. let s = showsTable.fetch(showId)
  103. if (s != null) { titleOk[s.id] = isPublicTitle(s) }
  104. return null
  105. }
  106. // built once: at boot (project.hl calls it), else by the first search
  107. static ensureIndex = () => {
  108. if (indexStats.built) { return true }
  109. indexStats.built = true
  110. let t0 = now()
  111. for (s of showsTable.find(null, null)) { indexShow(s) }
  112. for (p of personsTable.find(null, null)) {
  113. if (addEntry('person', p.id, p.name, '', 0)) { indexStats.people = indexStats.people + 1 }
  114. }
  115. indexStats.ms = now() - t0
  116. console.log('search index: ' + indexStats.titles + ' titles, ' + indexStats.people + ' people, ' + indexStats.keys + ' keys, ' + indexStats.ms + ' ms')
  117. return true
  118. }
  119. // tracker.worldapi.org#16 (people.hl): our show id for a TMDB title — `type` 'series' | 'movie' — or null
  120. static titleIdByTmdb = (type, tmdbId) => {
  121. ensureIndex()
  122. return haveTmdb[tmdbKeyOf(type, tmdbId)]
  123. }
  124. // ---- the query as it comes in the address ------------------------------------------------------------
  125. // `/search/<text>`: a full page load hands the page the path segment already DECODED by the server (measured: `%2F` splits
  126. // the segment → 404, so the page never puts one in the address), a client-side navigation the RAW one — decoded here.
  127. // decodeURIComponent stops the request on a malformed one ('%zz', a broken UTF-8 run), so every %-run is checked first; a
  128. // bad one is taken literally. (A text that itself looks like an escape, '%41', is decoded twice on a reload — accepted.)
  129. static hexDigits = '0123456789abcdef'
  130. static hexAt = (s, i) => {
  131. if (i + 3 > s.length) { return -1 }
  132. let a = hexDigits.indexOf(s.slice(i + 1, i + 2).toLowerCase())
  133. let b = hexDigits.indexOf(s.slice(i + 2, i + 3).toLowerCase())
  134. if (a < 0 || b < 0) { return -1 }
  135. return a * 16 + b
  136. }
  137. // the bytes of one run of %XX are valid UTF-8 (the rules decodeURIComponent applies)
  138. static utf8Ok = (bytes) => {
  139. let i = 0
  140. while (i < bytes.length) {
  141. let b = bytes[i]
  142. let need = 0
  143. let lo = 128
  144. let hi = 191
  145. if (b < 128) { need = 0 } else if (b >= 194 && b <= 223) { need = 1 } else if (b >= 224 && b <= 239) {
  146. need = 2
  147. if (b == 224) { lo = 160 }
  148. if (b == 237) { hi = 159 }
  149. } else if (b >= 240 && b <= 244) {
  150. need = 3
  151. if (b == 240) { lo = 144 }
  152. if (b == 244) { hi = 143 }
  153. } else { return false }
  154. if (i + need > bytes.length - 1) { return false }
  155. let k = 1
  156. while (k <= need) {
  157. let c = bytes[i + k]
  158. if (k == 1 && (c < lo || c > hi)) { return false }
  159. if (k > 1 && (c < 128 || c > 191)) { return false }
  160. k = k + 1
  161. }
  162. i = i + need + 1
  163. }
  164. return true
  165. }
  166. static decodableUri = (s) => {
  167. let i = 0
  168. while (i < s.length) {
  169. if (s.slice(i, i + 1) == '%') {
  170. let run = []
  171. while (i < s.length && s.slice(i, i + 1) == '%') {
  172. let v = hexAt(s, i)
  173. if (v < 0) { return false }
  174. run.push(v)
  175. i = i + 3
  176. }
  177. if (!utf8Ok(run)) { return false }
  178. } else {
  179. i = i + 1
  180. }
  181. }
  182. return true
  183. }
  184. static maxQuery = 100
  185. static queryOfParam = (raw) => {
  186. if (raw == null || hlTypeName(raw) != 'String' || raw == '') { return '' }
  187. let t = raw.includes('%') && decodableUri(raw) ? decodeURIComponent(raw) : raw
  188. return t.length > maxQuery ? t.slice(0, maxQuery) : t
  189. }
  190. // ---- our database ------------------------------------------------------------------------------------
  191. static titleLimit = 30
  192. static peopleLimit = 10
  193. // lower is better: the whole name (0), the name starts with the query (1), its first word does (2), anything else (3);
  194. // then the shorter name
  195. static scoreOf = (e, qw, qtext) => {
  196. if (e.text == qtext) { return 0 }
  197. if (e.text.startsWith(qtext)) { return 1 }
  198. if (e.words[0].startsWith(qw[0])) { return 2 }
  199. return 3
  200. }
  201. static matchesAll = (e, qw) => {
  202. for (q of qw) {
  203. let hit = false
  204. for (w of e.words) { if (!hit && w.startsWith(q)) { hit = true } }
  205. if (!hit) { return false }
  206. }
  207. return true
  208. }
  209. // keeps the best `limit` of { score, len, tie, i } in order (hand-written: no list sort(), hybriel#1)
  210. static insertTop = (top, item, limit) => {
  211. let out = []
  212. let inserted = false
  213. for (x of top) {
  214. if (!inserted && (item.score < x.score || (item.score == x.score && (item.len < x.len || (item.len == x.len && item.tie < x.tie))))) { out.push(item) inserted = true }
  215. if (out.length < limit) { out.push(x) }
  216. }
  217. if (!inserted && out.length < limit) { out.push(item) }
  218. return out
  219. }
  220. // { text, words, titles: [{ id }], people: [{ id }], titleCount, peopleCount, tooShort }
  221. static dbSearch = (text) => {
  222. ensureIndex()
  223. let out = { text = text titles = [] people = [] titleCount = 0 peopleCount = 0 tooShort = false ms = 0 }
  224. let t0 = now()
  225. let qw = wordsOf(text)
  226. let longest = ''
  227. for (w of qw) { if (w.length > longest.length) { longest = w } }
  228. if (longest.length < 2) { out.tooShort = true return out }
  229. let key = longest.slice(0, longest.length >= 3 ? 3 : 2)
  230. let cand = buckets[key]
  231. if (cand == null) { return out }
  232. let qtext = qw.join(' ')
  233. let titles = []
  234. let people = []
  235. for (i of cand) {
  236. let e = entries[i]
  237. if ((e.kind != 'title' || titleOk[e.id] == true) && matchesAll(e, qw)) {
  238. let item = { score = scoreOf(e, qw, qtext) len = e.text.length tie = e.tie i = i }
  239. if (e.kind == 'title') {
  240. out.titleCount = out.titleCount + 1
  241. titles = insertTop(titles, item, titleLimit)
  242. } else {
  243. out.peopleCount = out.peopleCount + 1
  244. people = insertTop(people, item, peopleLimit)
  245. }
  246. }
  247. }
  248. for (t of titles) { out.titles.push({ id = entries[t.i].id }) }
  249. for (p of people) { out.people.push({ id = entries[p.i].id }) }
  250. out.ms = now() - t0
  251. return out
  252. }
  253. // ---- the web (TMDB) -----------------------------------------------------------------------------------
  254. static yearOfDate = (d) => { return textOr(d) != null && d.length >= 4 ? d.slice(0, 4) : '' }
  255. // TMDB's search: [{ kind ('tv' | 'movie'), tmdbId, title, year, overview, posterUrl }] of what we DON'T have, or
  256. // { failed }. One request (search/multi, page 1 = up to 20 results); people and titles we have are left out.
  257. static webSearch = (text) => {
  258. ensureIndex()
  259. let qw = wordsOf(text)
  260. if (text == null || text.trim().length < 2 || qw.length == 0) { return { rows = [] left = 0 failed = null } }
  261. let j = tmdbGet('/search/multi?query=' + encodeURIComponent(text.trim()) + '&include_adult=false&page=1')
  262. if (j.failed != null) { return { rows = [] left = 0 failed = j.failed } }
  263. let rows = []
  264. let left = 0
  265. if (j.results != null) {
  266. for (r of j.results) {
  267. let kind = r.media_type
  268. // mission 054: only what TMDB says is not adult (include_adult=false already, and the result's own flag)
  269. if ((kind == 'tv' || kind == 'movie') && r.id != null && hlTypeName(r.id) == 'Number' && adultOf(r) == false) {
  270. if (haveTmdb[tmdbKeyOf(kind == 'tv' ? 'series' : 'movie', r.id)] != null) {
  271. left = left + 1
  272. } else {
  273. let title = kind == 'tv' ? textOr(r.name) : textOr(r.title)
  274. let year = yearOfDate(kind == 'tv' ? r.first_air_date : r.release_date)
  275. if (title != null) {
  276. rows.push({ kind = kind tmdbId = r.id title = title year = year
  277. overview = textOr(r.overview) != null ? r.overview : ''
  278. posterUrl = textOr(r.poster_path) != null ? tmdbImageBase + '/w92' + r.poster_path : '/posters/none' })
  279. }
  280. }
  281. }
  282. }
  283. }
  284. return { rows = rows left = left failed = null }
  285. }
  286. // ---- the import ------------------------------------------------------------------------------------
  287. static slugChars = 'abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789'
  288. // "Doctor Who" → "Doctor-Who" (the migrated urlSegment form: letters and digits, everything else one '-'); a slug in use
  289. // gets "-2", "-3", … like the migrated ones ("Raised-by-Wolves-2"); a title with no Latin letter at all → "title-<tmdbId>"
  290. static slugOf = (title, tmdbId) => {
  291. let t = normalizeSlugText(title)
  292. let out = ''
  293. let dash = false
  294. let i = 0
  295. while (i < t.length) {
  296. let c = t.slice(i, i + 1)
  297. if (slugChars.includes(c)) {
  298. if (dash && out != '') { out = out + '-' }
  299. out = out + c
  300. dash = false
  301. } else {
  302. dash = true
  303. }
  304. i = i + 1
  305. }
  306. if (out == '') { out = 'title-' + tmdbId }
  307. if (out.length > 80) { out = out.slice(0, 80) }
  308. if (slugTaken[out] != true) { return out }
  309. let n = 2
  310. while (slugTaken[out + '-' + n] == true) { n = n + 1 }
  311. return out + '-' + n
  312. }
  313. // the accents folded (keeping the case) and apostrophes dropped, so "Grey's Anatomy" → "Greys-Anatomy"
  314. static normalizeSlugText = (s) => {
  315. let t = '' + s
  316. for (p of apostrophes) { if (t.includes(p)) { t = t.replaceAll(p, '') } }
  317. for (f of folds) {
  318. if (t.includes(f[0])) { t = t.replaceAll(f[0], f[1]) }
  319. let up = f[0].toUpperCase()
  320. if (up != f[0] && t.includes(up)) { t = t.replaceAll(up, f[1].toUpperCase()) }
  321. }
  322. return t
  323. }
  324. // our genres by lower-case name (TMDB's "Drama" → ours); TMDB-only names ("Sci-Fi & Fantasy") are left out
  325. static genreIdsOf = (tmdbGenres) => {
  326. let out = []
  327. if (tmdbGenres == null) { return out }
  328. let byName = {}
  329. for (g of genresTable.find(null, null)) { if (g.name != null) { byName[('' + g.name).toLowerCase()] = g.id } }
  330. for (tg of tmdbGenres) {
  331. let id = tg.name != null ? byName[('' + tg.name).toLowerCase()] : null
  332. if (id != null && !out.includes(id)) { out.push(id) }
  333. }
  334. return out
  335. }
  336. // `kind` 'tv' | 'movie', `tmdbId` a number → { slug, show, created, sync } or { error }. A title we already have (same
  337. // TMDB id and kind) is not added again: its slug is answered.
  338. static importTitle = (kind, tmdbId) => {
  339. if (kind != 'tv' && kind != 'movie') { return { error = 'unknown kind' } }
  340. // `tid`, not `id`: inside the record literal below an entry named `id` would shadow it for every later entry
  341. let tid = toNumber('' + tmdbId)
  342. if (tid == null || tid <= 0 || tid % 1 != 0) { return { error = 'bad TMDB id' } }
  343. ensureIndex()
  344. let type = kind == 'tv' ? 'series' : 'movie'
  345. let have = haveTmdb[tmdbKeyOf(type, tid)]
  346. if (have != null) {
  347. let s = showsTable.fetch(have)
  348. if (s != null) { return { slug = s.urlSegment show = s.id created = false } }
  349. }
  350. let d = tmdbGet('/' + kind + '/' + tid)
  351. if (d.failed != null) { return { error = 'TMDB: ' + d.failed } }
  352. let title = kind == 'tv' ? textOr(d.name) : textOr(d.title)
  353. if (title == null) { return { error = 'TMDB has no title for ' + kind + ' ' + tid } }
  354. let release = textOr(kind == 'tv' ? d.first_air_date : d.release_date)
  355. let year = release != null && release.length >= 4 ? toNumber(release.slice(0, 4)) : null
  356. let overview = textOr(d.overview) != null ? d.overview : ''
  357. let record = {
  358. id = randomBytes(8, 'hex') title = title summary = overview tmdbSummary = textOr(d.overview) tmdbId = tid
  359. language = textOr(d.original_language) type = type release = release year = year image = null
  360. urlSegment = slugOf(title, tid) episodesCount = null homepage = textOr(d.homepage) imdbId = null seasonsCount = null
  361. status = textOr(d.status) tagline = textOr(d.tagline) tvdbId = null tvmzId = null genres = genreIdsOf(d.genres)
  362. cast = [] seasons = [] oldId = 'tmdb-' + kind + '-' + tid imported = now() adult = adultOf(d)
  363. }
  364. if (showsTable.put(record) == null) { return { error = 'could not store the title: ' + showsTable.lastError() } }
  365. indexShow(record)
  366. let r = syncShow(record.id)
  367. persistSync()
  368. // the public lists (catalog.hl, #13) get it at once, like a show the daily sync touched (merge of #13 + #14)
  369. refreshCatalogShow(record.id)
  370. console.log('search import: ' + kind + ' ' + tid + ' "' + title + '" → /shows/' + record.urlSegment + ' (newSeasons=' + r.newSeasons + ' newEpisodes=' + r.newEpisodes + ' posters=' + r.posters + ' ids=' + r.ids + ' requests=' + (r.requests + 1) + ' errors=' + r.errors + (r.error != '' ? ' ' + r.error : '') + ')')
  371. return { slug = record.urlSegment show = record.id created = true sync = r }
  372. }

Branches

Latest commits

  • 2b7fdd6cmission 056: signed-out header one row on phones ("Log in", nowrap), backfill skips adult titles' posters, gate checksmre
  • 2c53d5efMerge branch 't16-person' (tracker#16 person pages) into main; filmography shows only public titles (054 adult flag), gate race fix (backfill start line)mre
  • c171227emission 054: hide adult/unknown titles from the public lists and the search; in-app adult-flag backfill (TMDB details + poster per title, resumes), gate + real-data proofmre
  • 139fafd8tracker#16: short bio (4 lines, click = all), real-data check script, README + STATUSmre
  • 93be9476tracker#16: person pages /person/<slug> with the filmography fetched from TMDB on the first visit (step by step), gatemre
  • 47a3cae6STATUS: mission 053 merge commit idsmre
  • dcc5eecaMerge branch 't14-search'mre
  • 03edc783Merge branch 't15-tvmaze'mre
  • 71b46345tracker#15: numbering check by date or title, placeholder titles in other languages, docs + real-data proofmre
  • 6bb2daf1tracker#13: homepage (tiles, intro, latest movies/shows), /shows, /movies/page/N, /my/movies; lists cached in memorymre
  • b8bd1157tracker#14: README + STATUS (search, real-data numbers, gate, merge notes)mre
  • 65c694a8tracker#14: search — header magnifier, /search/<text> (in-memory word-prefix index over titles + people), Fetch from web (TMDB search/multi, ours left out), Add = import via syncShow; gate +25 checks, real-data scriptmre
  • 34f2c15btracker#15: TVmaze merge in the sync (gaps only: new episodes/seasons, empty titles/air dates; numbering check), fake TVmaze episodes + gatemre
  • cbdc4ea7tracker#12: link icons TMDB/IMDb/TVDB/TVmaze; sync fills missing ids (TVmaze lookup); movies fetched via /movie/mre
  • b105bcd8tracker#11: Hybriel master ff51cf46 (checks no longer vanish), mobile-first styles, carets, follow button, sign-in modal, inverted check, orange castmre
  • 31b758aatracker#10: installable app (manifest, service worker, offline shell), own icon + faviconmre
  • 2fa9d997tracker#9: TMDB sync (followed shows: seasons, episodes, posters), tools/sync-tmdb.hl + daily run 04:00 UTC, fake TMDB in gatemre
  • 49e1f61edeploy.sh: back up live storage/.sessions/.env before every deploy (newest 5 kept)mre
  • 54070a4etracker#8: /my/unwatched + /my/schedule (301 from old), S01E01, title (year), 1 episode, watched-set lookup (unwatched 15s -> 1s)mre
  • 3251488atracker#7: /my/shows (followed shows, newest follow first, poster, title, last watched SxxEyy); gate can take screenshots (TRACKER_GATE_SHOTS)mre