gitoriaLog in with ident

tracker

All repositories: gitoria

ReadmeCodePull requestsReleasesTicketsSettings
Commitc171227ec171227emission 054: hide adult/unknown titles from the public lists and the search; in-app adult-flag backfill (TMDB details + poster per title, resumes), gate + real-data proofmrec171227e/search.hl

17.1 KB

  1. // search.hl — THE SEARCH (tracker.worldapi.org#14): our own database first, then "Fetch from web" (TMDB) for titles
  2. // we don't have yet, and the import of the one picked. Used by components/search.hl (`/search/<text>`).
  3. //
  4. // dbSearch(text) our shows, movies and people — from an IN-MEMORY INDEX built once at boot (no table scan per
  5. // keystroke): every title/name split into words; each word's first 2 and first 3 characters are
  6. // a bucket key → the entries holding such a word. A query looks at ONE bucket (its longest word's
  7. // first 3 characters, or 2 for a 2-letter word) and keeps the entries where EVERY query word is
  8. // the start of one of the entry's words. A title's year is one of its words ("doctor who 2005").
  9. // webSearch(text) GET <TMDB_BASE_URL>/search/multi?query=… (one request; TV + movies only, people dropped) —
  10. // the titles we already have (same TMDB id and kind) are left out
  11. // importTitle(kind, id) a new show record from GET /tv/<id> (or /movie/<id>), then tmdbsync.hl `syncShow` fills it
  12. // exactly like the daily sync does (seasons, episodes, poster, external ids); added to the index
  13. //
  14. // ADULT TITLES (mission 054): our database lists only titles with `adult == false` (shows.hl isPublicTitle; unknown is
  15. // hidden) — `titleOk` holds that per show id, set by indexShow and by `searchTitleChanged` (the adult backfill and the daily
  16. // sync call it per title). The web list asks TMDB with include_adult=false AND keeps only results with `adult` false.
  17. //
  18. // ONE PROCESS PER STORAGE (hybriel#21): the index lives in the app's process and is kept up to date by importTitle; the
  19. // daily sync never changes a title. New rows get self-assigned 16-hex ids (hybriel#113), like tmdbsync.hl.
  20. import { now } from 'hl:time'
  21. import { randomBytes } from 'hl:crypto'
  22. import { showsTable, personsTable, genresTable, isPublicTitle } from './shows.hl'
  23. import { tmdbGet, tmdbImageBase, syncShow, persistSync, textOr, adultOf } from './tmdbsync.hl'
  24. import { refreshCatalogShow } from './catalog.hl'
  25. // ---- text → words ---------------------------------------------------------------------------------
  26. // lower case, apostrophes dropped ("Grey's" → "greys"), every other punctuation mark a word break, the common Latin
  27. // accents folded ("Amélie" → "amelie") — the same for the titles and the query
  28. static apostrophes = ["'", '’', '`', '´']
  29. static breaks = ['.', ',', ':', ';', '!', '?', '(', ')', '[', ']', '{', '}', '"', '“', '”', '„', '-', '–', '—', '/', '&', '+', '*', '#', '_', '|', '…', '·', '@', '=', '~', '<', '>']
  30. static folds = [['á', 'a'], ['à', 'a'], ['â', 'a'], ['ä', 'a'], ['ã', 'a'], ['å', 'a'], ['æ', 'ae'], ['ç', 'c'], ['é', 'e'], ['è', 'e'], ['ê', 'e'], ['ë', 'e'],
  31. ['í', 'i'], ['ì', 'i'], ['î', 'i'], ['ï', 'i'], ['ñ', 'n'], ['ó', 'o'], ['ò', 'o'], ['ô', 'o'], ['ö', 'o'], ['õ', 'o'], ['ø', 'o'], ['œ', 'oe'],
  32. ['ú', 'u'], ['ù', 'u'], ['û', 'u'], ['ü', 'u'], ['ý', 'y'], ['ÿ', 'y'], ['ß', 'ss'], ['š', 's'], ['ž', 'z'], ['č', 'c'], ['ł', 'l']]
  33. static normalize = (s) => {
  34. if (s == null) { return '' }
  35. let t = ('' + s).toLowerCase()
  36. for (p of apostrophes) { if (t.includes(p)) { t = t.replaceAll(p, '') } }
  37. for (p of breaks) { if (t.includes(p)) { t = t.replaceAll(p, ' ') } }
  38. for (f of folds) { if (t.includes(f[0])) { t = t.replaceAll(f[0], f[1]) } }
  39. return t
  40. }
  41. static wordsOf = (s) => {
  42. let out = []
  43. for (w of normalize(s).split(' ')) { if (w != '') { out.push(w) } }
  44. return out
  45. }
  46. // ---- the index --------------------------------------------------------------------------------------
  47. // entries[i] = { kind ('title' | 'person'), id, words, text (the words joined by one space, without the year) }
  48. static entries = []
  49. static buckets = {}
  50. // what we have, by TMDB id: 'series:<id>' / 'movie:<id>' → our show id (the web list leaves these out)
  51. static haveTmdb = {}
  52. // every show slug in use (an imported title gets a free one)
  53. static slugTaken = {}
  54. // show id → true when the title may be listed (not adult, mission 054)
  55. static titleOk = {}
  56. static indexStats = { built = false titles = 0 people = 0 keys = 0 ms = 0 }
  57. static bucketKeys = (words) => {
  58. let seen = {}
  59. let out = []
  60. for (w of words) {
  61. let n = 2
  62. while (n <= 3) {
  63. if (w.length >= n) {
  64. let k = w.slice(0, n)
  65. if (seen[k] != true) { seen[k] = true out.push(k) }
  66. }
  67. n = n + 1
  68. }
  69. }
  70. return out
  71. }
  72. // `tie` orders equally good matches: lower first (a title: series before movies, then the newer one)
  73. static addEntry = (kind, id, name, extra, tie) => {
  74. let words = wordsOf(name)
  75. if (words.length == 0) { return false }
  76. let i = entries.length
  77. let all = []
  78. for (w of words) { all.push(w) }
  79. if (extra != null && extra != '') { all.push(extra) }
  80. entries.push({ kind = kind id = id words = all text = words.join(' ') tie = tie })
  81. for (k of bucketKeys(all)) {
  82. if (buckets[k] == null) { buckets[k] = [i] indexStats.keys = indexStats.keys + 1 } else { buckets[k].push(i) }
  83. }
  84. return true
  85. }
  86. static tmdbKeyOf = (type, tmdbId) => { return (type == 'movie' ? 'movie:' : 'series:') + tmdbId }
  87. static yearWordOf = (show) => {
  88. if (show.year != null) { return '' + show.year }
  89. if (show.release != null && hlTypeName(show.release) == 'String' && show.release.length >= 4) { return show.release.slice(0, 4) }
  90. return ''
  91. }
  92. static indexShow = (s) => {
  93. if (s.urlSegment != null) { slugTaken[s.urlSegment] = true }
  94. if (s.tmdbId != null) { haveTmdb[tmdbKeyOf(s.type, s.tmdbId)] = s.id }
  95. titleOk[s.id] = isPublicTitle(s)
  96. let y = toNumber(yearWordOf(s))
  97. if (addEntry('title', s.id, s.title, yearWordOf(s), (s.type == 'movie' ? 10000 : 0) + (y != null ? 3000 - y : 2999))) { indexStats.titles = indexStats.titles + 1 }
  98. return null
  99. }
  100. // a title's adult flag changed (the backfill / the daily sync): listed from now on, or not
  101. static searchTitleChanged = (showId) => {
  102. let s = showsTable.fetch(showId)
  103. if (s != null) { titleOk[s.id] = isPublicTitle(s) }
  104. return null
  105. }
  106. // built once: at boot (project.hl calls it), else by the first search
  107. static ensureIndex = () => {
  108. if (indexStats.built) { return true }
  109. indexStats.built = true
  110. let t0 = now()
  111. for (s of showsTable.find(null, null)) { indexShow(s) }
  112. for (p of personsTable.find(null, null)) {
  113. if (addEntry('person', p.id, p.name, '', 0)) { indexStats.people = indexStats.people + 1 }
  114. }
  115. indexStats.ms = now() - t0
  116. console.log('search index: ' + indexStats.titles + ' titles, ' + indexStats.people + ' people, ' + indexStats.keys + ' keys, ' + indexStats.ms + ' ms')
  117. return true
  118. }
  119. // ---- the query as it comes in the address ------------------------------------------------------------
  120. // `/search/<text>`: a full page load hands the page the path segment already DECODED by the server (measured: `%2F` splits
  121. // the segment → 404, so the page never puts one in the address), a client-side navigation the RAW one — decoded here.
  122. // decodeURIComponent stops the request on a malformed one ('%zz', a broken UTF-8 run), so every %-run is checked first; a
  123. // bad one is taken literally. (A text that itself looks like an escape, '%41', is decoded twice on a reload — accepted.)
  124. static hexDigits = '0123456789abcdef'
  125. static hexAt = (s, i) => {
  126. if (i + 3 > s.length) { return -1 }
  127. let a = hexDigits.indexOf(s.slice(i + 1, i + 2).toLowerCase())
  128. let b = hexDigits.indexOf(s.slice(i + 2, i + 3).toLowerCase())
  129. if (a < 0 || b < 0) { return -1 }
  130. return a * 16 + b
  131. }
  132. // the bytes of one run of %XX are valid UTF-8 (the rules decodeURIComponent applies)
  133. static utf8Ok = (bytes) => {
  134. let i = 0
  135. while (i < bytes.length) {
  136. let b = bytes[i]
  137. let need = 0
  138. let lo = 128
  139. let hi = 191
  140. if (b < 128) { need = 0 } else if (b >= 194 && b <= 223) { need = 1 } else if (b >= 224 && b <= 239) {
  141. need = 2
  142. if (b == 224) { lo = 160 }
  143. if (b == 237) { hi = 159 }
  144. } else if (b >= 240 && b <= 244) {
  145. need = 3
  146. if (b == 240) { lo = 144 }
  147. if (b == 244) { hi = 143 }
  148. } else { return false }
  149. if (i + need > bytes.length - 1) { return false }
  150. let k = 1
  151. while (k <= need) {
  152. let c = bytes[i + k]
  153. if (k == 1 && (c < lo || c > hi)) { return false }
  154. if (k > 1 && (c < 128 || c > 191)) { return false }
  155. k = k + 1
  156. }
  157. i = i + need + 1
  158. }
  159. return true
  160. }
  161. static decodableUri = (s) => {
  162. let i = 0
  163. while (i < s.length) {
  164. if (s.slice(i, i + 1) == '%') {
  165. let run = []
  166. while (i < s.length && s.slice(i, i + 1) == '%') {
  167. let v = hexAt(s, i)
  168. if (v < 0) { return false }
  169. run.push(v)
  170. i = i + 3
  171. }
  172. if (!utf8Ok(run)) { return false }
  173. } else {
  174. i = i + 1
  175. }
  176. }
  177. return true
  178. }
  179. static maxQuery = 100
  180. static queryOfParam = (raw) => {
  181. if (raw == null || hlTypeName(raw) != 'String' || raw == '') { return '' }
  182. let t = raw.includes('%') && decodableUri(raw) ? decodeURIComponent(raw) : raw
  183. return t.length > maxQuery ? t.slice(0, maxQuery) : t
  184. }
  185. // ---- our database ------------------------------------------------------------------------------------
  186. static titleLimit = 30
  187. static peopleLimit = 10
  188. // lower is better: the whole name (0), the name starts with the query (1), its first word does (2), anything else (3);
  189. // then the shorter name
  190. static scoreOf = (e, qw, qtext) => {
  191. if (e.text == qtext) { return 0 }
  192. if (e.text.startsWith(qtext)) { return 1 }
  193. if (e.words[0].startsWith(qw[0])) { return 2 }
  194. return 3
  195. }
  196. static matchesAll = (e, qw) => {
  197. for (q of qw) {
  198. let hit = false
  199. for (w of e.words) { if (!hit && w.startsWith(q)) { hit = true } }
  200. if (!hit) { return false }
  201. }
  202. return true
  203. }
  204. // keeps the best `limit` of { score, len, tie, i } in order (hand-written: no list sort(), hybriel#1)
  205. static insertTop = (top, item, limit) => {
  206. let out = []
  207. let inserted = false
  208. for (x of top) {
  209. if (!inserted && (item.score < x.score || (item.score == x.score && (item.len < x.len || (item.len == x.len && item.tie < x.tie))))) { out.push(item) inserted = true }
  210. if (out.length < limit) { out.push(x) }
  211. }
  212. if (!inserted && out.length < limit) { out.push(item) }
  213. return out
  214. }
  215. // { text, words, titles: [{ id }], people: [{ id }], titleCount, peopleCount, tooShort }
  216. static dbSearch = (text) => {
  217. ensureIndex()
  218. let out = { text = text titles = [] people = [] titleCount = 0 peopleCount = 0 tooShort = false ms = 0 }
  219. let t0 = now()
  220. let qw = wordsOf(text)
  221. let longest = ''
  222. for (w of qw) { if (w.length > longest.length) { longest = w } }
  223. if (longest.length < 2) { out.tooShort = true return out }
  224. let key = longest.slice(0, longest.length >= 3 ? 3 : 2)
  225. let cand = buckets[key]
  226. if (cand == null) { return out }
  227. let qtext = qw.join(' ')
  228. let titles = []
  229. let people = []
  230. for (i of cand) {
  231. let e = entries[i]
  232. if ((e.kind != 'title' || titleOk[e.id] == true) && matchesAll(e, qw)) {
  233. let item = { score = scoreOf(e, qw, qtext) len = e.text.length tie = e.tie i = i }
  234. if (e.kind == 'title') {
  235. out.titleCount = out.titleCount + 1
  236. titles = insertTop(titles, item, titleLimit)
  237. } else {
  238. out.peopleCount = out.peopleCount + 1
  239. people = insertTop(people, item, peopleLimit)
  240. }
  241. }
  242. }
  243. for (t of titles) { out.titles.push({ id = entries[t.i].id }) }
  244. for (p of people) { out.people.push({ id = entries[p.i].id }) }
  245. out.ms = now() - t0
  246. return out
  247. }
  248. // ---- the web (TMDB) -----------------------------------------------------------------------------------
  249. static yearOfDate = (d) => { return textOr(d) != null && d.length >= 4 ? d.slice(0, 4) : '' }
  250. // TMDB's search: [{ kind ('tv' | 'movie'), tmdbId, title, year, overview, posterUrl }] of what we DON'T have, or
  251. // { failed }. One request (search/multi, page 1 = up to 20 results); people and titles we have are left out.
  252. static webSearch = (text) => {
  253. ensureIndex()
  254. let qw = wordsOf(text)
  255. if (text == null || text.trim().length < 2 || qw.length == 0) { return { rows = [] left = 0 failed = null } }
  256. let j = tmdbGet('/search/multi?query=' + encodeURIComponent(text.trim()) + '&include_adult=false&page=1')
  257. if (j.failed != null) { return { rows = [] left = 0 failed = j.failed } }
  258. let rows = []
  259. let left = 0
  260. if (j.results != null) {
  261. for (r of j.results) {
  262. let kind = r.media_type
  263. // mission 054: only what TMDB says is not adult (include_adult=false already, and the result's own flag)
  264. if ((kind == 'tv' || kind == 'movie') && r.id != null && hlTypeName(r.id) == 'Number' && adultOf(r) == false) {
  265. if (haveTmdb[tmdbKeyOf(kind == 'tv' ? 'series' : 'movie', r.id)] != null) {
  266. left = left + 1
  267. } else {
  268. let title = kind == 'tv' ? textOr(r.name) : textOr(r.title)
  269. let year = yearOfDate(kind == 'tv' ? r.first_air_date : r.release_date)
  270. if (title != null) {
  271. rows.push({ kind = kind tmdbId = r.id title = title year = year
  272. overview = textOr(r.overview) != null ? r.overview : ''
  273. posterUrl = textOr(r.poster_path) != null ? tmdbImageBase + '/w92' + r.poster_path : '/posters/none' })
  274. }
  275. }
  276. }
  277. }
  278. }
  279. return { rows = rows left = left failed = null }
  280. }
  281. // ---- the import ------------------------------------------------------------------------------------
  282. static slugChars = 'abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789'
  283. // "Doctor Who" → "Doctor-Who" (the migrated urlSegment form: letters and digits, everything else one '-'); a slug in use
  284. // gets "-2", "-3", … like the migrated ones ("Raised-by-Wolves-2"); a title with no Latin letter at all → "title-<tmdbId>"
  285. static slugOf = (title, tmdbId) => {
  286. let t = normalizeSlugText(title)
  287. let out = ''
  288. let dash = false
  289. let i = 0
  290. while (i < t.length) {
  291. let c = t.slice(i, i + 1)
  292. if (slugChars.includes(c)) {
  293. if (dash && out != '') { out = out + '-' }
  294. out = out + c
  295. dash = false
  296. } else {
  297. dash = true
  298. }
  299. i = i + 1
  300. }
  301. if (out == '') { out = 'title-' + tmdbId }
  302. if (out.length > 80) { out = out.slice(0, 80) }
  303. if (slugTaken[out] != true) { return out }
  304. let n = 2
  305. while (slugTaken[out + '-' + n] == true) { n = n + 1 }
  306. return out + '-' + n
  307. }
  308. // the accents folded (keeping the case) and apostrophes dropped, so "Grey's Anatomy" → "Greys-Anatomy"
  309. static normalizeSlugText = (s) => {
  310. let t = '' + s
  311. for (p of apostrophes) { if (t.includes(p)) { t = t.replaceAll(p, '') } }
  312. for (f of folds) {
  313. if (t.includes(f[0])) { t = t.replaceAll(f[0], f[1]) }
  314. let up = f[0].toUpperCase()
  315. if (up != f[0] && t.includes(up)) { t = t.replaceAll(up, f[1].toUpperCase()) }
  316. }
  317. return t
  318. }
  319. // our genres by lower-case name (TMDB's "Drama" → ours); TMDB-only names ("Sci-Fi & Fantasy") are left out
  320. static genreIdsOf = (tmdbGenres) => {
  321. let out = []
  322. if (tmdbGenres == null) { return out }
  323. let byName = {}
  324. for (g of genresTable.find(null, null)) { if (g.name != null) { byName[('' + g.name).toLowerCase()] = g.id } }
  325. for (tg of tmdbGenres) {
  326. let id = tg.name != null ? byName[('' + tg.name).toLowerCase()] : null
  327. if (id != null && !out.includes(id)) { out.push(id) }
  328. }
  329. return out
  330. }
  331. // `kind` 'tv' | 'movie', `tmdbId` a number → { slug, show, created, sync } or { error }. A title we already have (same
  332. // TMDB id and kind) is not added again: its slug is answered.
  333. static importTitle = (kind, tmdbId) => {
  334. if (kind != 'tv' && kind != 'movie') { return { error = 'unknown kind' } }
  335. // `tid`, not `id`: inside the record literal below an entry named `id` would shadow it for every later entry
  336. let tid = toNumber('' + tmdbId)
  337. if (tid == null || tid <= 0 || tid % 1 != 0) { return { error = 'bad TMDB id' } }
  338. ensureIndex()
  339. let type = kind == 'tv' ? 'series' : 'movie'
  340. let have = haveTmdb[tmdbKeyOf(type, tid)]
  341. if (have != null) {
  342. let s = showsTable.fetch(have)
  343. if (s != null) { return { slug = s.urlSegment show = s.id created = false } }
  344. }
  345. let d = tmdbGet('/' + kind + '/' + tid)
  346. if (d.failed != null) { return { error = 'TMDB: ' + d.failed } }
  347. let title = kind == 'tv' ? textOr(d.name) : textOr(d.title)
  348. if (title == null) { return { error = 'TMDB has no title for ' + kind + ' ' + tid } }
  349. let release = textOr(kind == 'tv' ? d.first_air_date : d.release_date)
  350. let year = release != null && release.length >= 4 ? toNumber(release.slice(0, 4)) : null
  351. let overview = textOr(d.overview) != null ? d.overview : ''
  352. let record = {
  353. id = randomBytes(8, 'hex') title = title summary = overview tmdbSummary = textOr(d.overview) tmdbId = tid
  354. language = textOr(d.original_language) type = type release = release year = year image = null
  355. urlSegment = slugOf(title, tid) episodesCount = null homepage = textOr(d.homepage) imdbId = null seasonsCount = null
  356. status = textOr(d.status) tagline = textOr(d.tagline) tvdbId = null tvmzId = null genres = genreIdsOf(d.genres)
  357. cast = [] seasons = [] oldId = 'tmdb-' + kind + '-' + tid imported = now() adult = adultOf(d)
  358. }
  359. if (showsTable.put(record) == null) { return { error = 'could not store the title: ' + showsTable.lastError() } }
  360. indexShow(record)
  361. let r = syncShow(record.id)
  362. persistSync()
  363. // the public lists (catalog.hl, #13) get it at once, like a show the daily sync touched (merge of #13 + #14)
  364. refreshCatalogShow(record.id)
  365. console.log('search import: ' + kind + ' ' + tid + ' "' + title + '" → /shows/' + record.urlSegment + ' (newSeasons=' + r.newSeasons + ' newEpisodes=' + r.newEpisodes + ' posters=' + r.posters + ' ids=' + r.ids + ' requests=' + (r.requests + 1) + ' errors=' + r.errors + (r.error != '' ? ' ' + r.error : '') + ')')
  366. return { slug = record.urlSegment show = record.id created = true sync = r }
  367. }

Branches

Latest commits

  • c171227emission 054: hide adult/unknown titles from the public lists and the search; in-app adult-flag backfill (TMDB details + poster per title, resumes), gate + real-data proofmre
  • 47a3cae6STATUS: mission 053 merge commit idsmre
  • dcc5eecaMerge branch 't14-search'mre
  • 03edc783Merge branch 't15-tvmaze'mre
  • 71b46345tracker#15: numbering check by date or title, placeholder titles in other languages, docs + real-data proofmre
  • 6bb2daf1tracker#13: homepage (tiles, intro, latest movies/shows), /shows, /movies/page/N, /my/movies; lists cached in memorymre
  • b8bd1157tracker#14: README + STATUS (search, real-data numbers, gate, merge notes)mre
  • 65c694a8tracker#14: search — header magnifier, /search/<text> (in-memory word-prefix index over titles + people), Fetch from web (TMDB search/multi, ours left out), Add = import via syncShow; gate +25 checks, real-data scriptmre
  • 34f2c15btracker#15: TVmaze merge in the sync (gaps only: new episodes/seasons, empty titles/air dates; numbering check), fake TVmaze episodes + gatemre
  • cbdc4ea7tracker#12: link icons TMDB/IMDb/TVDB/TVmaze; sync fills missing ids (TVmaze lookup); movies fetched via /movie/mre
  • b105bcd8tracker#11: Hybriel master ff51cf46 (checks no longer vanish), mobile-first styles, carets, follow button, sign-in modal, inverted check, orange castmre
  • 31b758aatracker#10: installable app (manifest, service worker, offline shell), own icon + faviconmre
  • 2fa9d997tracker#9: TMDB sync (followed shows: seasons, episodes, posters), tools/sync-tmdb.hl + daily run 04:00 UTC, fake TMDB in gatemre
  • 49e1f61edeploy.sh: back up live storage/.sessions/.env before every deploy (newest 5 kept)mre
  • 54070a4etracker#8: /my/unwatched + /my/schedule (301 from old), S01E01, title (year), 1 episode, watched-set lookup (unwatched 15s -> 1s)mre
  • 3251488atracker#7: /my/shows (followed shows, newest follow first, poster, title, last watched SxxEyy); gate can take screenshots (TRACKER_GATE_SHOTS)mre
  • 44b7d9f9tracker#6: /schedule — upcoming episodes of followed shows, soonest firstmre
  • 91c9fc8ctracker#5: /unwatched — unwatched released episodes of followed shows, newest firstmre
  • a97c0295tracker#4: show page /shows/:slug (header, seasons, episodes, watch checks) + tools/relink-episode-seasons.hlmre
  • 05f407c5tracker: no border on any button except inverted ones (Log out, ident status and identities too); header brand weight 100mre