gitoriaLog in with ident

tracker

All repositories: gitoria

ReadmeCodePull requestsReleasesTicketsSettings
Commit65c694a865c694a8tracker#14: search — header magnifier, /search/<text> (in-memory word-prefix index over titles + people), Fetch from web (TMDB search/multi, ours left out), Add = import via syncShow; gate +25 checks, real-data scriptmre65c694a8/search.hl

16.0 KB

  1. // search.hl — THE SEARCH (tracker.worldapi.org#14): our own database first, then "Fetch from web" (TMDB) for titles
  2. // we don't have yet, and the import of the one picked. Used by components/search.hl (`/search/<text>`).
  3. //
  4. // dbSearch(text) our shows, movies and people — from an IN-MEMORY INDEX built once at boot (no table scan per
  5. // keystroke): every title/name split into words; each word's first 2 and first 3 characters are
  6. // a bucket key → the entries holding such a word. A query looks at ONE bucket (its longest word's
  7. // first 3 characters, or 2 for a 2-letter word) and keeps the entries where EVERY query word is
  8. // the start of one of the entry's words. A title's year is one of its words ("doctor who 2005").
  9. // webSearch(text) GET <TMDB_BASE_URL>/search/multi?query=… (one request; TV + movies only, people dropped) —
  10. // the titles we already have (same TMDB id and kind) are left out
  11. // importTitle(kind, id) a new show record from GET /tv/<id> (or /movie/<id>), then tmdbsync.hl `syncShow` fills it
  12. // exactly like the daily sync does (seasons, episodes, poster, external ids); added to the index
  13. //
  14. // ONE PROCESS PER STORAGE (hybriel#21): the index lives in the app's process and is kept up to date by importTitle; the
  15. // daily sync never changes a title. New rows get self-assigned 16-hex ids (hybriel#113), like tmdbsync.hl.
  16. import { now } from 'hl:time'
  17. import { randomBytes } from 'hl:crypto'
  18. import { showsTable, personsTable, genresTable } from './shows.hl'
  19. import { tmdbGet, tmdbImageBase, syncShow, persistSync, textOr } from './tmdbsync.hl'
  20. // ---- text → words ---------------------------------------------------------------------------------
  21. // lower case, apostrophes dropped ("Grey's" → "greys"), every other punctuation mark a word break, the common Latin
  22. // accents folded ("Amélie" → "amelie") — the same for the titles and the query
  23. static apostrophes = ["'", '’', '`', '´']
  24. static breaks = ['.', ',', ':', ';', '!', '?', '(', ')', '[', ']', '{', '}', '"', '“', '”', '„', '-', '–', '—', '/', '&', '+', '*', '#', '_', '|', '…', '·', '@', '=', '~', '<', '>']
  25. static folds = [['á', 'a'], ['à', 'a'], ['â', 'a'], ['ä', 'a'], ['ã', 'a'], ['å', 'a'], ['æ', 'ae'], ['ç', 'c'], ['é', 'e'], ['è', 'e'], ['ê', 'e'], ['ë', 'e'],
  26. ['í', 'i'], ['ì', 'i'], ['î', 'i'], ['ï', 'i'], ['ñ', 'n'], ['ó', 'o'], ['ò', 'o'], ['ô', 'o'], ['ö', 'o'], ['õ', 'o'], ['ø', 'o'], ['œ', 'oe'],
  27. ['ú', 'u'], ['ù', 'u'], ['û', 'u'], ['ü', 'u'], ['ý', 'y'], ['ÿ', 'y'], ['ß', 'ss'], ['š', 's'], ['ž', 'z'], ['č', 'c'], ['ł', 'l']]
  28. static normalize = (s) => {
  29. if (s == null) { return '' }
  30. let t = ('' + s).toLowerCase()
  31. for (p of apostrophes) { if (t.includes(p)) { t = t.replaceAll(p, '') } }
  32. for (p of breaks) { if (t.includes(p)) { t = t.replaceAll(p, ' ') } }
  33. for (f of folds) { if (t.includes(f[0])) { t = t.replaceAll(f[0], f[1]) } }
  34. return t
  35. }
  36. static wordsOf = (s) => {
  37. let out = []
  38. for (w of normalize(s).split(' ')) { if (w != '') { out.push(w) } }
  39. return out
  40. }
  41. // ---- the index --------------------------------------------------------------------------------------
  42. // entries[i] = { kind ('title' | 'person'), id, words, text (the words joined by one space, without the year) }
  43. static entries = []
  44. static buckets = {}
  45. // what we have, by TMDB id: 'series:<id>' / 'movie:<id>' → our show id (the web list leaves these out)
  46. static haveTmdb = {}
  47. // every show slug in use (an imported title gets a free one)
  48. static slugTaken = {}
  49. static indexStats = { built = false titles = 0 people = 0 keys = 0 ms = 0 }
  50. static bucketKeys = (words) => {
  51. let seen = {}
  52. let out = []
  53. for (w of words) {
  54. let n = 2
  55. while (n <= 3) {
  56. if (w.length >= n) {
  57. let k = w.slice(0, n)
  58. if (seen[k] != true) { seen[k] = true out.push(k) }
  59. }
  60. n = n + 1
  61. }
  62. }
  63. return out
  64. }
  65. // `tie` orders equally good matches: lower first (a title: series before movies, then the newer one)
  66. static addEntry = (kind, id, name, extra, tie) => {
  67. let words = wordsOf(name)
  68. if (words.length == 0) { return false }
  69. let i = entries.length
  70. let all = []
  71. for (w of words) { all.push(w) }
  72. if (extra != null && extra != '') { all.push(extra) }
  73. entries.push({ kind = kind id = id words = all text = words.join(' ') tie = tie })
  74. for (k of bucketKeys(all)) {
  75. if (buckets[k] == null) { buckets[k] = [i] indexStats.keys = indexStats.keys + 1 } else { buckets[k].push(i) }
  76. }
  77. return true
  78. }
  79. static tmdbKeyOf = (type, tmdbId) => { return (type == 'movie' ? 'movie:' : 'series:') + tmdbId }
  80. static yearWordOf = (show) => {
  81. if (show.year != null) { return '' + show.year }
  82. if (show.release != null && hlTypeName(show.release) == 'String' && show.release.length >= 4) { return show.release.slice(0, 4) }
  83. return ''
  84. }
  85. static indexShow = (s) => {
  86. if (s.urlSegment != null) { slugTaken[s.urlSegment] = true }
  87. if (s.tmdbId != null) { haveTmdb[tmdbKeyOf(s.type, s.tmdbId)] = s.id }
  88. let y = toNumber(yearWordOf(s))
  89. if (addEntry('title', s.id, s.title, yearWordOf(s), (s.type == 'movie' ? 10000 : 0) + (y != null ? 3000 - y : 2999))) { indexStats.titles = indexStats.titles + 1 }
  90. return null
  91. }
  92. // built once: at boot (project.hl calls it), else by the first search
  93. static ensureIndex = () => {
  94. if (indexStats.built) { return true }
  95. indexStats.built = true
  96. let t0 = now()
  97. for (s of showsTable.find(null, null)) { indexShow(s) }
  98. for (p of personsTable.find(null, null)) {
  99. if (addEntry('person', p.id, p.name, '', 0)) { indexStats.people = indexStats.people + 1 }
  100. }
  101. indexStats.ms = now() - t0
  102. console.log('search index: ' + indexStats.titles + ' titles, ' + indexStats.people + ' people, ' + indexStats.keys + ' keys, ' + indexStats.ms + ' ms')
  103. return true
  104. }
  105. // ---- the query as it comes in the address ------------------------------------------------------------
  106. // `/search/<text>`: a full page load hands the page the path segment already DECODED by the server (measured: `%2F` splits
  107. // the segment → 404, so the page never puts one in the address), a client-side navigation the RAW one — decoded here.
  108. // decodeURIComponent stops the request on a malformed one ('%zz', a broken UTF-8 run), so every %-run is checked first; a
  109. // bad one is taken literally. (A text that itself looks like an escape, '%41', is decoded twice on a reload — accepted.)
  110. static hexDigits = '0123456789abcdef'
  111. static hexAt = (s, i) => {
  112. if (i + 3 > s.length) { return -1 }
  113. let a = hexDigits.indexOf(s.slice(i + 1, i + 2).toLowerCase())
  114. let b = hexDigits.indexOf(s.slice(i + 2, i + 3).toLowerCase())
  115. if (a < 0 || b < 0) { return -1 }
  116. return a * 16 + b
  117. }
  118. // the bytes of one run of %XX are valid UTF-8 (the rules decodeURIComponent applies)
  119. static utf8Ok = (bytes) => {
  120. let i = 0
  121. while (i < bytes.length) {
  122. let b = bytes[i]
  123. let need = 0
  124. let lo = 128
  125. let hi = 191
  126. if (b < 128) { need = 0 } else if (b >= 194 && b <= 223) { need = 1 } else if (b >= 224 && b <= 239) {
  127. need = 2
  128. if (b == 224) { lo = 160 }
  129. if (b == 237) { hi = 159 }
  130. } else if (b >= 240 && b <= 244) {
  131. need = 3
  132. if (b == 240) { lo = 144 }
  133. if (b == 244) { hi = 143 }
  134. } else { return false }
  135. if (i + need > bytes.length - 1) { return false }
  136. let k = 1
  137. while (k <= need) {
  138. let c = bytes[i + k]
  139. if (k == 1 && (c < lo || c > hi)) { return false }
  140. if (k > 1 && (c < 128 || c > 191)) { return false }
  141. k = k + 1
  142. }
  143. i = i + need + 1
  144. }
  145. return true
  146. }
  147. static decodableUri = (s) => {
  148. let i = 0
  149. while (i < s.length) {
  150. if (s.slice(i, i + 1) == '%') {
  151. let run = []
  152. while (i < s.length && s.slice(i, i + 1) == '%') {
  153. let v = hexAt(s, i)
  154. if (v < 0) { return false }
  155. run.push(v)
  156. i = i + 3
  157. }
  158. if (!utf8Ok(run)) { return false }
  159. } else {
  160. i = i + 1
  161. }
  162. }
  163. return true
  164. }
  165. static maxQuery = 100
  166. static queryOfParam = (raw) => {
  167. if (raw == null || hlTypeName(raw) != 'String' || raw == '') { return '' }
  168. let t = raw.includes('%') && decodableUri(raw) ? decodeURIComponent(raw) : raw
  169. return t.length > maxQuery ? t.slice(0, maxQuery) : t
  170. }
  171. // ---- our database ------------------------------------------------------------------------------------
  172. static titleLimit = 30
  173. static peopleLimit = 10
  174. // lower is better: the whole name (0), the name starts with the query (1), its first word does (2), anything else (3);
  175. // then the shorter name
  176. static scoreOf = (e, qw, qtext) => {
  177. if (e.text == qtext) { return 0 }
  178. if (e.text.startsWith(qtext)) { return 1 }
  179. if (e.words[0].startsWith(qw[0])) { return 2 }
  180. return 3
  181. }
  182. static matchesAll = (e, qw) => {
  183. for (q of qw) {
  184. let hit = false
  185. for (w of e.words) { if (!hit && w.startsWith(q)) { hit = true } }
  186. if (!hit) { return false }
  187. }
  188. return true
  189. }
  190. // keeps the best `limit` of { score, len, tie, i } in order (hand-written: no list sort(), hybriel#1)
  191. static insertTop = (top, item, limit) => {
  192. let out = []
  193. let inserted = false
  194. for (x of top) {
  195. if (!inserted && (item.score < x.score || (item.score == x.score && (item.len < x.len || (item.len == x.len && item.tie < x.tie))))) { out.push(item) inserted = true }
  196. if (out.length < limit) { out.push(x) }
  197. }
  198. if (!inserted && out.length < limit) { out.push(item) }
  199. return out
  200. }
  201. // { text, words, titles: [{ id }], people: [{ id }], titleCount, peopleCount, tooShort }
  202. static dbSearch = (text) => {
  203. ensureIndex()
  204. let out = { text = text titles = [] people = [] titleCount = 0 peopleCount = 0 tooShort = false ms = 0 }
  205. let t0 = now()
  206. let qw = wordsOf(text)
  207. let longest = ''
  208. for (w of qw) { if (w.length > longest.length) { longest = w } }
  209. if (longest.length < 2) { out.tooShort = true return out }
  210. let key = longest.slice(0, longest.length >= 3 ? 3 : 2)
  211. let cand = buckets[key]
  212. if (cand == null) { return out }
  213. let qtext = qw.join(' ')
  214. let titles = []
  215. let people = []
  216. for (i of cand) {
  217. let e = entries[i]
  218. if (matchesAll(e, qw)) {
  219. let item = { score = scoreOf(e, qw, qtext) len = e.text.length tie = e.tie i = i }
  220. if (e.kind == 'title') {
  221. out.titleCount = out.titleCount + 1
  222. titles = insertTop(titles, item, titleLimit)
  223. } else {
  224. out.peopleCount = out.peopleCount + 1
  225. people = insertTop(people, item, peopleLimit)
  226. }
  227. }
  228. }
  229. for (t of titles) { out.titles.push({ id = entries[t.i].id }) }
  230. for (p of people) { out.people.push({ id = entries[p.i].id }) }
  231. out.ms = now() - t0
  232. return out
  233. }
  234. // ---- the web (TMDB) -----------------------------------------------------------------------------------
  235. static yearOfDate = (d) => { return textOr(d) != null && d.length >= 4 ? d.slice(0, 4) : '' }
  236. // TMDB's search: [{ kind ('tv' | 'movie'), tmdbId, title, year, overview, posterUrl }] of what we DON'T have, or
  237. // { failed }. One request (search/multi, page 1 = up to 20 results); people and titles we have are left out.
  238. static webSearch = (text) => {
  239. ensureIndex()
  240. let qw = wordsOf(text)
  241. if (text == null || text.trim().length < 2 || qw.length == 0) { return { rows = [] left = 0 failed = null } }
  242. let j = tmdbGet('/search/multi?query=' + encodeURIComponent(text.trim()) + '&include_adult=false&page=1')
  243. if (j.failed != null) { return { rows = [] left = 0 failed = j.failed } }
  244. let rows = []
  245. let left = 0
  246. if (j.results != null) {
  247. for (r of j.results) {
  248. let kind = r.media_type
  249. if ((kind == 'tv' || kind == 'movie') && r.id != null && hlTypeName(r.id) == 'Number') {
  250. if (haveTmdb[tmdbKeyOf(kind == 'tv' ? 'series' : 'movie', r.id)] != null) {
  251. left = left + 1
  252. } else {
  253. let title = kind == 'tv' ? textOr(r.name) : textOr(r.title)
  254. let year = yearOfDate(kind == 'tv' ? r.first_air_date : r.release_date)
  255. if (title != null) {
  256. rows.push({ kind = kind tmdbId = r.id title = title year = year
  257. overview = textOr(r.overview) != null ? r.overview : ''
  258. posterUrl = textOr(r.poster_path) != null ? tmdbImageBase + '/w92' + r.poster_path : '/posters/none' })
  259. }
  260. }
  261. }
  262. }
  263. }
  264. return { rows = rows left = left failed = null }
  265. }
  266. // ---- the import ------------------------------------------------------------------------------------
  267. static slugChars = 'abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789'
  268. // "Doctor Who" → "Doctor-Who" (the migrated urlSegment form: letters and digits, everything else one '-'); a slug in use
  269. // gets "-2", "-3", … like the migrated ones ("Raised-by-Wolves-2"); a title with no Latin letter at all → "title-<tmdbId>"
  270. static slugOf = (title, tmdbId) => {
  271. let t = normalizeSlugText(title)
  272. let out = ''
  273. let dash = false
  274. let i = 0
  275. while (i < t.length) {
  276. let c = t.slice(i, i + 1)
  277. if (slugChars.includes(c)) {
  278. if (dash && out != '') { out = out + '-' }
  279. out = out + c
  280. dash = false
  281. } else {
  282. dash = true
  283. }
  284. i = i + 1
  285. }
  286. if (out == '') { out = 'title-' + tmdbId }
  287. if (out.length > 80) { out = out.slice(0, 80) }
  288. if (slugTaken[out] != true) { return out }
  289. let n = 2
  290. while (slugTaken[out + '-' + n] == true) { n = n + 1 }
  291. return out + '-' + n
  292. }
  293. // the accents folded (keeping the case) and apostrophes dropped, so "Grey's Anatomy" → "Greys-Anatomy"
  294. static normalizeSlugText = (s) => {
  295. let t = '' + s
  296. for (p of apostrophes) { if (t.includes(p)) { t = t.replaceAll(p, '') } }
  297. for (f of folds) {
  298. if (t.includes(f[0])) { t = t.replaceAll(f[0], f[1]) }
  299. let up = f[0].toUpperCase()
  300. if (up != f[0] && t.includes(up)) { t = t.replaceAll(up, f[1].toUpperCase()) }
  301. }
  302. return t
  303. }
  304. // our genres by lower-case name (TMDB's "Drama" → ours); TMDB-only names ("Sci-Fi & Fantasy") are left out
  305. static genreIdsOf = (tmdbGenres) => {
  306. let out = []
  307. if (tmdbGenres == null) { return out }
  308. let byName = {}
  309. for (g of genresTable.find(null, null)) { if (g.name != null) { byName[('' + g.name).toLowerCase()] = g.id } }
  310. for (tg of tmdbGenres) {
  311. let id = tg.name != null ? byName[('' + tg.name).toLowerCase()] : null
  312. if (id != null && !out.includes(id)) { out.push(id) }
  313. }
  314. return out
  315. }
  316. // `kind` 'tv' | 'movie', `tmdbId` a number → { slug, show, created, sync } or { error }. A title we already have (same
  317. // TMDB id and kind) is not added again: its slug is answered.
  318. static importTitle = (kind, tmdbId) => {
  319. if (kind != 'tv' && kind != 'movie') { return { error = 'unknown kind' } }
  320. // `tid`, not `id`: inside the record literal below an entry named `id` would shadow it for every later entry
  321. let tid = toNumber('' + tmdbId)
  322. if (tid == null || tid <= 0 || tid % 1 != 0) { return { error = 'bad TMDB id' } }
  323. ensureIndex()
  324. let type = kind == 'tv' ? 'series' : 'movie'
  325. let have = haveTmdb[tmdbKeyOf(type, tid)]
  326. if (have != null) {
  327. let s = showsTable.fetch(have)
  328. if (s != null) { return { slug = s.urlSegment show = s.id created = false } }
  329. }
  330. let d = tmdbGet('/' + kind + '/' + tid)
  331. if (d.failed != null) { return { error = 'TMDB: ' + d.failed } }
  332. let title = kind == 'tv' ? textOr(d.name) : textOr(d.title)
  333. if (title == null) { return { error = 'TMDB has no title for ' + kind + ' ' + tid } }
  334. let release = textOr(kind == 'tv' ? d.first_air_date : d.release_date)
  335. let year = release != null && release.length >= 4 ? toNumber(release.slice(0, 4)) : null
  336. let overview = textOr(d.overview) != null ? d.overview : ''
  337. let record = {
  338. id = randomBytes(8, 'hex') title = title summary = overview tmdbSummary = textOr(d.overview) tmdbId = tid
  339. language = textOr(d.original_language) type = type release = release year = year image = null
  340. urlSegment = slugOf(title, tid) episodesCount = null homepage = textOr(d.homepage) imdbId = null seasonsCount = null
  341. status = textOr(d.status) tagline = textOr(d.tagline) tvdbId = null tvmzId = null genres = genreIdsOf(d.genres)
  342. cast = [] seasons = [] oldId = 'tmdb-' + kind + '-' + tid imported = now()
  343. }
  344. if (showsTable.put(record) == null) { return { error = 'could not store the title: ' + showsTable.lastError() } }
  345. indexShow(record)
  346. let r = syncShow(record.id)
  347. persistSync()
  348. console.log('search import: ' + kind + ' ' + tid + ' "' + title + '" → /shows/' + record.urlSegment + ' (newSeasons=' + r.newSeasons + ' newEpisodes=' + r.newEpisodes + ' posters=' + r.posters + ' ids=' + r.ids + ' requests=' + (r.requests + 1) + ' errors=' + r.errors + (r.error != '' ? ' ' + r.error : '') + ')')
  349. return { slug = record.urlSegment show = record.id created = true sync = r }
  350. }

Branches

Latest commits

  • 65c694a8tracker#14: search — header magnifier, /search/<text> (in-memory word-prefix index over titles + people), Fetch from web (TMDB search/multi, ours left out), Add = import via syncShow; gate +25 checks, real-data scriptmre
  • cbdc4ea7tracker#12: link icons TMDB/IMDb/TVDB/TVmaze; sync fills missing ids (TVmaze lookup); movies fetched via /movie/mre
  • b105bcd8tracker#11: Hybriel master ff51cf46 (checks no longer vanish), mobile-first styles, carets, follow button, sign-in modal, inverted check, orange castmre
  • 31b758aatracker#10: installable app (manifest, service worker, offline shell), own icon + faviconmre
  • 2fa9d997tracker#9: TMDB sync (followed shows: seasons, episodes, posters), tools/sync-tmdb.hl + daily run 04:00 UTC, fake TMDB in gatemre
  • 49e1f61edeploy.sh: back up live storage/.sessions/.env before every deploy (newest 5 kept)mre
  • 54070a4etracker#8: /my/unwatched + /my/schedule (301 from old), S01E01, title (year), 1 episode, watched-set lookup (unwatched 15s -> 1s)mre
  • 3251488atracker#7: /my/shows (followed shows, newest follow first, poster, title, last watched SxxEyy); gate can take screenshots (TRACKER_GATE_SHOTS)mre
  • 44b7d9f9tracker#6: /schedule — upcoming episodes of followed shows, soonest firstmre
  • 91c9fc8ctracker#5: /unwatched — unwatched released episodes of followed shows, newest firstmre
  • a97c0295tracker#4: show page /shows/:slug (header, seasons, episodes, watch checks) + tools/relink-episode-seasons.hlmre
  • 05f407c5tracker: no border on any button except inverted ones (Log out, ident status and identities too); header brand weight 100mre
  • 17375427tracker#2: tools/migrate.hl + tools/verify.hl — old MongoDB data into mpackdb with new idsmre
  • 31aac936tracker#3: no border on the header and on filled buttons; inverted buttons keep theirsmre
  • 2ad9d29ctracker#1: login exactly like calendar (identity selector in the header, empty homepage)mre
  • 3691e176tracker#1: empty tracker with the ident login (state of 2026-09-27)mre