tracker
All repositories: gitoria
19.5 KB
// search.hl — THE SEARCH (tracker.worldapi.org#14): our own database first, then "Fetch from web" (TMDB) for titles// we don't have yet, and the import of the one picked. Used by components/search.hl (`/search/<text>`).//// dbSearch(text) our shows, movies and people — from an IN-MEMORY INDEX built once at boot (no table scan per// keystroke): every title/name split into words; each word's first 2 and first 3 characters are// a bucket key → the entries holding such a word. A query looks at ONE bucket (its longest word's// first 3 characters, or 2 for a 2-letter word) and keeps the entries where EVERY query word is// the start of one of the entry's words. A title's year is one of its words ("doctor who 2005").// webSearch(text) GET <TMDB_BASE_URL>/search/multi?query=… (one request; TV + movies only, people dropped) —// the titles we already have (same TMDB id and kind) are left out// importTitle(kind, id) a new show record from GET /tv/<id>?append_to_response=external_ids,aggregate_credits (a movie:// /movie/<id>?…,credits), then tmdbsync.hl `syncShowWith` fills it from that same answer exactly like// the daily sync does (seasons, episodes, poster, external ids); added to the index. The answer goes// back as `details`: the caller (components/search.hl) stores its cast + crew (details.hl applyCredits,// tracker#28 — details.hl imports this file, so it cannot be called from here)//// ADULT TITLES (mission 054): our database lists only titles with `adult == false` (shows.hl isPublicTitle; unknown is// hidden) — `titleOk` holds that per show id, set by indexShow and by `searchTitleChanged` (the adult backfill and the daily// sync call it per title). The web list asks TMDB with include_adult=false AND keeps only results with `adult` false.//// ONE PROCESS PER STORAGE (hybriel#21): the index lives in the app's process and is kept up to date by importTitle; the// daily sync never changes a title. New rows get self-assigned 16-hex ids (hybriel#113), like tmdbsync.hl.import { now } from 'hl:time'import { randomBytes } from 'hl:crypto'import { showsTable, personsTable, genresTable, isPublicTitle } from './shows.hl'import { tmdbGet, tmdbImageBase, syncShowWith, persistSync, textOr, adultOf } from './tmdbsync.hl'import { refreshCatalogShow } from './catalog.hl'// ---- text → words ---------------------------------------------------------------------------------// lower case, apostrophes dropped ("Grey's" → "greys"), every other punctuation mark a word break, the common Latin// accents folded ("Amélie" → "amelie") — the same for the titles and the querystatic apostrophes = ["'", '’', '`', '´']static breaks = ['.', ',', ':', ';', '!', '?', '(', ')', '[', ']', '{', '}', '"', '“', '”', '„', '-', '–', '—', '/', '&', '+', '*', '#', '_', '|', '…', '·', '@', '=', '~', '<', '>']static folds = [['á', 'a'], ['à', 'a'], ['â', 'a'], ['ä', 'a'], ['ã', 'a'], ['å', 'a'], ['æ', 'ae'], ['ç', 'c'], ['é', 'e'], ['è', 'e'], ['ê', 'e'], ['ë', 'e'],['í', 'i'], ['ì', 'i'], ['î', 'i'], ['ï', 'i'], ['ñ', 'n'], ['ó', 'o'], ['ò', 'o'], ['ô', 'o'], ['ö', 'o'], ['õ', 'o'], ['ø', 'o'], ['œ', 'oe'],['ú', 'u'], ['ù', 'u'], ['û', 'u'], ['ü', 'u'], ['ý', 'y'], ['ÿ', 'y'], ['ß', 'ss'], ['š', 's'], ['ž', 'z'], ['č', 'c'], ['ł', 'l']]static normalize = (s) => {if (s == null) { return '' }let t = ('' + s).toLowerCase()for (p of apostrophes) { if (t.includes(p)) { t = t.replaceAll(p, '') } }for (p of breaks) { if (t.includes(p)) { t = t.replaceAll(p, ' ') } }for (f of folds) { if (t.includes(f[0])) { t = t.replaceAll(f[0], f[1]) } }return t}static wordsOf = (s) => {let out = []for (w of normalize(s).split(' ')) { if (w != '') { out.push(w) } }return out}// ---- the index --------------------------------------------------------------------------------------// entries[i] = { kind ('title' | 'person'), id, words, text (the words joined by one space, without the year) }static entries = []static buckets = {}// what we have, by TMDB id: 'series:<id>' / 'movie:<id>' → our show id (the web list leaves these out)static haveTmdb = {}// tracker#18 (the daily sync's TVmaze change list, deltasync.hl): TVmaze id → our show id (kept by indexShow and// searchTitleChanged — the sync fills TVmaze ids)static haveTvmz = {}// every show slug in use (an imported title gets a free one)static slugTaken = {}// show id → true when the title may be listed (not adult, mission 054)static titleOk = {}// tracker#18 (a title renamed by the daily sync): show id → its entry's index; an entry replaced by a newer one is `stale`// (skipped by dbSearch — the buckets only ever grow)static titleEntry = {}static staleEntry = {}static indexStats = { built = false titles = 0 people = 0 keys = 0 ms = 0 }static bucketKeys = (words) => {let seen = {}let out = []for (w of words) {let n = 2while (n <= 3) {if (w.length >= n) {let k = w.slice(0, n)if (seen[k] != true) { seen[k] = true out.push(k) }}n = n + 1}}return out}// `tie` orders equally good matches: lower first (a title: series before movies, then the newer one)static addEntry = (kind, id, name, extra, tie) => {let words = wordsOf(name)if (words.length == 0) { return false }let i = entries.lengthlet all = []for (w of words) { all.push(w) }if (extra != null && extra != '') { all.push(extra) }entries.push({ kind = kind id = id words = all text = words.join(' ') tie = tie })if (kind == 'title') { titleEntry[id] = i }for (k of bucketKeys(all)) {if (buckets[k] == null) { buckets[k] = [i] indexStats.keys = indexStats.keys + 1 } else { buckets[k].push(i) }}return true}static tmdbKeyOf = (type, tmdbId) => { return (type == 'movie' ? 'movie:' : 'series:') + tmdbId }static yearWordOf = (show) => {if (show.year != null) { return '' + show.year }if (show.release != null && hlTypeName(show.release) == 'String' && show.release.length >= 4) { return show.release.slice(0, 4) }return ''}static indexShow = (s) => {if (s.urlSegment != null) { slugTaken[s.urlSegment] = true }if (s.tmdbId != null) { haveTmdb[tmdbKeyOf(s.type, s.tmdbId)] = s.id }if (s.tvmzId != null) { haveTvmz['' + s.tvmzId] = s.id }titleOk[s.id] = isPublicTitle(s)let y = toNumber(yearWordOf(s))if (addEntry('title', s.id, s.title, yearWordOf(s), (s.type == 'movie' ? 10000 : 0) + (y != null ? 3000 - y : 2999))) { indexStats.titles = indexStats.titles + 1 }return null}// tracker#26 (details.hl): a person added from a title's cast — findable by the search at oncestatic indexPerson = (p) => {ensureIndex()if (addEntry('person', p.id, p.name, '', 0)) { indexStats.people = indexStats.people + 1 }return null}// a title's adult flag changed (the backfill / the daily sync): listed from now on, or notstatic searchTitleChanged = (showId) => {let s = showsTable.fetch(showId)if (s != null) { titleOk[s.id] = isPublicTitle(s) }if (s != null && s.tvmzId != null) { haveTvmz['' + s.tvmzId] = s.id }return null}// a title's name changed (the daily sync, tracker#18): its old entry goes stale, a new one is added — found by the new name at oncestatic searchTitleRenamed = (showId) => {if (!indexStats.built) { return null }let s = showsTable.fetch(showId)if (s == null) { return null }let old = titleEntry[s.id]if (old != null && entries[old].text == wordsOf(s.title).join(' ')) { return null }if (old != null) { staleEntry['' + old] = true }indexShow(s)return null}// built once: at boot (project.hl calls it), else by the first searchstatic ensureIndex = () => {if (indexStats.built) { return true }indexStats.built = truelet t0 = now()for (s of showsTable.find(null, null)) { indexShow(s) }for (p of personsTable.find(null, null)) {if (addEntry('person', p.id, p.name, '', 0)) { indexStats.people = indexStats.people + 1 }}indexStats.ms = now() - t0console.log('search index: ' + indexStats.titles + ' titles, ' + indexStats.people + ' people, ' + indexStats.keys + ' keys, ' + indexStats.ms + ' ms')return true}// tracker.worldapi.org#16 (people.hl): our show id for a TMDB title — `type` 'series' | 'movie' — or nullstatic titleIdByTmdb = (type, tmdbId) => {ensureIndex()return haveTmdb[tmdbKeyOf(type, tmdbId)]}// tracker#18: our show id for a TVmaze id, or nullstatic titleIdByTvmaze = (tvmzId) => {ensureIndex()return haveTvmz['' + tvmzId]}// ---- the query as it comes in the address ------------------------------------------------------------// `/search/<text>`: a full page load hands the page the path segment already DECODED by the server (measured: `%2F` splits// the segment → 404, so the page never puts one in the address), a client-side navigation the RAW one — decoded here.// decodeURIComponent stops the request on a malformed one ('%zz', a broken UTF-8 run), so every %-run is checked first; a// bad one is taken literally. (A text that itself looks like an escape, '%41', is decoded twice on a reload — accepted.)static hexDigits = '0123456789abcdef'static hexAt = (s, i) => {if (i + 3 > s.length) { return -1 }let a = hexDigits.indexOf(s.slice(i + 1, i + 2).toLowerCase())let b = hexDigits.indexOf(s.slice(i + 2, i + 3).toLowerCase())if (a < 0 || b < 0) { return -1 }return a * 16 + b}// the bytes of one run of %XX are valid UTF-8 (the rules decodeURIComponent applies)static utf8Ok = (bytes) => {let i = 0while (i < bytes.length) {let b = bytes[i]let need = 0let lo = 128let hi = 191if (b < 128) { need = 0 } else if (b >= 194 && b <= 223) { need = 1 } else if (b >= 224 && b <= 239) {need = 2if (b == 224) { lo = 160 }if (b == 237) { hi = 159 }} else if (b >= 240 && b <= 244) {need = 3if (b == 240) { lo = 144 }if (b == 244) { hi = 143 }} else { return false }if (i + need > bytes.length - 1) { return false }let k = 1while (k <= need) {let c = bytes[i + k]if (k == 1 && (c < lo || c > hi)) { return false }if (k > 1 && (c < 128 || c > 191)) { return false }k = k + 1}i = i + need + 1}return true}static decodableUri = (s) => {let i = 0while (i < s.length) {if (s.slice(i, i + 1) == '%') {let run = []while (i < s.length && s.slice(i, i + 1) == '%') {let v = hexAt(s, i)if (v < 0) { return false }run.push(v)i = i + 3}if (!utf8Ok(run)) { return false }} else {i = i + 1}}return true}static maxQuery = 100static queryOfParam = (raw) => {if (raw == null || hlTypeName(raw) != 'String' || raw == '') { return '' }let t = raw.includes('%') && decodableUri(raw) ? decodeURIComponent(raw) : rawreturn t.length > maxQuery ? t.slice(0, maxQuery) : t}// ---- our database ------------------------------------------------------------------------------------static titleLimit = 30static peopleLimit = 10// lower is better: the whole name (0), the name starts with the query (1), its first word does (2), anything else (3);// then the shorter namestatic scoreOf = (e, qw, qtext) => {if (e.text == qtext) { return 0 }if (e.text.startsWith(qtext)) { return 1 }if (e.words[0].startsWith(qw[0])) { return 2 }return 3}static matchesAll = (e, qw) => {for (q of qw) {let hit = falsefor (w of e.words) { if (!hit && w.startsWith(q)) { hit = true } }if (!hit) { return false }}return true}// keeps the best `limit` of { score, len, tie, i } in order (hand-written: no list sort(), hybriel#1)static insertTop = (top, item, limit) => {let out = []let inserted = falsefor (x of top) {if (!inserted && (item.score < x.score || (item.score == x.score && (item.len < x.len || (item.len == x.len && item.tie < x.tie))))) { out.push(item) inserted = true }if (out.length < limit) { out.push(x) }}if (!inserted && out.length < limit) { out.push(item) }return out}// { text, words, titles: [{ id }], people: [{ id }], titleCount, peopleCount, tooShort }static dbSearch = (text) => {ensureIndex()let out = { text = text titles = [] people = [] titleCount = 0 peopleCount = 0 tooShort = false ms = 0 }let t0 = now()let qw = wordsOf(text)let longest = ''for (w of qw) { if (w.length > longest.length) { longest = w } }if (longest.length < 2) { out.tooShort = true return out }let key = longest.slice(0, longest.length >= 3 ? 3 : 2)let cand = buckets[key]if (cand == null) { return out }let qtext = qw.join(' ')let titles = []let people = []for (i of cand) {let e = entries[i]if ((e.kind != 'title' || (titleOk[e.id] == true && staleEntry['' + i] != true)) && matchesAll(e, qw)) {let item = { score = scoreOf(e, qw, qtext) len = e.text.length tie = e.tie i = i }if (e.kind == 'title') {out.titleCount = out.titleCount + 1titles = insertTop(titles, item, titleLimit)} else {out.peopleCount = out.peopleCount + 1people = insertTop(people, item, peopleLimit)}}}for (t of titles) { out.titles.push({ id = entries[t.i].id }) }for (p of people) { out.people.push({ id = entries[p.i].id }) }out.ms = now() - t0return out}// ---- the web (TMDB) -----------------------------------------------------------------------------------static yearOfDate = (d) => { return textOr(d) != null && d.length >= 4 ? d.slice(0, 4) : '' }// TMDB's search: [{ kind ('tv' | 'movie'), tmdbId, title, year, overview, posterUrl }] of what we DON'T have, or// { failed }. One request (search/multi, page 1 = up to 20 results); people and titles we have are left out.static webSearch = (text) => {ensureIndex()let qw = wordsOf(text)if (text == null || text.trim().length < 2 || qw.length == 0) { return { rows = [] left = 0 failed = null } }let j = tmdbGet('/search/multi?query=' + encodeURIComponent(text.trim()) + '&include_adult=false&page=1')if (j.failed != null) { return { rows = [] left = 0 failed = j.failed } }let rows = []let left = 0if (j.results != null) {for (r of j.results) {let kind = r.media_type// mission 054: only what TMDB says is not adult (include_adult=false already, and the result's own flag)if ((kind == 'tv' || kind == 'movie') && r.id != null && hlTypeName(r.id) == 'Number' && adultOf(r) == false) {if (haveTmdb[tmdbKeyOf(kind == 'tv' ? 'series' : 'movie', r.id)] != null) {left = left + 1} else {let title = kind == 'tv' ? textOr(r.name) : textOr(r.title)let year = yearOfDate(kind == 'tv' ? r.first_air_date : r.release_date)if (title != null) {rows.push({ kind = kind tmdbId = r.id title = title year = yearoverview = textOr(r.overview) != null ? r.overview : ''posterUrl = textOr(r.poster_path) != null ? tmdbImageBase + '/w92' + r.poster_path : '/posters/none' })}}}}}return { rows = rows left = left failed = null }}// ---- the import ------------------------------------------------------------------------------------static slugChars = 'abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789'// "Doctor Who" → "Doctor-Who" (the migrated urlSegment form: letters and digits, everything else one '-'); a slug in use// gets "-2", "-3", … like the migrated ones ("Raised-by-Wolves-2"); a title with no Latin letter at all → "title-<tmdbId>"static slugOf = (title, tmdbId) => {let out = slugBaseOf(title, 'title-' + tmdbId)if (slugTaken[out] != true) { return out }let n = 2while (slugTaken[out + '-' + n] == true) { n = n + 1 }return out + '-' + n}// the slug of a text without the "in use" check — `fallback` when it has no Latin letter at all (people.hl personSlugOf// uses it for new people, tracker#26)static slugBaseOf = (title, fallback) => {let t = normalizeSlugText(title)let out = ''let dash = falselet i = 0while (i < t.length) {let c = t.slice(i, i + 1)if (slugChars.includes(c)) {if (dash && out != '') { out = out + '-' }out = out + cdash = false} else {dash = true}i = i + 1}if (out == '') { out = fallback }if (out.length > 80) { out = out.slice(0, 80) }return out}// the accents folded (keeping the case) and apostrophes dropped, so "Grey's Anatomy" → "Greys-Anatomy"static normalizeSlugText = (s) => {let t = '' + sfor (p of apostrophes) { if (t.includes(p)) { t = t.replaceAll(p, '') } }for (f of folds) {if (t.includes(f[0])) { t = t.replaceAll(f[0], f[1]) }let up = f[0].toUpperCase()if (up != f[0] && t.includes(up)) { t = t.replaceAll(up, f[1].toUpperCase()) }}return t}// our genres by lower-case name (TMDB's "Drama" → ours); TMDB-only names ("Sci-Fi & Fantasy") are left outstatic genreIdsOf = (tmdbGenres) => {let out = []if (tmdbGenres == null) { return out }let byName = {}for (g of genresTable.find(null, null)) { if (g.name != null) { byName[('' + g.name).toLowerCase()] = g.id } }for (tg of tmdbGenres) {let id = tg.name != null ? byName[('' + tg.name).toLowerCase()] : nullif (id != null && !out.includes(id)) { out.push(id) }}return out}// `kind` 'tv' | 'movie', `tmdbId` a number → { slug, show, created, sync, details } or { error }. A title we already have (same// TMDB id and kind) is not added again: its slug is answered.static importTitle = (kind, tmdbId) => {if (kind != 'tv' && kind != 'movie') { return { error = 'unknown kind' } }// `tid`, not `id`: inside the record literal below an entry named `id` would shadow it for every later entrylet tid = toNumber('' + tmdbId)if (tid == null || tid <= 0 || tid % 1 != 0) { return { error = 'bad TMDB id' } }ensureIndex()let type = kind == 'tv' ? 'series' : 'movie'let have = haveTmdb[tmdbKeyOf(type, tid)]if (have != null) {let s = showsTable.fetch(have)if (s != null) { return { slug = s.urlSegment show = s.id created = false } }}let d = tmdbGet('/' + kind + '/' + tid + '?append_to_response=external_ids,' + (kind == 'tv' ? 'aggregate_credits' : 'credits'))if (d.failed != null) { return { error = 'TMDB: ' + d.failed } }let title = kind == 'tv' ? textOr(d.name) : textOr(d.title)if (title == null) { return { error = 'TMDB has no title for ' + kind + ' ' + tid } }let release = textOr(kind == 'tv' ? d.first_air_date : d.release_date)let year = release != null && release.length >= 4 ? toNumber(release.slice(0, 4)) : null// tracker#18: `summary` is the creator's own text — TMDB's goes to tmdbSummary onlylet record = {id = randomBytes(8, 'hex') title = title summary = null tmdbSummary = textOr(d.overview) tmdbId = tidlanguage = textOr(d.original_language) type = type release = release year = year image = nullurlSegment = slugOf(title, tid) episodesCount = null homepage = textOr(d.homepage) imdbId = null seasonsCount = nullstatus = textOr(d.status) tagline = textOr(d.tagline) tvdbId = null tvmzId = null genres = genreIdsOf(d.genres)cast = [] seasons = [] oldId = 'tmdb-' + kind + '-' + tid imported = now() adult = adultOf(d)}if (showsTable.put(record) == null) { return { error = 'could not store the title: ' + showsTable.lastError() } }indexShow(record)let r = syncShowWith(record.id, d)persistSync()// the public lists (catalog.hl, #13) get it at once, like a show the daily sync touched (merge of #13 + #14)refreshCatalogShow(record.id)console.log('search import: ' + kind + ' ' + tid + ' "' + title + '" → /shows/' + record.urlSegment + ' (newSeasons=' + r.newSeasons + ' newEpisodes=' + r.newEpisodes + ' posters=' + r.posters + ' ids=' + r.ids + ' requests=' + r.requests + ' errors=' + r.errors + (r.error != '' ? ' ' + r.error : '') + ')')return { slug = record.urlSegment show = record.id created = true sync = r details = d }}
Branches
- mainmain branch
Latest commits
- f83571c3tracker#18 (mission 067): daily sync by change lists — TMDB /tv|movie/changes (since the stored day, paged) + TVmaze /updates/shows → only our changed titles (followed: full step, unfollowed: light step — changed seasons, no TVmaze), full walk on first run / gap > 14 days / failed list; show record refreshed (title, tmdbSummary, tagline, status, genres …; renamed titles re-indexed); summary = the creator's own text (page: summary > tmdbSummary > tvmazeSummary), one-time clear of copied summaries (9,647 on the live copy); gate 261, tests/realdata-018*.mjs, README + STATUSmre
- 10bb3f93tracker#28 (mission 061): full cast (all seasons, main cast by episodes, guest stars) + crew (created by, directed by, written by, screenplay, story, music) — stored by the details completion, the daily sync, the search import (one details request) and a background credits job (resumes, RSS limit); show page collapsed after 20 with client-side Show all; showBySlug via a slug map; gate 249, tests/realdata-028.mjs, README + STATUSmre
- 25a50bc4tracker#26 (mission 059): titles from a filmography are completed — on open (skeleton, step-wise face showComplete, no reload) and by the in-app details repair (resumes, TMDB-paced, series in parts); cast from TMDB credits; gate 231, tests/realdata-026.mjs, README + STATUSmre
- 1704ec45tracker#17 (mission 057): season caret down/up, skeleton rows while a season loads, sessionless showSeasonEpisodes face (no page re-mount), client-only close; gate 214, tests/realdata-057.mjs, README + STATUSmre
- f2fe3e36mission 056: README + STATUS (merge, fixes, Hybriel 8590df63, real-data check), tests/realdata-056.mjs, tools/check-public-slugs.hlmre
- f40c250emission 056: re-vendor hybriel master 8590df63 (#121, #122); an adult title's page is Not found for non-followers; gate: leave the page before stopping the servermre
- 2b7fdd6cmission 056: signed-out header one row on phones ("Log in", nowrap), backfill skips adult titles' posters, gate checksmre
- 2c53d5efMerge branch 't16-person' (tracker#16 person pages) into main; filmography shows only public titles (054 adult flag), gate race fix (backfill start line)mre
- c171227emission 054: hide adult/unknown titles from the public lists and the search; in-app adult-flag backfill (TMDB details + poster per title, resumes), gate + real-data proofmre
- 139fafd8tracker#16: short bio (4 lines, click = all), real-data check script, README + STATUSmre
- 93be9476tracker#16: person pages /person/<slug> with the filmography fetched from TMDB on the first visit (step by step), gatemre
- 47a3cae6STATUS: mission 053 merge commit idsmre
- dcc5eecaMerge branch 't14-search'mre
- 03edc783Merge branch 't15-tvmaze'mre
- 71b46345tracker#15: numbering check by date or title, placeholder titles in other languages, docs + real-data proofmre
- 6bb2daf1tracker#13: homepage (tiles, intro, latest movies/shows), /shows, /movies/page/N, /my/movies; lists cached in memorymre
- b8bd1157tracker#14: README + STATUS (search, real-data numbers, gate, merge notes)mre
- 65c694a8tracker#14: search — header magnifier, /search/<text> (in-memory word-prefix index over titles + people), Fetch from web (TMDB search/multi, ours left out), Add = import via syncShow; gate +25 checks, real-data scriptmre
- 34f2c15btracker#15: TVmaze merge in the sync (gaps only: new episodes/seasons, empty titles/air dates; numbering check), fake TVmaze episodes + gatemre
- cbdc4ea7tracker#12: link icons TMDB/IMDb/TVDB/TVmaze; sync fills missing ids (TVmaze lookup); movies fetched via /movie/mre