tracker
All repositories: gitoria
15.2 KB
// lib/search.hl — THE SEARCH (tracker.worldapi.org#14): our own database first, then "Fetch from web" (TMDB) for titles// we don't have yet, and the import of the one picked. Used by components/search.hl (`/search/<text>`).//// dbSearch(text) our shows, movies and people — from an IN-MEMORY INDEX built once at boot (no table scan per// keystroke): every title/name split into words; each word's first 2 and first 3 characters are// a bucket key → the entries holding such a word. A query looks at ONE bucket (its longest word's// first 3 characters, or 2 for a 2-letter word) and keeps the entries where EVERY query word is// the start of one of the entry's words. A title's year is one of its words ("doctor who 2005").// webSearch(text) GET <TMDB_BASE_URL>/search/multi?query=… (one request; TV + movies only, people dropped) —// the titles we already have (same TMDB id and kind) are left out// importTitle(kind, id) a new show record from GET /tv/<id>?append_to_response=external_ids,aggregate_credits (a movie:// /movie/<id>?…,credits), then sync.hl `syncShowWith` fills it from that same answer exactly like// the daily sync does (seasons, episodes, poster, external ids); added to the index. The answer goes// back as `details`: the caller (components/search.hl) stores its cast + crew (credits.hl applyCredits,// tracker#28 — details.hl imports this file, so it cannot be called from here)//// ADULT TITLES (mission 054): our database lists only titles with `adult == false` (shows.hl isPublicTitle; unknown is// hidden) — `titleOk` holds that per show id, set by indexShow and by `searchTitleChanged` (the adult backfill and the daily// sync call it per title). The web list asks TMDB with include_adult=false AND keeps only results with `adult` false.//// ONE PROCESS PER STORAGE (hybriel#21): the index lives in the app's process and is kept up to date by importTitle; the// daily sync never changes a title. New rows get self-assigned 16-hex ids (hybriel#113), like sync.hl.import { now } from 'hl:time'import { textOr, newId } from './util.hl'import { showsTable, personsTable, genresTable, isPublicTitle, kindOfTitle, titlePath, liveShowOf } from './shows.hl'import { tmdbGet, tmdbImageBase, adultOf } from './tmdb.hl'import { syncShowWith, persistSync } from './sync.hl'import { refreshCatalogShow } from './catalog.hl'import { claimShortId } from './shortids.hl'import { wordsOf, bucketKeys, tmdbKeyOf, yearWordOf, scoreOf, matchesAll, insertTop, yearOfDate, slugBaseOf } from './search-helpers.hl'// ---- the index --------------------------------------------------------------------------------------// entries[i] = { kind ('title' | 'person'), id, words, text (the words joined by one space, without the year) }static entries = []static buckets = {}// what we have, by TMDB id: 'series:<id>' / 'movie:<id>' → our show id (the web list leaves these out)static haveTmdb = {}// tracker#18 (the daily sync's TVmaze change list, dailysync.hl): TVmaze id → our show id (kept by indexShow and// searchTitleChanged — the sync fills TVmaze ids)static haveTvmz = {}// every show slug in use (an imported title gets a free one)static slugTaken = {}// show id → true when the title may be listed (not adult, mission 054)static titleOk = {}// tracker#18 (a title renamed by the daily sync): show id → its entry's index; an entry replaced by a newer one is `stale`// (skipped by dbSearch — the buckets only ever grow)static titleEntry = {}static staleEntry = {}static indexStats = { built = false titles = 0 people = 0 keys = 0 ms = 0 }// `tie` orders equally good matches: lower first (a title: series before movies, then the newer one)static addEntry = (kind, id, &name, &extra, tie) => {words = wordsOf(name)if (words.length == 0) { return false }i = entries.lengthall = []for (w of words) { all.push(w) }if (extra != null && extra != '') { all.push(extra) }entries.push({ kind = kind id = id words = all text = words.join(' ') tie = tie })if (kind == 'title') { titleEntry[id] = i }for (k of bucketKeys(all)) {if (buckets[k] == null) { buckets[k] = [i] indexStats.keys = indexStats.keys + 1 } else { buckets[k].push(i) }}return true}static indexShow = (&s) => {if (s.urlSegment != null) { slugTaken[s.urlSegment] = true }// tracker#31: a merged tombstone keeps its slug taken, nothing else (its keeper is the title)if (s.mergedInto != null) {titleOk[s.id] = falsereturn null}if (s.tmdbId != null) { haveTmdb[tmdbKeyOf(s.type, s.tmdbId)] = s.id }if (s.tvmzId != null) { haveTvmz['' + s.tvmzId] = s.id }titleOk[s.id] = isPublicTitle(s)y = toNumber(yearWordOf(s))if (addEntry('title', s.id, s.title, yearWordOf(s), (s.type == 'movie' ? 10000 : 0) + (y != null ? 3000 - y : 2999))) { indexStats.titles = indexStats.titles + 1 }return null}// tracker#26 (details.hl): a person added from a title's cast — findable by the search at oncestatic indexPerson = (&p) => {ensureIndex()if (addEntry('person', p.id, p.name, '', 0)) { indexStats.people = indexStats.people + 1 }return null}// a title's adult flag changed (the backfill / the daily sync): listed from now on, or notstatic searchTitleChanged = (&showId) => {s = showsTable.fetch(showId)if (s != null) { titleOk[s.id] = isPublicTitle(s) }if (s != null && s.tvmzId != null) { haveTvmz['' + s.tvmzId] = s.id }return null}// tracker#31 (merge.hl): a duplicate became a tombstone of `keeperId` — its entry goes stale, its TMDB id (and TVmaze id) now// answer the keeper, so no import or filmography adds the title againstatic searchTitleMerged = (&tombId, &keeperId, &type, &tmdbId) => {ensureIndex()old = titleEntry[tombId]if (old != null) { staleEntry['' + old] = true }titleOk[tombId] = falseif (tmdbId != null) { haveTmdb[tmdbKeyOf(type, tmdbId)] = keeperId }k = showsTable.fetch(keeperId)if (k != null) {if (k.urlSegment != null) { slugTaken[k.urlSegment] = true }if (k.tvmzId != null) { haveTvmz['' + k.tvmzId] = k.id }titleOk[k.id] = isPublicTitle(k)}return null}// a title's name changed (the daily sync, tracker#18): its old entry goes stale, a new one is added — found by the new name at oncestatic searchTitleRenamed = (&showId) => {if (!indexStats.built) { return null }s = showsTable.fetch(showId)if (s == null) { return null }old = titleEntry[s.id]if (old != null && entries[old].text == wordsOf(s.title).join(' ')) { return null }if (old != null) { staleEntry['' + old] = true }indexShow(s)return null}// /people (tracker#32, people.hl peopleOfLetter): every person entry of the index — { id, text (the name's words, lower case) }static indexedPeople = () => {ensureIndex()out = []for (e of entries) { if (e.kind == 'person') { out.push(e) } }return out}// built once: at boot (project.hl calls it), else by the first searchstatic ensureIndex = () => {if (indexStats.built) { return true }indexStats.built = truet0 = now()for (s of showsTable.find(null, null)) { indexShow(s) }for (p of personsTable.find(null, null)) {if (addEntry('person', p.id, p.name, '', 0)) { indexStats.people = indexStats.people + 1 }}indexStats.ms = now() - t0console.log('search index: ' + indexStats.titles + ' titles, ' + indexStats.people + ' people, ' + indexStats.keys + ' keys, ' + indexStats.ms + ' ms')return true}// tracker.worldapi.org#16 (people.hl): our show id for a TMDB title — `type` 'series' | 'movie' — or nullstatic titleIdByTmdb = (&type, &tmdbId) => {ensureIndex()return haveTmdb[tmdbKeyOf(type, tmdbId)]}// tracker#18: our show id for a TVmaze id, or nullstatic titleIdByTvmaze = (&tvmzId) => {ensureIndex()return haveTvmz['' + tvmzId]}// ---- our database ------------------------------------------------------------------------------------static titleLimit = 30static peopleLimit = 10// { text, words, titles: [{ id }], people: [{ id }], titleCount, peopleCount, tooShort }static dbSearch = (text) => {ensureIndex()out = { text = text titles = [] people = [] titleCount = 0 peopleCount = 0 tooShort = false ms = 0 }t0 = now()qw = wordsOf(text)let longest = ''for (w of qw) { if (w.length > longest.length) { longest = w } }if (longest.length < 2) { out.tooShort = true return out }key = longest.slice(0, longest.length >= 3 ? 3 : 2)cand = buckets[key]if (cand == null) { return out }qtext = qw.join(' ')let titles = []let people = []for (i of cand) {let e = entries[i]if ((e.kind != 'title' || (titleOk[e.id] == true && staleEntry['' + i] != true)) && matchesAll(e, qw)) {let item = { score = scoreOf(e, qw, qtext) len = e.text.length tie = e.tie i = i }if (e.kind == 'title') {out.titleCount = out.titleCount + 1titles = insertTop(titles, item, titleLimit)} else {out.peopleCount = out.peopleCount + 1people = insertTop(people, item, peopleLimit)}}}for (t of titles) { out.titles.push({ id = entries[t.i].id }) }for (p of people) { out.people.push({ id = entries[p.i].id }) }out.ms = now() - t0return out}// ---- the web (TMDB) -----------------------------------------------------------------------------------// TMDB's search: [{ kind ('tv' | 'movie'), tmdbId, title, year, overview, posterUrl }] of what we DON'T have, or// { failed }. One request (search/multi, page 1 = up to 20 results); people and titles we have are left out.static webSearch = (&text) => {ensureIndex()qw = wordsOf(text)if (text == null || text.trim().length < 2 || qw.length == 0) { return { rows = [] left = 0 failed = null } }j = tmdbGet('/search/multi?query=' + encodeURIComponent(text.trim()) + '&include_adult=false&page=1')if (j.failed != null) { return { rows = [] left = 0 failed = j.failed } }rows = []let left = 0if (j.results != null) {for (r of j.results) {let kind = r.media_type// mission 054: only what TMDB says is not adult (include_adult=false already, and the result's own flag)if ((kind == 'tv' || kind == 'movie') && r.id != null && hlTypeName(r.id) == 'Number' && adultOf(r) == false) {if (haveTmdb[tmdbKeyOf(kind == 'tv' ? 'series' : 'movie', r.id)] != null) {left = left + 1} else {let title = kind == 'tv' ? textOr(r.name) : textOr(r.title)let year = yearOfDate(kind == 'tv' ? r.first_air_date : r.release_date)if (title != null) {rows.push({ kind = kind tmdbId = r.id title = title year = yearoverview = textOr(r.overview) != null ? r.overview : ''posterUrl = textOr(r.poster_path) != null ? tmdbImageBase + '/w92' + r.poster_path : '/posters/none' })}}}}}return { rows = rows left = left failed = null }}// ---- the import ------------------------------------------------------------------------------------// "Doctor Who" → "Doctor-Who" (the migrated urlSegment form: letters and digits, everything else one '-'); a slug in use// gets "-2", "-3", … like the migrated ones ("Raised-by-Wolves-2"); a title with no Latin letter at all → "title-<tmdbId>"static slugOf = (&title, &tmdbId) => {out = slugBaseOf(title, 'title-' + tmdbId)if (slugTaken[out] != true) { return out }let n = 2while (slugTaken[out + '-' + n] == true) { n = n + 1 }return out + '-' + n}// our genres by lower-case name (TMDB's "Drama" → ours); TMDB-only names ("Sci-Fi & Fantasy") are left outstatic genreIdsOf = (&tmdbGenres) => {out = []if (tmdbGenres == null) { return out }byName = {}for (g of genresTable.find(null, null)) { if (g.name != null) { byName[('' + g.name).toLowerCase()] = g.id } }for (tg of tmdbGenres) {let id = tg.name != null ? byName[('' + tg.name).toLowerCase()] : nullif (id != null && !out.includes(id)) { out.push(id) }}return out}// `kind` 'tv' | 'movie', `tmdbId` a number → { slug, path (its page, shows.hl titlePath), show, created, sync, details } or { error }. A title we already have (same// TMDB id and kind) is not added again: its slug is answered.static importTitle = (&kind, tmdbId) => {if (kind != 'tv' && kind != 'movie') { return { error = 'unknown kind' } }// `tid`, not `id`: inside the record literal below an entry named `id` would shadow it for every later entrytid = toNumber('' + tmdbId)if (tid == null || tid <= 0 || tid % 1 != 0) { return { error = 'bad TMDB id' } }ensureIndex()type = kind == 'tv' ? 'series' : 'movie'have = haveTmdb[tmdbKeyOf(type, tid)]if (have != null) {// tracker#31: a merged tombstone answers its keeperlet s = liveShowOf(showsTable.fetch(have))if (s != null) { return { slug = s.urlSegment path = titlePath(s) show = s.id created = false } }}d = tmdbGet('/' + kind + '/' + tid + '?append_to_response=external_ids,' + (kind == 'tv' ? 'aggregate_credits' : 'credits'))if (d.failed != null) { return { error = 'TMDB: ' + d.failed } }let title = kind == 'tv' ? textOr(d.name) : textOr(d.title)if (title == null) { return { error = 'TMDB has no title for ' + kind + ' ' + tid } }release = textOr(kind == 'tv' ? d.first_air_date : d.release_date)year = release != null && release.length >= 4 ? toNumber(release.slice(0, 4)) : null// tracker#18: `summary` is the creator's own text — TMDB's goes to tmdbSummary onlyrecord = {id = newId() title = title summary = null tmdbSummary = textOr(d.overview) tmdbId = tidlanguage = textOr(d.original_language) type = type release = release year = year image = nullurlSegment = slugOf(title, tid) episodesCount = null homepage = textOr(d.homepage) imdbId = null seasonsCount = nullstatus = textOr(d.status) tagline = textOr(d.tagline) tvdbId = null tvmzId = null genres = genreIdsOf(d.genres)cast = [] seasons = [] oldId = 'tmdb-' + kind + '-' + tid imported = now() adult = adultOf(d)}// tracker.worldapi.org#21: a TV title's TMDB type and kind (set after the literal: an entry named `kind` would shadow the// parameter for the entries after it)if (kind == 'tv') {record.tmdbType = textOr(d.type)record.kind = kindOfTitle(record.tmdbType, record.genres)}// tracker#31: looked up again right before the put — the TMDB request above may have taken a whileagain = titleIdByTmdb(type, tid)if (again != null) {let s = liveShowOf(showsTable.fetch(again))if (s != null) { return { slug = s.urlSegment path = titlePath(s) show = s.id created = false } }}// mission 060: its short id (shortids.hl — unique across shows and persons)record.shortId = claimShortId('show', record.id, record.oldId)if (showsTable.put(record) == null) { return { error = 'could not store the title: ' + showsTable.lastError() } }indexShow(record)r = syncShowWith(record.id, d)persistSync()// the public lists (catalog.hl, #13) get it at once, like a show the daily sync touched (merge of #13 + #14)refreshCatalogShow(record.id)console.log('search import: ' + kind + ' ' + tid + ' "' + title + '" → ' + titlePath(record) + ' (newSeasons=' + r.newSeasons + ' newEpisodes=' + r.newEpisodes + ' posters=' + r.posters + ' ids=' + r.ids + ' requests=' + r.requests + ' errors=' + r.errors + (r.error != '' ? ' ' + r.error : '') + ')')return { slug = record.urlSegment path = titlePath(record) show = record.id created = true sync = r details = d }}
Branches
- mainmain branch
Latest commits
- 9448d643tracker#35 follow-up (mission 033): TVmaze 'Running' is labelled 'Returning', or 'Airing' while a non-special episode of the two newest seasons is released within today +-7 days (data unchanged); gate fixtures Running/Airing/Aired + a special; browser 366/0, kinds 32/0, franchises 52/0, pager 229/0, check-theme 0mre
- 03ec792ftracker#38 (mission 033): pagination goes exactly to the clicked page — tilelist read the clicked button's text after pagination.hl's own handler had rebuilt the buttons (real clicks only); now li.current, else the button's own text; new gate tests/pager.mjs (5 lists x 11 pages, 390/1280, real + script clicks, Back/Forward) 229/0; browser 365/0, kinds 32/0, franchises 52/0, check-theme 0mre
- 96ba683adeploy.sh: a gate without a 'passed,' line (check-theme) no longer ends the scriptmre
- eb3b9205tracker: report 031mre
- 9b5d2e89tracker mission 031: README (What it does, Test: four gates + the #32 checks, Files: theme/, new pages), STATUS (real copy, A/B load, how to repeat, open points), LOGmre
- 39950e4ctracker#32 (mission 031): the WorldAPI theme (theme/ vendored verbatim from layouts.worldapi.org 85b5654; styles.hl inherits it: accent green-dark, type colours 1-6; own base/header rules, row lines, genre-pill and inverted-button frames removed, the season foldable keeps its line; check-theme 21 -> 0, 4th deploy gate; main actions class primary) and the #32 header (theme AppHeader/MainMenu/UserMenu/Sidebar/ContentFirst: desktop brand, search, Series|Shows|Movies|Genres|People, user icon with Unwatched..Settings, Logout; signed out the ident selector, phone the iD icon dropdown; phone menu in the sidebar overlay; marked entry by :has); /find -> /search/<q>, /genres, /people(/<letter>), /settings; main { ContentFirst { slot } } works around the hl:web one-line slot bug; gates 365/0, 32/0, 52/0, check-theme 0mre
- a386dc92tracker: reports 029 + 030mre
- 71e0fd7dtracker missions 029 + 030: README (What it does, Files, gate count), STATUS (real-copy numbers, how to repeat, open points), LOGmre
- d36ea6eatracker#34 + #35 (mission 030): Follow directly under the poster, as wide as the poster (show.hl, styles.hl); the status pill next to a series' title — TVmaze's status (new tvmazeStatus, stored by the sync's TVmaze merge) else TMDB's, TVmaze Ended + TMDB Canceled = Canceled, inverted (filled, dark text, no border), green running / yellow pending / red canceled / muted ended (shows.hl statusOf); the daily delta asks TVmaze's status of an unfollowed series TVmaze's change list names (dailysync.hl syncRunStep, sync.hl syncTvmazeStatus); the status backfill after the details repair (backfill.hl, jobs.hl statusTick; resumable, 550 ms per TVmaze request); gates 354/0, 32/0, 52/0mre
- 7d7d4487tracker#33 (mission 029): reduced titles — every title TMDB's details never went through this app (no detailsAt, no tmdbSync) is incomplete (shows.hl isIncomplete; the old tracker's migrated rows passed #26's test: 5,697 non-adult on the live copy, 691 series without seasons); the repair job does the visibly reduced first (shows.hl missingParts), the page completes one on open; a title TMDB has no poster for (The Remaining) shows the placeholder; tools/count-incomplete.hl; gate fixtures stand for synced titles (tmdbSync), tests/seed-reduced.hl + #33 checks; gates 347/0, 32/0, 52/0mre
- 661c2592tracker: report 028mre
- 27c916fatracker mission 028: README ("Code order", the new file map), STATUS (counts before/after, tests, how to repeat, open), LOGmre
- d924f398tracker mission 028: comments name the new files (sync.hl, dailysync.hl, backfill.hl, credits.hl, jobs.hl, images.hl …); tools/ref-params.py + tools/lambda-audit.py also scan lib/ (they globbed the root only), lambda-audit counts a plain `x = p` alias like `let x = p`mre
- 2e89b968tracker mission 028 (code order) 5/5 let: `let` only where a variable is reassigned — 667 never-reassigned lets became plain declarations (project.hl, lib/, components/, tools/, tests/); kept: 264 in loop bodies (a plain declaration there is 'Cannot reassign' on the 2nd pass), 234 reassigned, 27 whose name is also a member/outer/free name (a plain write would rebind it); tools/let-audit.py decides and fixes (README 'Code order'); tests/realdata-m028.{sh,mjs} = the page-output diff on a real copy; gates 342/0, 32/0, 52/0, real-copy pages identicalmre
- 54796ff2tracker mission 028 (code order) 4/5 thin faces + last copies: the show page's check/follow faces call lib/watches.hl toggleWatched / toggleSeasonWatched (seasonAllWatched moved there) and lib/follows.hl toggleFollowed; both logins (header selector face, /login/callback) share lib/users.hl userOfCode; todayStr/listOf copies in components and the export readers copied into tools/migrate.hl + tools/old-short-ids.hl now once (lib/util.hl, lib/export.hl); gates 342/0, 32/0, 52/0; old-short-ids output byte-identical, migrate output identicalmre
- 06b078e3tracker mission 028 (code order) 3/5 project.hl is the map: config, routes, wiring and a feature → file index (914 → 258 lines); the background jobs (daily sync run, backfills, details repair, credits job, merge, short ids, collection seed) moved unchanged into lib/jobs.hl (a class: their state is reassigned every step, a static cannot be; one instance made after the server), the login callback into lib/users.hl, poster/photo serving into lib/images.hl, the /shows/<slug> rule into lib/shows.hl showsMovedPath; route handlers are thin wrappers; gates 342/0, 32/0, 52/0, real-copy pages identicalmre
- 94716fd2tracker mission 028 (code order) 2/5 util + topics: lib/util.hl holds envOr, storageDir, postersDir, profilesDir, newId, hexDigits, todayStr, dateOr, textOr, hasId, listOr, firstOf, sortDesc once (were copied into up to 5 files); tmdbsync.hl split into tmdb.hl (TMDB/TVmaze requests), sync.hl (one title's sync), sync-helpers.hl, backfill.hl; details.hl split into details.hl, credits.hl, credits-helpers.hl (isIncomplete to shows.hl); search-helpers.hl (words, query, ranking, slugs); collections.hl (the TMDB collection seed, out of franchises.hl); deltasync.hl renamed dailysync.hl; no behaviour change: gates 342/0, 32/0, 52/0, real-copy pages identicalmre
- 186079b0tracker mission 028 (code order) 1/5 move: every root .hl except project.hl into lib/ (styles.hl into components/), import paths only; gates 342/0, 32/0, 52/0; real-copy pages identicalmre
- 4f47f181tracker: report 027mre
- dc1d4be4tracker mission 027: Hybriel master 06617221 vendored (plugin allocator fixes 3a781359 + 413f60e4); real copy RSS through first-start jobs + 400 loads flat ~2.55 GB (190aa11d 2.3 -> 5.6 GB), page times <= 1.1x; gates 342/0, 32/0, 52/0mre