gitoriaLog in with ident

tracker

All repositories: gitoria

ReadmeCodePull requestsReleasesTicketsSettings
Commitf83571c3f83571c3tracker#18 (mission 067): daily sync by change lists — TMDB /tv|movie/changes (since the stored day, paged) + TVmaze /updates/shows → only our changed titles (followed: full step, unfollowed: light step — changed seasons, no TVmaze), full walk on first run / gap > 14 days / failed list; show record refreshed (title, tmdbSummary, tagline, status, genres …; renamed titles re-indexed); summary = the creator's own text (page: summary > tmdbSummary > tvmazeSummary), one-time clear of copied summaries (9,647 on the live copy); gate 261, tests/realdata-018*.mjs, README + STATUSmref83571c3/people.hl

17.4 KB

  1. // people.hl — THE PERSON PAGE's data (tracker.worldapi.org#16, components/person.hl `/person/<slug>`): a person by slug,
  2. // their facts and filmography, and the ONE-TIME fetch of their whole filmography from TMDB on the first visit. Creator:
  3. // "the movies come from artist scrolls. when one clicks on an artist that has no dataset yet the old version fetched that
  4. // artist and its entire bibilogrqphy so its profile was complete".
  5. //
  6. // personBySlug(slug) the person (slug = `urlSegment`) — a slug → id map built once (15k people, one scan)
  7. // personInfoOf(p) name, photo URL, born/died lines, bio
  8. // filmographyOf(p) their credits (`p.shows`): poster, title, year, role, kind — newest first, adult titles hidden
  9. // personFillStep(id) ONE step of the fetch, called again and again by the page until it answers `done`:
  10. // 1st call: GET <TMDB>/person/<tmdbId>?append_to_response=combined_credits (ONE request: details +
  11. // every cast/crew credit) + the photo (w185, like the posters); the person's facts are updated;
  12. // every next call stores the next `fillBatch` titles; the last one links the credits and marks
  13. // the person complete (`filmographyAt`, ms). hl:web serves one request at a time — between two
  14. // steps every other request is served, so a person with 300 credits never holds the server.
  15. //
  16. // COMPLETE = `filmographyAt` set: never fetched again (the daily sync may refresh followed people later — not built).
  17. // Migrated people are NOT complete (153 of them carry the old tracker's credits — kept, the fetch adds the rest).
  18. // A title we don't have → a MINIMAL show record (title, year, type, tmdbId, `adult`, overview, genres by TMDB id, language;
  19. // no seasons, no poster file; `minimal = true` — details.hl completes it when it is opened or by the repair job, tracker#26), shaped like search.hl importTitle's
  20. // (`oldId tmdb-<tv|movie>-<id>`, a free slug), indexed for the search and put into the public lists (catalog.hl).
  21. // The person's credits (`shows`, the migrated field): [{ show, character, posterPath, adult }] — `character` is the role
  22. // text (the characters, then crew jobs like "Director", " / " between them); `posterPath` = TMDB's, so the page can show
  23. // TMDB's small poster (w92, like search's web rows) until we have the file. ADULT (TMDB's flag, mission 054): a row is
  24. // shown only for a PUBLIC title (shows.hl `isPublicTitle`: `adult == false`, unknown hidden like the public lists) whose
  25. // credit is not adult either (mission 056 merge).
  26. // ONE PROCESS PER STORAGE (hybriel#21); new rows get self-assigned 16-hex ids (hybriel#113), like tmdbsync.hl.
  27. import { now } from 'hl:time'
  28. import { fetch } from 'hl:fetch'
  29. import { randomBytes } from 'hl:crypto'
  30. import { writeFile, mkDir, exists } from 'hl:fs'
  31. import { personsTable, showsTable, genresTable, storageDir, showById, showYear, posterUrlOf, isPublicTitle } from './shows.hl'
  32. import { tmdbGet, tmdbImageBase, textOr } from './tmdbsync.hl'
  33. import { titleIdByTmdb, indexShow, slugOf, slugBaseOf, indexPerson } from './search.hl'
  34. import { refreshCatalogShow, sortDesc } from './catalog.hl'
  35. static profilesDir = storageDir + '/profiles'
  36. // titles stored per step (real data: see STATUS.md ticket #16 for the per-step time)
  37. static fillBatch = 40
  38. static listOr = (v) => { return v == null ? [] : v }
  39. // ---- the person -------------------------------------------------------------------------------------------------
  40. static slugIndex = { built = false ids = {} }
  41. static ensureSlugIndex = () => {
  42. if (!slugIndex.built) {
  43. let m = {}
  44. for (p of personsTable.find(null, null)) { if (p.urlSegment != null) { m[p.urlSegment] = p.id } }
  45. slugIndex.ids = m
  46. slugIndex.built = true
  47. }
  48. return true
  49. }
  50. static personBySlug = (slug) => {
  51. if (slug == null || hlTypeName(slug) != 'String' || slug == '') { return null }
  52. ensureSlugIndex()
  53. let id = slugIndex.ids[slug]
  54. return id == null ? null : personsTable.fetch(id)
  55. }
  56. static isComplete = (p) => { return p.filmographyAt != null }
  57. // the page asks TMDB for this person: not complete yet, and TMDB knows them
  58. static needsFill = (p) => { return !isComplete(p) && p.tmdbId != null }
  59. // the photo: /profiles/<oldId>.<ext> (project.hl profileRoute) — only when the file is there (the migrated `image` = 'jpg'
  60. // of 129 people never came with a file)
  61. static profileName = (p) => {
  62. if (p.oldId == null || p.image == null || p.image == '') { return null }
  63. return p.oldId + '.' + p.image
  64. }
  65. static photoUrlOf = (p) => {
  66. let name = profileName(p)
  67. return name != null && exists(profilesDir + '/' + name) ? '/profiles/' + name : ''
  68. }
  69. static monthNames = ['January', 'February', 'March', 'April', 'May', 'June', 'July', 'August', 'September', 'October', 'November', 'December']
  70. // '1986-03-04' → '4 March 1986'; a bare year or anything else as it is
  71. static dateText = (d) => {
  72. let t = textOr(d)
  73. if (t == null) { return null }
  74. if (t.length < 10) { return t }
  75. let m = toNumber(t.slice(5, 7))
  76. let day = toNumber(t.slice(8, 10))
  77. if (m == null || day == null || m < 1 || m > 12) { return t }
  78. return day + ' ' + monthNames[m - 1] + ' ' + t.slice(0, 4)
  79. }
  80. // { name, photoUrl, born, died, bio } — every field a plain string ('' = none)
  81. static personInfoOf = (p) => {
  82. let born = dateText(p.birthDate)
  83. let place = textOr(p.birthPlace)
  84. let bornText = ''
  85. if (born != null && place != null) { bornText = 'Born ' + born + ' in ' + place } else if (born != null) { bornText = 'Born ' + born } else if (place != null) { bornText = 'Born in ' + place }
  86. let died = dateText(p.deathDate)
  87. return { name = p.name != null ? p.name : '' photoUrl = photoUrlOf(p) born = bornText died = died != null ? 'Died ' + died : ''
  88. bio = textOr(p.biography) != null ? p.biography : '' }
  89. }
  90. // ---- the filmography --------------------------------------------------------------------------------------------
  91. // the credits, one row per title (a title twice — the migrated credit and TMDB's — is ONE row, roles merged), newest
  92. // first by the title's release ('YYYY-MM-DD', else its year), undated last; only public titles (no adult/unknown)
  93. static filmographyOf = (p) => {
  94. let byShow = {}
  95. let order = []
  96. for (c of listOr(p.shows)) {
  97. if (c != null && c.show != null) {
  98. let role = textOr(c.character)
  99. let prev = byShow[c.show]
  100. if (prev == null) {
  101. byShow[c.show] = { role = role != null ? role : '' posterPath = textOr(c.posterPath) adult = c.adult == true }
  102. order.push(c.show)
  103. } else {
  104. let x = prev
  105. if (role != null && !(' / ' + x.role + ' / ').includes(' / ' + role + ' / ')) { x.role = x.role == '' ? role : x.role + ' / ' + role }
  106. if (x.posterPath == null) { x.posterPath = textOr(c.posterPath) }
  107. if (c.adult == true) { x.adult = true }
  108. byShow[c.show] = x
  109. }
  110. }
  111. }
  112. let list = []
  113. for (sid of order) {
  114. let c = byShow[sid]
  115. let s = showById(sid)
  116. if (isPublicTitle(s) && !c.adult) {
  117. let y = showYear(s)
  118. let release = textOr(s.release)
  119. let key = release != null && release.length >= 4 ? release : (y != null ? y : '')
  120. let poster = posterUrlOf(s)
  121. if (poster == '/posters/none' && c.posterPath != null) { poster = tmdbImageBase + '/w92' + c.posterPath }
  122. list.push({ key = key id = s.id title = s.title year = y != null ? y : '' href = '/shows/' + s.urlSegment posterUrl = poster
  123. role = c.role typeLabel = s.type == 'movie' ? 'Movie' : 'Series' })
  124. }
  125. }
  126. let out = []
  127. let undated = []
  128. for (r of sortDesc(list)) { if (r.key == '') { undated.push(r) } else { out.push(r) } }
  129. for (r of undated) { out.push(r) }
  130. return out
  131. }
  132. // ---- the fetch ------------------------------------------------------------------------------------------------
  133. // person id → { items, pos, credits, added, requests, started, steps } while a fetch is under way (in memory: a restart
  134. // loses it — the titles stored so far are found again by TMDB id, nothing is duplicated)
  135. static jobs = {}
  136. // TMDB genre id → ours (our genres carry TMDB's id)
  137. static genreMap = { built = false ids = {} }
  138. static genreIdsOfTmdb = (tmdbIds) => {
  139. if (!genreMap.built) {
  140. let m = {}
  141. for (g of genresTable.find(null, null)) { if (g.tmdbId != null) { m['' + g.tmdbId] = g.id } }
  142. genreMap.ids = m
  143. genreMap.built = true
  144. }
  145. let out = []
  146. for (t of listOr(tmdbIds)) {
  147. let id = genreMap.ids['' + t]
  148. if (id != null && !out.includes(id)) { out.push(id) }
  149. }
  150. return out
  151. }
  152. // TMDB's combined_credits → one item per title (cast AND crew of the same title merged), TV + movies only
  153. static creditItems = (cc) => {
  154. let byKey = {}
  155. let order = []
  156. let all = []
  157. if (cc != null) {
  158. for (c of listOr(cc.cast)) { all.push({ credit = c role = textOr(c.character) }) }
  159. for (c of listOr(cc.crew)) { all.push({ credit = c role = textOr(c.job) }) }
  160. }
  161. for (a of all) {
  162. let c = a.credit
  163. let kind = c.media_type
  164. if ((kind == 'tv' || kind == 'movie') && c.id != null && hlTypeName(c.id) == 'Number') {
  165. let k = kind + ':' + c.id
  166. let known = byKey[k]
  167. if (known == null) {
  168. let name = kind == 'tv' ? textOr(c.name) : textOr(c.title)
  169. if (name != null) {
  170. byKey[k] = { kind = kind tmdbId = c.id title = name release = textOr(kind == 'tv' ? c.first_air_date : c.release_date)
  171. overview = textOr(c.overview) language = textOr(c.original_language) genreIds = listOr(c.genre_ids)
  172. posterPath = textOr(c.poster_path) adult = c.adult == true role = a.role != null ? a.role : '' }
  173. order.push(k)
  174. }
  175. } else if (a.role != null && !(' / ' + known.role + ' / ').includes(' / ' + a.role + ' / ')) {
  176. let x = known
  177. x.role = x.role == '' ? a.role : x.role + ' / ' + a.role
  178. byKey[k] = x
  179. }
  180. }
  181. }
  182. let out = []
  183. for (k of order) { out.push(byKey[k]) }
  184. return out
  185. }
  186. // the photo: w185 → storage/mpackdb/profiles/<oldId>.<ext> — { requests, image, tmdbProfile, failed }
  187. static syncProfile = (p, profilePath) => {
  188. let out = { requests = 0 image = p.image tmdbProfile = p.tmdbProfile failed = null }
  189. if (textOr(profilePath) == null || p.oldId == null) { return out }
  190. let parts = profilePath.split('.')
  191. let ext = parts[parts.length - 1].toLowerCase()
  192. if (ext != 'jpg' && ext != 'jpeg' && ext != 'png' && ext != 'webp') { return out }
  193. let file = profilesDir + '/' + p.oldId + '.' + ext
  194. if (exists(file) && p.tmdbProfile == profilePath && p.image == ext) { return out }
  195. out.requests = 1
  196. let r = fetch(tmdbImageBase + '/w185' + profilePath, { headers = { 'user-agent' = 'tracker.worldapi.org (person page)' } timeoutMs = 20000 })
  197. if (r == null || r.status != 200 || r.body == null || r.body.length == 0) {
  198. out.failed = 'photo ' + profilePath + ': ' + (r == null ? 'no answer' : 'HTTP ' + r.status)
  199. return out
  200. }
  201. mkDir(profilesDir)
  202. writeFile(file, toBytes(r.body), 420)
  203. out.image = ext
  204. out.tmdbProfile = profilePath
  205. return out
  206. }
  207. // step 1: TMDB's person + credits; the facts (a TMDB value that is empty never overwrites ours) and the photo
  208. static startFill = (p) => {
  209. let t0 = now()
  210. let d = tmdbGet('/person/' + p.tmdbId + '?append_to_response=combined_credits')
  211. if (d.failed != null) { return { error = 'TMDB: ' + d.failed } }
  212. let items = creditItems(d.combined_credits)
  213. let photo = syncProfile(p, d.profile_path)
  214. let x = p
  215. if (textOr(d.biography) != null) { x.biography = d.biography }
  216. if (textOr(d.birthday) != null) { x.birthDate = d.birthday }
  217. if (textOr(d.deathday) != null) { x.deathDate = d.deathday }
  218. if (textOr(d.place_of_birth) != null) { x.birthPlace = d.place_of_birth }
  219. if (textOr(x.imdbId) == null && textOr(d.imdb_id) != null) { x.imdbId = d.imdb_id }
  220. if (d.adult != null) { x.adult = d.adult == true }
  221. x.image = photo.image
  222. x.tmdbProfile = photo.tmdbProfile
  223. if (personsTable.update(x.id, x) == null) { return { error = 'could not store the person: ' + personsTable.lastError() } }
  224. if (photo.failed != null) { console.log('person fill: ' + p.name + ': ' + photo.failed) }
  225. jobs[p.id] = { items = items pos = 0 credits = [] added = 0 requests = 1 + photo.requests started = t0 steps = 1 tmdbMs = now() - t0 }
  226. return { done = false pos = 0 total = items.length }
  227. }
  228. // a title we don't have: the minimal record (shape: search.hl importTitle) → its id, or null
  229. static addMinimalTitle = (cr, pid, pname) => {
  230. let kindType = cr.kind == 'tv' ? 'series' : 'movie'
  231. let rel = cr.release
  232. let yr = rel != null && rel.length >= 4 ? toNumber(rel.slice(0, 4)) : null
  233. let slug = slugOf(cr.title, cr.tmdbId)
  234. let record = {
  235. id = randomBytes(8, 'hex') title = cr.title summary = null tmdbSummary = cr.overview tmdbId = cr.tmdbId
  236. language = cr.language type = kindType release = rel year = yr image = null
  237. urlSegment = slug episodesCount = null homepage = null imdbId = null seasonsCount = null
  238. status = null tagline = null tvdbId = null tvmzId = null genres = genreIdsOfTmdb(cr.genreIds)
  239. cast = [{ person = pid character = cr.role actor = pname }] seasons = [] oldId = 'tmdb-' + cr.kind + '-' + cr.tmdbId
  240. adult = cr.adult imported = now() minimal = true
  241. }
  242. if (showsTable.put(record) == null) {
  243. console.log('person fill: "' + cr.title + '" not stored: ' + showsTable.lastError())
  244. return null
  245. }
  246. indexShow(record)
  247. refreshCatalogShow(record.id)
  248. return record.id
  249. }
  250. // every next step: the next `fillBatch` titles; the last one links the credits and marks the person complete
  251. static fillStep = (p, jobIn) => {
  252. let job = jobIn
  253. let total = job.items.length
  254. let end = job.pos + fillBatch
  255. if (end > total) { end = total }
  256. let credits = job.credits
  257. let added = job.added
  258. let pid = p.id
  259. let pname = p.name
  260. let i = job.pos
  261. while (i < end) {
  262. let item = job.items[i]
  263. let sid = titleIdByTmdb(item.kind == 'tv' ? 'series' : 'movie', item.tmdbId)
  264. if (sid == null) {
  265. sid = addMinimalTitle(item, pid, pname)
  266. if (sid != null) { added = added + 1 }
  267. }
  268. if (sid != null) { credits.push({ show = sid character = item.role posterPath = item.posterPath adult = item.adult }) }
  269. i = i + 1
  270. }
  271. job.pos = end
  272. job.credits = credits
  273. job.added = added
  274. job.steps = job.steps + 1
  275. jobs[pid] = job
  276. if (end < total) { return { done = false pos = end total = total } }
  277. // done: the credits (+ the migrated ones TMDB no longer lists), complete
  278. let have = {}
  279. let all = []
  280. for (c of credits) { have[c.show] = true all.push(c) }
  281. for (c of listOr(p.shows)) { if (c != null && c.show != null && have[c.show] != true) { all.push(c) } }
  282. let x = p
  283. x.shows = all
  284. x.filmographyAt = now()
  285. if (personsTable.update(x.id, x) == null) { return { error = 'could not store the filmography: ' + personsTable.lastError() } }
  286. personsTable.persist()
  287. showsTable.persist()
  288. jobs[pid] = null
  289. console.log('person fill: ' + pname + ' (tmdb ' + p.tmdbId + ') credits=' + total + ' added=' + added + ' requests=' + job.requests + ' steps=' + job.steps + ' tmdbMs=' + job.tmdbMs + ' ms=' + (now() - job.started))
  290. return { done = true pos = total total = total added = added }
  291. }
  292. // one step for the person with this id: { done, pos, total } or { error }
  293. static personFillStep = (personId) => {
  294. let p = personsTable.fetch(personId)
  295. if (p == null) { return { error = 'no such person' } }
  296. if (isComplete(p)) { return { done = true } }
  297. if (p.tmdbId == null) { return { error = 'TMDB does not know this person' } }
  298. let job = jobs[personId]
  299. if (job == null) { return startFill(p) }
  300. return fillStep(p, job)
  301. }
  302. // ---- people from a title's cast (tracker#26, details.hl completeStep) ------------------------------------------
  303. // TMDB person id → our person id, built once (15k people, one scan), kept up to date by castPersonId
  304. // (flat statics, written key by key like search.hl's haveTmdb: `let m = …ids` would copy the whole map — plain data
  305. // deep-copies on assignment)
  306. static personTmdbBuilt = { yes = false }
  307. static personTmdbIds = {}
  308. static personIdByTmdb = (tmdbId) => {
  309. if (!personTmdbBuilt.yes) {
  310. for (p of personsTable.find(null, null)) { if (p.tmdbId != null) { personTmdbIds['' + p.tmdbId] = p.id } }
  311. personTmdbBuilt.yes = true
  312. }
  313. return personTmdbIds['' + tmdbId]
  314. }
  315. // a free person slug: "Maya-Rudolph", taken → "-2", … (the migrated people's form)
  316. static personSlugOf = (name, tmdbId) => {
  317. ensureSlugIndex()
  318. let base = slugBaseOf(name, 'person-' + tmdbId)
  319. if (slugIndex.ids[base] == null) { return base }
  320. let n = 2
  321. while (slugIndex.ids[base + '-' + n] != null) { n = n + 1 }
  322. return base + '-' + n
  323. }
  324. // how many people castPersonId added since the start (details.hl counts a title's new people by the difference, tracker#28)
  325. static castPeopleAdded = { n = 0 }
  326. static castPeopleCount = () => { return castPeopleAdded.n }
  327. // our person for one cast entry of TMDB ({ tmdbId, name, character, gender }) — the one we have by TMDB id, else a NEW
  328. // person (name, tmdbId, a free slug, this one credit; not complete: their page fetches the rest on the first visit like any
  329. // migrated person). Answers the id, or null when it could not be stored. Also a CREW person (tracker#28: `character` = the
  330. // job, "Director").
  331. static castPersonId = (c, showId) => {
  332. let have = personIdByTmdb(c.tmdbId)
  333. if (have != null) { return have }
  334. let record = { id = randomBytes(8, 'hex') name = c.name tmdbId = c.tmdbId urlSegment = personSlugOf(c.name, c.tmdbId) oldId = 'tmdb-person-' + c.tmdbId
  335. gender = c.gender image = null tmdbProfile = null adult = false birthDate = null birthPlace = null deathDate = null imdbId = null
  336. shows = [{ show = showId character = c.character }] added = now() }
  337. if (personsTable.put(record) == null) {
  338. console.log('details: person "' + c.name + '" not stored: ' + personsTable.lastError())
  339. return null
  340. }
  341. personTmdbIds['' + c.tmdbId] = record.id
  342. slugIndex.ids[record.urlSegment] = record.id
  343. indexPerson(record)
  344. castPeopleAdded.n = castPeopleAdded.n + 1
  345. return record.id
  346. }

Branches

Latest commits

  • f83571c3tracker#18 (mission 067): daily sync by change lists — TMDB /tv|movie/changes (since the stored day, paged) + TVmaze /updates/shows → only our changed titles (followed: full step, unfollowed: light step — changed seasons, no TVmaze), full walk on first run / gap > 14 days / failed list; show record refreshed (title, tmdbSummary, tagline, status, genres …; renamed titles re-indexed); summary = the creator's own text (page: summary > tmdbSummary > tvmazeSummary), one-time clear of copied summaries (9,647 on the live copy); gate 261, tests/realdata-018*.mjs, README + STATUSmre
  • 10bb3f93tracker#28 (mission 061): full cast (all seasons, main cast by episodes, guest stars) + crew (created by, directed by, written by, screenplay, story, music) — stored by the details completion, the daily sync, the search import (one details request) and a background credits job (resumes, RSS limit); show page collapsed after 20 with client-side Show all; showBySlug via a slug map; gate 249, tests/realdata-028.mjs, README + STATUSmre
  • 25a50bc4tracker#26 (mission 059): titles from a filmography are completed — on open (skeleton, step-wise face showComplete, no reload) and by the in-app details repair (resumes, TMDB-paced, series in parts); cast from TMDB credits; gate 231, tests/realdata-026.mjs, README + STATUSmre
  • 1704ec45tracker#17 (mission 057): season caret down/up, skeleton rows while a season loads, sessionless showSeasonEpisodes face (no page re-mount), client-only close; gate 214, tests/realdata-057.mjs, README + STATUSmre
  • f2fe3e36mission 056: README + STATUS (merge, fixes, Hybriel 8590df63, real-data check), tests/realdata-056.mjs, tools/check-public-slugs.hlmre
  • f40c250emission 056: re-vendor hybriel master 8590df63 (#121, #122); an adult title's page is Not found for non-followers; gate: leave the page before stopping the servermre
  • 2b7fdd6cmission 056: signed-out header one row on phones ("Log in", nowrap), backfill skips adult titles' posters, gate checksmre
  • 2c53d5efMerge branch 't16-person' (tracker#16 person pages) into main; filmography shows only public titles (054 adult flag), gate race fix (backfill start line)mre
  • c171227emission 054: hide adult/unknown titles from the public lists and the search; in-app adult-flag backfill (TMDB details + poster per title, resumes), gate + real-data proofmre
  • 139fafd8tracker#16: short bio (4 lines, click = all), real-data check script, README + STATUSmre
  • 93be9476tracker#16: person pages /person/<slug> with the filmography fetched from TMDB on the first visit (step by step), gatemre
  • 47a3cae6STATUS: mission 053 merge commit idsmre
  • dcc5eecaMerge branch 't14-search'mre
  • 03edc783Merge branch 't15-tvmaze'mre
  • 71b46345tracker#15: numbering check by date or title, placeholder titles in other languages, docs + real-data proofmre
  • 6bb2daf1tracker#13: homepage (tiles, intro, latest movies/shows), /shows, /movies/page/N, /my/movies; lists cached in memorymre
  • b8bd1157tracker#14: README + STATUS (search, real-data numbers, gate, merge notes)mre
  • 65c694a8tracker#14: search — header magnifier, /search/<text> (in-memory word-prefix index over titles + people), Fetch from web (TMDB search/multi, ours left out), Add = import via syncShow; gate +25 checks, real-data scriptmre
  • 34f2c15btracker#15: TVmaze merge in the sync (gaps only: new episodes/seasons, empty titles/air dates; numbering check), fake TVmaze episodes + gatemre
  • cbdc4ea7tracker#12: link icons TMDB/IMDb/TVDB/TVmaze; sync fills missing ids (TVmaze lookup); movies fetched via /movie/mre