# Phase 2 — researchowl → shortsmith integration **Handoff document for Claude Code.** Prerequisite: shortsmith v1 deployed and healthy (`shortsmith-svc.shortsmith.svc.cluster.local:8080`). - **Repo touched:** `git.chemavx.xyz/chemavx/researchowl` only. shortsmith is not modified. - **Deliverable:** `/generate short_en` produces a rendered MP4 from a research session and delivers it to Telegram for human review. - **Explicitly out of scope:** YouTube upload. That is phase 3. See §11. --- ## 1. What changes ``` /research ↓ /generate blog en → Ghost article published, URL stored ↓ /generate short_en → Haiku writes a shot spec (JSON) → grounding check against source chunks → POST to shortsmith, poll, fetch MP4 → Telegram: video + claims report → human reviews, uploads to YouTube manually ``` The Ghost step comes first and is a hard dependency: the Short's description links to the article, so the article URL must exist before the spec is generated. §6 covers what happens when it doesn't. --- ## 2. Why this is not like generating prose Every other output type in `generator.py` produces text a human reads and judges. A shot spec is different in three ways, and each needs its own mitigation: | Property | Consequence | Mitigation | |---|---|---| | It's a typed contract, not prose | Malformed output is unusable, not merely poor | Validation retry loop, §4 | | It contains figures and quotes | These are exactly what an LLM fabricates | Grounding check, §5 | | It becomes a published video | An error is public and hard to retract | Human review gate, §8 | The grounding check is the one that matters most. The channel's entire premise is that its numbers come from primary sources. A fabricated radar figure in a 40-second video is worse than no video. --- ## 3. Prompt construction — fetch the contract, don't hardcode it shortsmith exposes `GET /templates`, which returns each template's prop schema. **Fetch it at generation time and inject it into the prompt.** Do not copy the schemas into researchowl. This means adding a template to shortsmith makes it immediately available to the generator with no change here. Hardcoding the schemas would create a second source of truth that silently drifts — the same class of failure as the ffmpeg version delta that caused the v1 OOM. Cache the response for the lifetime of the process; refetch on validation failure, in case the renderer was updated mid-run. ### Narrative shapes Free template choice produces mush. Give the model three shapes matching the three article types actually published, and have it pick one: | Shape | Arc | Fits | |---|---|---| | `case_file` | hook → date/place → witness credentials → escalation → evidence → official explanation and its problem → close | JAL 1628, Belgium, Ariel School | | `debunk` | the claim → why it spread → the method → the finding → what it means → close | Roswell crater video, Yellow Sea star | | `document_drop` | what was released → the standout item → context → what's still missing → close | PURSUE releases | `examples/jal1628.json` in the shortsmith repo is a worked `case_file`. Include it in the prompt as a full example — one concrete example is worth more than any amount of description. ### Constraints to state explicitly - Total duration **20–45 s**. Not the 180 s contract ceiling; that is a hard limit, not a target. - Per-template `max_length` limits exist and are enforced. Listing them in the prompt turns a rejection into a non-event. - Colours are palette names (`ink`, `amber`, `amber_dark`, `muted`, `dim`, `red`), never hex. - Every figure and quote must come from the supplied chunks. No outside knowledge, even if correct. --- ## 4. Validation retry loop ```python for attempt in range(3): spec = await _generate_spec(prompt, feedback) try: validated = await client.validate(spec) # POST /render dry-run or local pydantic break except ValidationError as e: feedback = _format_errors(e) # feed the exact paths back else: return _fallback(spec) # §6 ``` shortsmith's discriminated union produces precise error locations (`shots.0.radar_sweep.props.sweeeps`, or `shots.0` with the valid template names listed). **Feed those paths back verbatim.** They are more useful to the model than any paraphrase. Cap at 3 attempts. Log attempts-to-valid as a metric — if it trends above 1.5, the prompt needs work, not the retry limit. --- ## 5. Grounding check — the important part After the spec validates and **before** rendering, verify every factual string in it appears in the session's source material. ```python def check_grounding(spec, chunks) -> list[Ungrounded]: """Extract figures and quoted strings from spec props, confirm each appears in at least one source chunk.""" ``` **What to extract from the spec:** - Every quoted string (`quote_a`, `quote_b`, `quote`, anything in `“ ”`) - Every number with a unit or magnitude (`35,000 FT`, `1,500`, `~1,600 – 2,000 FT`, `50 minutes`) - Every date (`17 NOV 1986`, `5 MARCH 1987`) - Proper nouns in `label`/`key` positions (`ELMENDORF ROCC`, `CAPT. KENJU TERAUCHI`) **Matching:** normalise both sides — case-fold, strip thousands separators, collapse whitespace, normalise quote glyphs and dashes. Then substring match against chunk text. No LLM in this path: it must be deterministic and free. **On failure:** do not silently drop the shot and do not retry blindly. Return the spec plus the list of ungrounded strings, and surface them in Telegram (§8). A human decides whether it is a real fabrication or a formatting artefact. Expect false positives at first — `"twice the size of an aircraft carrier"` appears in the source but a rephrasing would not match. That is the correct bias: a false positive costs a glance, a false negative costs the channel's credibility. This is the automated version of the fact-check table that was written by hand for the first Short. That table is in `short-01-jal1628-script.md` if you want the shape of the output. --- ## 6. ShortsmithClient Mirror the `GhostPublisher` shape in `generator.py`. Layer rule holds: `generator/` does not import from `bot/`; progress is reported through a generic callable. ```python class ShortsmithClient: def __init__(self, base_url: str, timeout: float = 600.0) async def templates(self) -> dict async def render(self, spec: dict) -> str # -> job_id async def poll(self, job_id, on_progress=None) -> JobResult async def fetch_video(self, job_id) -> bytes ``` **Polling:** 2 s interval, 10 min ceiling. A 42 s Short renders in ~32 s; the 180 s ceiling takes ~138 s. Anything past 10 min is a stuck job, not a slow one. **Fallbacks always** (repo convention). If shortsmith is unreachable, or the job errors, or grounding fails hard — **return the spec JSON to Telegram as a file**. The expensive part is the generation, not the render. Never discard it. **Config** (`src/config.py`, Pydantic Settings, env-direct — no secret): ``` SHORTSMITH_URL = http://shortsmith-svc.shortsmith.svc.cluster.local:8080 SHORTSMITH_TIMEOUT = 600 SHORTSMITH_ENABLED = true ``` `SHORTSMITH_ENABLED=false` must make `/generate short_en` reply that the feature is off, not crash. This is the kill switch if the renderer misbehaves while nobody is watching. --- ## 7. Database No migrations (`CREATE TABLE IF NOT EXISTS` convention holds). - `outputs` takes `output_type='short_en'`, `content` = the spec JSON as text. - **New:** the Ghost article URL must be retrievable. Check whether `GhostPublisher` already persists it; if not, store it on the `outputs` row for the blog post, or add a `published_url` column to `outputs` (nullable, `ALTER TABLE` guarded by a column check). The spec generator needs it for the description. - Store the rendered MP4 **on disk**, not in SQLite. `/data/shorts/{session_id}.mp4`. Blobs in SQLite will make the WAL pathological. --- ## 8. Telegram flow `/generate short_en` — reuse `ProgressReporter`, editing a single message: ``` 🎬 Writing shot spec… (Haiku, ~5 s) 🔍 Checking claims against sources… 🎞 Rendering… 40% (progress from shortsmith poll) 📤 Uploading… ``` Then send the MP4 as a **video message** (not a document, so it plays inline), with a caption carrying the title and the article URL. **Immediately after, send the claims report as a separate message.** This is the review gate and it must be impossible to miss: ``` ✅ 11 claims matched to sources ⚠️ 2 not found: • "roughly 1,600 feet across" • "NORAD confirmed" Sources: 14 chunks from 9 URLs Cost: $0.004 ``` Zero ungrounded claims still sends the report, saying so. A silent success trains the reader to stop looking. Also add `/short_spec` to return the last spec JSON as a file, for hand-editing and re-rendering without regenerating. --- ## 9. Cost One Haiku call over the top-scored chunks. ~$0.003–0.008, plus retries. Rendering is free (own hardware). A Short costs roughly what a `/generate blog` costs, which for practical purposes is nothing — the constraint on volume is review time, not money. --- ## 10. Tests | Area | Assert | |---|---| | Spec generation | Mocked Haiku response validates; malformed response triggers retry with error paths fed back; 3 failures fall through to fallback | | Grounding | Known-good spec over known chunks yields zero ungrounded; a spec with an injected fabricated figure flags exactly that string; normalisation handles thousands separators, curly quotes, en-dashes | | Client | Poll loop handles queued→running→done, error status, timeout, connection refused | | Fallback | Every failure path returns the spec JSON rather than nothing | | Layer separation | `grep` that `generator/` does not import from `bot/` | **Golden eval, worth building once:** run the generator against the stored JAL 1628 session and compare the output structurally to `examples/jal1628.json` — shape count, templates chosen, total duration, zero ungrounded claims. Not string equality; the model will phrase differently. It answers "could this pipeline have produced the video we already know is good?" --- ## 11. Out of scope — phase 3 YouTube upload via Data API v3. Deliberately excluded: it needs OAuth with a stored refresh token in `researchowl-secrets` (managed imperatively), a new failure surface, and it removes the human from the loop at exactly the point where the human is most valuable. Ship phase 2, publish five or six Shorts by hand, then decide whether the review step is actually a bottleneck. It probably is not. `short_es` for Zona de Exclusión is nearly free once this works — shortsmith draws whatever strings it is given and does not care about language. Only the prompt and the narrative shapes need translating. Do it after `short_en` has produced something worth publishing, not before. --- ## 12. Implementation order One change at a time, verified before the next. 1. `ShortsmithClient` + config + tests, against the live service. No generation yet — prove the plumbing by POSTing `examples/jal1628.json` and getting the MP4 back. 2. Grounding checker + tests, standalone. Test it against the known-good JAL 1628 spec and against a deliberately corrupted copy. 3. Spec generation: prompt, `GET /templates` injection, retry loop. 4. Wire `output_type='short_en'` into `generator.py`; article URL retrieval. 5. Telegram `/generate short_en` and `/short_spec`. 6. Golden eval against the JAL 1628 session. **Step 2 before step 3 is deliberate.** Build the check before the thing it checks, so the first generated spec is graded by a checker that was written without knowledge of it.