Files
researchowl/docs/shortsmith-phase2-spec.md
T
ChemaVXandClaude Opus 5 20c8d03aa7
Build & Deploy ResearchOwl / build-and-push (push) Successful in 9s
feat(short): generación y render de Shorts vía shortsmith
Añade /generate short_en y /short_spec. El pipeline genera un shot spec
con Haiku, verifica cada cifra, fecha y cita contra los chunks de la
sesión, lo renderiza en shortsmith y entrega el MP4 por Telegram junto
a un informe de claims.

- ShortsmithClient con sondeo y fallback al spec JSON si el render falla
- Contrato de plantillas obtenido de GET /templates, no codificado
- Comprobación de fundamento determinista, sin LLM
- outputs.published_url para enlazar el artículo de Ghost
- Normalización de comillas rectas a tipográficas (ver KNOWN-ISSUES.md)

Lo que no aparece en los chunks se contrasta contra el ejemplo del
prompt: si casa ahí es fuga, no invención, y se informa como tal. El
purgado de sesiones se lleva también su MP4.

La subida a YouTube queda fuera a propósito: fase 3.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 21:55:42 +00:00

292 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Phase 2 — researchowl → shortsmith integration
**Handoff document for Claude Code.** Prerequisite: shortsmith v1 deployed and healthy
(`shortsmith-svc.shortsmith.svc.cluster.local:8080`).
- **Repo touched:** `git.chemavx.xyz/chemavx/researchowl` only. shortsmith is not modified.
- **Deliverable:** `/generate short_en` produces a rendered MP4 from a research session
and delivers it to Telegram for human review.
- **Explicitly out of scope:** YouTube upload. That is phase 3. See §11.
---
## 1. What changes
```
/research <case>
/generate blog en → Ghost article published, URL stored
/generate short_en → Haiku writes a shot spec (JSON)
→ grounding check against source chunks
→ POST to shortsmith, poll, fetch MP4
→ Telegram: video + claims report
→ human reviews, uploads to YouTube manually
```
The Ghost step comes first and is a hard dependency: the Short's description links to
the article, so the article URL must exist before the spec is generated. §6 covers what
happens when it doesn't.
---
## 2. Why this is not like generating prose
Every other output type in `generator.py` produces text a human reads and judges. A shot
spec is different in three ways, and each needs its own mitigation:
| Property | Consequence | Mitigation |
|---|---|---|
| It's a typed contract, not prose | Malformed output is unusable, not merely poor | Validation retry loop, §4 |
| It contains figures and quotes | These are exactly what an LLM fabricates | Grounding check, §5 |
| It becomes a published video | An error is public and hard to retract | Human review gate, §8 |
The grounding check is the one that matters most. The channel's entire premise is that
its numbers come from primary sources. A fabricated radar figure in a 40-second video is
worse than no video.
---
## 3. Prompt construction — fetch the contract, don't hardcode it
shortsmith exposes `GET /templates`, which returns each template's prop schema. **Fetch
it at generation time and inject it into the prompt.** Do not copy the schemas into
researchowl.
This means adding a template to shortsmith makes it immediately available to the
generator with no change here. Hardcoding the schemas would create a second source of
truth that silently drifts — the same class of failure as the ffmpeg version delta that
caused the v1 OOM.
Cache the response for the lifetime of the process; refetch on validation failure, in
case the renderer was updated mid-run.
### Narrative shapes
Free template choice produces mush. Give the model three shapes matching the three
article types actually published, and have it pick one:
| Shape | Arc | Fits |
|---|---|---|
| `case_file` | hook → date/place → witness credentials → escalation → evidence → official explanation and its problem → close | JAL 1628, Belgium, Ariel School |
| `debunk` | the claim → why it spread → the method → the finding → what it means → close | Roswell crater video, Yellow Sea star |
| `document_drop` | what was released → the standout item → context → what's still missing → close | PURSUE releases |
`examples/jal1628.json` in the shortsmith repo is a worked `case_file`. Include it in the
prompt as a full example — one concrete example is worth more than any amount of
description.
### Constraints to state explicitly
- Total duration **2045 s**. Not the 180 s contract ceiling; that is a hard limit, not a
target.
- Per-template `max_length` limits exist and are enforced. Listing them in the prompt
turns a rejection into a non-event.
- Colours are palette names (`ink`, `amber`, `amber_dark`, `muted`, `dim`, `red`), never
hex.
- Every figure and quote must come from the supplied chunks. No outside knowledge, even
if correct.
---
## 4. Validation retry loop
```python
for attempt in range(3):
spec = await _generate_spec(prompt, feedback)
try:
validated = await client.validate(spec) # POST /render dry-run or local pydantic
break
except ValidationError as e:
feedback = _format_errors(e) # feed the exact paths back
else:
return _fallback(spec) # §6
```
shortsmith's discriminated union produces precise error locations
(`shots.0.radar_sweep.props.sweeeps`, or `shots.0` with the valid template names listed).
**Feed those paths back verbatim.** They are more useful to the model than any
paraphrase.
Cap at 3 attempts. Log attempts-to-valid as a metric — if it trends above 1.5, the prompt
needs work, not the retry limit.
---
## 5. Grounding check — the important part
After the spec validates and **before** rendering, verify every factual string in it
appears in the session's source material.
```python
def check_grounding(spec, chunks) -> list[Ungrounded]:
"""Extract figures and quoted strings from spec props, confirm each
appears in at least one source chunk."""
```
**What to extract from the spec:**
- Every quoted string (`quote_a`, `quote_b`, `quote`, anything in `“ ”`)
- Every number with a unit or magnitude (`35,000 FT`, `1,500`, `~1,600 2,000 FT`, `50 minutes`)
- Every date (`17 NOV 1986`, `5 MARCH 1987`)
- Proper nouns in `label`/`key` positions (`ELMENDORF ROCC`, `CAPT. KENJU TERAUCHI`)
**Matching:** normalise both sides — case-fold, strip thousands separators, collapse
whitespace, normalise quote glyphs and dashes. Then substring match against chunk text.
No LLM in this path: it must be deterministic and free.
**On failure:** do not silently drop the shot and do not retry blindly. Return the spec
plus the list of ungrounded strings, and surface them in Telegram (§8). A human decides
whether it is a real fabrication or a formatting artefact.
Expect false positives at first — `"twice the size of an aircraft carrier"` appears in
the source but a rephrasing would not match. That is the correct bias: a false positive
costs a glance, a false negative costs the channel's credibility.
This is the automated version of the fact-check table that was written by hand for the
first Short. That table is in `short-01-jal1628-script.md` if you want the shape of the
output.
---
## 6. ShortsmithClient
Mirror the `GhostPublisher` shape in `generator.py`. Layer rule holds: `generator/` does
not import from `bot/`; progress is reported through a generic callable.
```python
class ShortsmithClient:
def __init__(self, base_url: str, timeout: float = 600.0)
async def templates(self) -> dict
async def render(self, spec: dict) -> str # -> job_id
async def poll(self, job_id, on_progress=None) -> JobResult
async def fetch_video(self, job_id) -> bytes
```
**Polling:** 2 s interval, 10 min ceiling. A 42 s Short renders in ~32 s; the 180 s
ceiling takes ~138 s. Anything past 10 min is a stuck job, not a slow one.
**Fallbacks always** (repo convention). If shortsmith is unreachable, or the job errors,
or grounding fails hard — **return the spec JSON to Telegram as a file**. The expensive
part is the generation, not the render. Never discard it.
**Config** (`src/config.py`, Pydantic Settings, env-direct — no secret):
```
SHORTSMITH_URL = http://shortsmith-svc.shortsmith.svc.cluster.local:8080
SHORTSMITH_TIMEOUT = 600
SHORTSMITH_ENABLED = true
```
`SHORTSMITH_ENABLED=false` must make `/generate short_en` reply that the feature is off,
not crash. This is the kill switch if the renderer misbehaves while nobody is watching.
---
## 7. Database
No migrations (`CREATE TABLE IF NOT EXISTS` convention holds).
- `outputs` takes `output_type='short_en'`, `content` = the spec JSON as text.
- **New:** the Ghost article URL must be retrievable. Check whether `GhostPublisher`
already persists it; if not, store it on the `outputs` row for the blog post, or add a
`published_url` column to `outputs` (nullable, `ALTER TABLE` guarded by a column check).
The spec generator needs it for the description.
- Store the rendered MP4 **on disk**, not in SQLite. `/data/shorts/{session_id}.mp4`.
Blobs in SQLite will make the WAL pathological.
---
## 8. Telegram flow
`/generate short_en` — reuse `ProgressReporter`, editing a single message:
```
🎬 Writing shot spec… (Haiku, ~5 s)
🔍 Checking claims against sources…
🎞 Rendering… 40% (progress from shortsmith poll)
📤 Uploading…
```
Then send the MP4 as a **video message** (not a document, so it plays inline), with a
caption carrying the title and the article URL.
**Immediately after, send the claims report as a separate message.** This is the review
gate and it must be impossible to miss:
```
✅ 11 claims matched to sources
⚠️ 2 not found:
• "roughly 1,600 feet across"
• "NORAD confirmed"
Sources: 14 chunks from 9 URLs
Cost: $0.004
```
Zero ungrounded claims still sends the report, saying so. A silent success trains the
reader to stop looking.
Also add `/short_spec` to return the last spec JSON as a file, for hand-editing and
re-rendering without regenerating.
---
## 9. Cost
One Haiku call over the top-scored chunks. ~$0.0030.008, plus retries. Rendering is free
(own hardware). A Short costs roughly what a `/generate blog` costs, which for practical
purposes is nothing — the constraint on volume is review time, not money.
---
## 10. Tests
| Area | Assert |
|---|---|
| Spec generation | Mocked Haiku response validates; malformed response triggers retry with error paths fed back; 3 failures fall through to fallback |
| Grounding | Known-good spec over known chunks yields zero ungrounded; a spec with an injected fabricated figure flags exactly that string; normalisation handles thousands separators, curly quotes, en-dashes |
| Client | Poll loop handles queued→running→done, error status, timeout, connection refused |
| Fallback | Every failure path returns the spec JSON rather than nothing |
| Layer separation | `grep` that `generator/` does not import from `bot/` |
**Golden eval, worth building once:** run the generator against the stored JAL 1628
session and compare the output structurally to `examples/jal1628.json` — shape count,
templates chosen, total duration, zero ungrounded claims. Not string equality; the model
will phrase differently. It answers "could this pipeline have produced the video we
already know is good?"
---
## 11. Out of scope — phase 3
YouTube upload via Data API v3. Deliberately excluded: it needs OAuth with a stored
refresh token in `researchowl-secrets` (managed imperatively), a new failure surface, and
it removes the human from the loop at exactly the point where the human is most valuable.
Ship phase 2, publish five or six Shorts by hand, then decide whether the review step is
actually a bottleneck. It probably is not.
`short_es` for Zona de Exclusión is nearly free once this works — shortsmith draws
whatever strings it is given and does not care about language. Only the prompt and the
narrative shapes need translating. Do it after `short_en` has produced something worth
publishing, not before.
---
## 12. Implementation order
One change at a time, verified before the next.
1. `ShortsmithClient` + config + tests, against the live service. No generation yet —
prove the plumbing by POSTing `examples/jal1628.json` and getting the MP4 back.
2. Grounding checker + tests, standalone. Test it against the known-good JAL 1628 spec
and against a deliberately corrupted copy.
3. Spec generation: prompt, `GET /templates` injection, retry loop.
4. Wire `output_type='short_en'` into `generator.py`; article URL retrieval.
5. Telegram `/generate short_en` and `/short_spec`.
6. Golden eval against the JAL 1628 session.
**Step 2 before step 3 is deliberate.** Build the check before the thing it checks, so
the first generated spec is graded by a checker that was written without knowledge of it.