Files
researchowl/docs/shortsmith-phase2-spec.md
T
ChemaVXandClaude Opus 5 20c8d03aa7
Build & Deploy ResearchOwl / build-and-push (push) Successful in 9s
feat(short): generación y render de Shorts vía shortsmith
Añade /generate short_en y /short_spec. El pipeline genera un shot spec
con Haiku, verifica cada cifra, fecha y cita contra los chunks de la
sesión, lo renderiza en shortsmith y entrega el MP4 por Telegram junto
a un informe de claims.

- ShortsmithClient con sondeo y fallback al spec JSON si el render falla
- Contrato de plantillas obtenido de GET /templates, no codificado
- Comprobación de fundamento determinista, sin LLM
- outputs.published_url para enlazar el artículo de Ghost
- Normalización de comillas rectas a tipográficas (ver KNOWN-ISSUES.md)

Lo que no aparece en los chunks se contrasta contra el ejemplo del
prompt: si casa ahí es fuga, no invención, y se informa como tal. El
purgado de sesiones se lleva también su MP4.

La subida a YouTube queda fuera a propósito: fase 3.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 21:55:42 +00:00

12 KiB
Raw Blame History

Phase 2 — researchowl → shortsmith integration

Handoff document for Claude Code. Prerequisite: shortsmith v1 deployed and healthy (shortsmith-svc.shortsmith.svc.cluster.local:8080).

  • Repo touched: git.chemavx.xyz/chemavx/researchowl only. shortsmith is not modified.
  • Deliverable: /generate short_en produces a rendered MP4 from a research session and delivers it to Telegram for human review.
  • Explicitly out of scope: YouTube upload. That is phase 3. See §11.

1. What changes

/research <case>
        ↓
/generate blog en   →  Ghost article published, URL stored
        ↓
/generate short_en  →  Haiku writes a shot spec (JSON)
                    →  grounding check against source chunks
                    →  POST to shortsmith, poll, fetch MP4
                    →  Telegram: video + claims report
                    →  human reviews, uploads to YouTube manually

The Ghost step comes first and is a hard dependency: the Short's description links to the article, so the article URL must exist before the spec is generated. §6 covers what happens when it doesn't.


2. Why this is not like generating prose

Every other output type in generator.py produces text a human reads and judges. A shot spec is different in three ways, and each needs its own mitigation:

Property Consequence Mitigation
It's a typed contract, not prose Malformed output is unusable, not merely poor Validation retry loop, §4
It contains figures and quotes These are exactly what an LLM fabricates Grounding check, §5
It becomes a published video An error is public and hard to retract Human review gate, §8

The grounding check is the one that matters most. The channel's entire premise is that its numbers come from primary sources. A fabricated radar figure in a 40-second video is worse than no video.


3. Prompt construction — fetch the contract, don't hardcode it

shortsmith exposes GET /templates, which returns each template's prop schema. Fetch it at generation time and inject it into the prompt. Do not copy the schemas into researchowl.

This means adding a template to shortsmith makes it immediately available to the generator with no change here. Hardcoding the schemas would create a second source of truth that silently drifts — the same class of failure as the ffmpeg version delta that caused the v1 OOM.

Cache the response for the lifetime of the process; refetch on validation failure, in case the renderer was updated mid-run.

Narrative shapes

Free template choice produces mush. Give the model three shapes matching the three article types actually published, and have it pick one:

Shape Arc Fits
case_file hook → date/place → witness credentials → escalation → evidence → official explanation and its problem → close JAL 1628, Belgium, Ariel School
debunk the claim → why it spread → the method → the finding → what it means → close Roswell crater video, Yellow Sea star
document_drop what was released → the standout item → context → what's still missing → close PURSUE releases

examples/jal1628.json in the shortsmith repo is a worked case_file. Include it in the prompt as a full example — one concrete example is worth more than any amount of description.

Constraints to state explicitly

  • Total duration 2045 s. Not the 180 s contract ceiling; that is a hard limit, not a target.
  • Per-template max_length limits exist and are enforced. Listing them in the prompt turns a rejection into a non-event.
  • Colours are palette names (ink, amber, amber_dark, muted, dim, red), never hex.
  • Every figure and quote must come from the supplied chunks. No outside knowledge, even if correct.

4. Validation retry loop

for attempt in range(3):
    spec = await _generate_spec(prompt, feedback)
    try:
        validated = await client.validate(spec)   # POST /render dry-run or local pydantic
        break
    except ValidationError as e:
        feedback = _format_errors(e)              # feed the exact paths back
else:
    return _fallback(spec)                        # §6

shortsmith's discriminated union produces precise error locations (shots.0.radar_sweep.props.sweeeps, or shots.0 with the valid template names listed). Feed those paths back verbatim. They are more useful to the model than any paraphrase.

Cap at 3 attempts. Log attempts-to-valid as a metric — if it trends above 1.5, the prompt needs work, not the retry limit.


5. Grounding check — the important part

After the spec validates and before rendering, verify every factual string in it appears in the session's source material.

def check_grounding(spec, chunks) -> list[Ungrounded]:
    """Extract figures and quoted strings from spec props, confirm each
       appears in at least one source chunk."""

What to extract from the spec:

  • Every quoted string (quote_a, quote_b, quote, anything in “ ”)
  • Every number with a unit or magnitude (35,000 FT, 1,500, ~1,600 2,000 FT, 50 minutes)
  • Every date (17 NOV 1986, 5 MARCH 1987)
  • Proper nouns in label/key positions (ELMENDORF ROCC, CAPT. KENJU TERAUCHI)

Matching: normalise both sides — case-fold, strip thousands separators, collapse whitespace, normalise quote glyphs and dashes. Then substring match against chunk text. No LLM in this path: it must be deterministic and free.

On failure: do not silently drop the shot and do not retry blindly. Return the spec plus the list of ungrounded strings, and surface them in Telegram (§8). A human decides whether it is a real fabrication or a formatting artefact.

Expect false positives at first — "twice the size of an aircraft carrier" appears in the source but a rephrasing would not match. That is the correct bias: a false positive costs a glance, a false negative costs the channel's credibility.

This is the automated version of the fact-check table that was written by hand for the first Short. That table is in short-01-jal1628-script.md if you want the shape of the output.


6. ShortsmithClient

Mirror the GhostPublisher shape in generator.py. Layer rule holds: generator/ does not import from bot/; progress is reported through a generic callable.

class ShortsmithClient:
    def __init__(self, base_url: str, timeout: float = 600.0)
    async def templates(self) -> dict
    async def render(self, spec: dict) -> str            # -> job_id
    async def poll(self, job_id, on_progress=None) -> JobResult
    async def fetch_video(self, job_id) -> bytes

Polling: 2 s interval, 10 min ceiling. A 42 s Short renders in ~32 s; the 180 s ceiling takes ~138 s. Anything past 10 min is a stuck job, not a slow one.

Fallbacks always (repo convention). If shortsmith is unreachable, or the job errors, or grounding fails hard — return the spec JSON to Telegram as a file. The expensive part is the generation, not the render. Never discard it.

Config (src/config.py, Pydantic Settings, env-direct — no secret):

SHORTSMITH_URL = http://shortsmith-svc.shortsmith.svc.cluster.local:8080
SHORTSMITH_TIMEOUT = 600
SHORTSMITH_ENABLED = true

SHORTSMITH_ENABLED=false must make /generate short_en reply that the feature is off, not crash. This is the kill switch if the renderer misbehaves while nobody is watching.


7. Database

No migrations (CREATE TABLE IF NOT EXISTS convention holds).

  • outputs takes output_type='short_en', content = the spec JSON as text.
  • New: the Ghost article URL must be retrievable. Check whether GhostPublisher already persists it; if not, store it on the outputs row for the blog post, or add a published_url column to outputs (nullable, ALTER TABLE guarded by a column check). The spec generator needs it for the description.
  • Store the rendered MP4 on disk, not in SQLite. /data/shorts/{session_id}.mp4. Blobs in SQLite will make the WAL pathological.

8. Telegram flow

/generate short_en — reuse ProgressReporter, editing a single message:

🎬 Writing shot spec…            (Haiku, ~5 s)
🔍 Checking claims against sources…
🎞  Rendering… 40%               (progress from shortsmith poll)
📤 Uploading…

Then send the MP4 as a video message (not a document, so it plays inline), with a caption carrying the title and the article URL.

Immediately after, send the claims report as a separate message. This is the review gate and it must be impossible to miss:

✅ 11 claims matched to sources
⚠️ 2 not found:
   • "roughly 1,600 feet across"
   • "NORAD confirmed"

Sources: 14 chunks from 9 URLs
Cost: $0.004

Zero ungrounded claims still sends the report, saying so. A silent success trains the reader to stop looking.

Also add /short_spec to return the last spec JSON as a file, for hand-editing and re-rendering without regenerating.


9. Cost

One Haiku call over the top-scored chunks. ~$0.0030.008, plus retries. Rendering is free (own hardware). A Short costs roughly what a /generate blog costs, which for practical purposes is nothing — the constraint on volume is review time, not money.


10. Tests

Area Assert
Spec generation Mocked Haiku response validates; malformed response triggers retry with error paths fed back; 3 failures fall through to fallback
Grounding Known-good spec over known chunks yields zero ungrounded; a spec with an injected fabricated figure flags exactly that string; normalisation handles thousands separators, curly quotes, en-dashes
Client Poll loop handles queued→running→done, error status, timeout, connection refused
Fallback Every failure path returns the spec JSON rather than nothing
Layer separation grep that generator/ does not import from bot/

Golden eval, worth building once: run the generator against the stored JAL 1628 session and compare the output structurally to examples/jal1628.json — shape count, templates chosen, total duration, zero ungrounded claims. Not string equality; the model will phrase differently. It answers "could this pipeline have produced the video we already know is good?"


11. Out of scope — phase 3

YouTube upload via Data API v3. Deliberately excluded: it needs OAuth with a stored refresh token in researchowl-secrets (managed imperatively), a new failure surface, and it removes the human from the loop at exactly the point where the human is most valuable.

Ship phase 2, publish five or six Shorts by hand, then decide whether the review step is actually a bottleneck. It probably is not.

short_es for Zona de Exclusión is nearly free once this works — shortsmith draws whatever strings it is given and does not care about language. Only the prompt and the narrative shapes need translating. Do it after short_en has produced something worth publishing, not before.


12. Implementation order

One change at a time, verified before the next.

  1. ShortsmithClient + config + tests, against the live service. No generation yet — prove the plumbing by POSTing examples/jal1628.json and getting the MP4 back.
  2. Grounding checker + tests, standalone. Test it against the known-good JAL 1628 spec and against a deliberately corrupted copy.
  3. Spec generation: prompt, GET /templates injection, retry loop.
  4. Wire output_type='short_en' into generator.py; article URL retrieval.
  5. Telegram /generate short_en and /short_spec.
  6. Golden eval against the JAL 1628 session.

Step 2 before step 3 is deliberate. Build the check before the thing it checks, so the first generated spec is graded by a checker that was written without knowledge of it.