Files
researchowl/README.md
T
ChemaVXandClaude Opus 5 ec3e6d05c0 feat(youtube): subir Shorts al canal con /upload_short
Fase 3, con una corrección sobre lo que decía la §11 de la spec de fase 2.

El bloqueo no es OAuth. Los vídeos subidos por videos.insert desde un
proyecto de API sin auditar quedan restringidos a privado, y el candado es
del proyecto, no del vídeo: no se abre desde Studio, se abre pasando la
auditoría de cumplimiento de Google. Así que esto no publica. Deja el vídeo
en el canal con los metadatos puestos y devuelve el enlace de Studio para
que una persona lo revise y le dé a publicar — la misma forma que /publish
con los borradores de Ghost, y por la misma razón: el informe de fundamento
no sirve de nada si el vídeo ya está subido cuando lo lees.

Comando aparte, no un paso de /generate short_en.

- src/generator/youtube.py: refresco de token contra oauth2.googleapis.com,
  subida resumable en dos pasos y metadatos derivados del shot spec ya
  guardado (título, enlace al artículo, fuentes que el Short cita en
  pantalla, etiquetas del tema). Sin google-api-python-client: es síncrono
  y bloquearía el loop del bot; son dos peticiones HTTP y el repo ya firma
  los JWT de Ghost a mano. aiohttp con SAFE_ACCEPT_ENCODING como todo lo
  demás.
- Scope youtube.upload y nada más: un token filtrado no puede leer ni
  borrar nada del canal, sólo subir.
- forced_private detecta que YouTube devolvió "private" cuando se pidió
  otra cosa, y el aviso lo dice. Es la firma del candado, y tragárselo
  haría creer que salió publicado.
- invalid_grant se traduce a su causa real: la pantalla de consentimiento
  quedó en "Testing" y Google revoca esos tokens a los siete días. Es el
  fallo que menos se adivina y el que más probable es encontrarse.
- get_article_url ahora excluye las filas short_en. Su published_url pasa a
  ser la URL de YouTube, y sin el filtro el siguiente Short de la sesión
  enlazaría al Short anterior: un bucle silencioso, porque la URL es válida
  y nadie la mira dos veces.
- scripts/youtube_oauth.py, sólo stdlib: corre en el portátil, no en el
  contenedor, y no debería exigir instalar nada.
- Las tres claves van optional:true en el Deployment. Sin eso, una clave que
  aún no está en Infisical deja el pod en CreateContainerConfigError y tira
  el bot entero por una función que nadie ha pedido todavía.

30 tests nuevos contra un servidor falso. No hay test en vivo a propósito:
cualquier ejecución real sube un vídeo a un canal de verdad, y eso no es
algo que deba pasar por teclear pytest.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 14:43:53 +00:00

214 lines
8.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 🦉 ResearchOwl
**Exhaustive research engine with Telegram interface.**
Recursively discovers, scrapes, and processes sources from across the web,
then generates podcast scripts, blog posts, reports, or social threads using Ollama.
## Architecture
```
Telegram (/research <topic>)
ExhaustiveScraper
├── DuckDuckGo (8 queries × 5 results)
├── Wikipedia + recursive internal links
├── Reddit (top posts + top comments)
├── YouTube (transcripts)
├── PDFs (public documents)
└── Web scraping (trafilatura)
↓ recursive expansion (depth 1-3)
ContentProcessor (Ollama qwen2.5:7b + bge-m3 embeddings)
├── Chunking (800 token chunks, 100 overlap)
├── Quality scoring (0-10 per chunk)
├── Embeddings (cosine similarity RAG)
└── Deduplication
OutputGenerator (Ollama)
├── 🎙️ Podcast script (20-30 min)
├── 📝 Blog post (1500-2500 words)
├── 📊 Research report (structured)
└── 🐦 Social thread (15-25 tweets)
```
## Telegram Commands
| Command | Description |
|---------|-------------|
| `/research <topic>` | Start exhaustive research |
| `/status` | Check progress |
| `/finish` | Stop early, proceed to generation |
| `/generate podcast\|blog\|report\|thread` | Generate output |
| `/generate short_en` | Vertical Short: shot spec → grounding check → MP4 |
| `/short_spec` | Last shot spec as a JSON file, to hand-edit and re-render |
| `/upload_short` | Upload the rendered Short to YouTube (private, for review) |
| `/sources` | List all sources found |
| `/cancel` | Cancel current research |
## Shorts (`/generate short_en`)
Claude writes a **shot spec** — typed JSON, not prose — which
[shortsmith](https://git.chemavx.xyz/chemavx/shortsmith) renders into a 1080×1920
MP4. The bot sends the video and, in a separate message, a **claims report**.
```
/research JAL 1628 Alaska 1986 …
/generate blog en → Ghost draft, article URL stored on the output row
/generate short_en → spec → grounding → render → video + claims report
/upload_short → uploads to YouTube as PRIVATE, with metadata filled
in; publishing stays a human click in Studio
```
Three things make this different from generating text, and each has its own
mitigation:
- **It is a contract, not prose.** The template schemas are fetched live from
`GET /templates` and never copied here, so a template added to shortsmith is
available immediately. A spec is validated locally against those schemas
before anything renders, and the exact error paths
(`shots.0.radar_sweep.props.sweeeps`) go back to the model verbatim — up to 3
attempts.
- **It contains figures and quotes.** `grounding.py` extracts every quote,
figure, date and proper noun and checks it against the exact chunks the model
was given. No LLM in that path: normalisation plus substring, deterministic
and free. Whatever is not in the chunks is checked against the worked example
that travels in the prompt, so a figure lifted from it is reported as a
**prompt leak**, not as an invention — different diagnosis, different fix.
Neither ever blocks the render: both are surfaced next to the video and a
human decides.
- **It becomes a published video.** `/generate short_en` uploads nothing: the
MP4 lands in Telegram for review and in `/data/shorts/{session_id}.mp4`.
Getting it onto the channel is a separate, explicit `/upload_short`.
Fallbacks hold throughout: if shortsmith is unreachable, the job errors, or the
spec never validates, the spec JSON comes back as a file. The expensive part is
the generation, not the render.
Full spec of the phase: `docs/shortsmith-phase2-spec.md`.
## YouTube (`/upload_short`)
Uploads `/data/shorts/{session_id}.mp4` to the channel with the title from the
spec, a description carrying the article link and the sources the Short cites on
screen, and tags derived from the topic. The YouTube URL is written back to the
output row, so a second `/upload_short` on the same session refuses unless you
say `/upload_short force`.
**Read this before setting it up.** Videos uploaded through `videos.insert` from
an **unaudited API project** are [restricted to private viewing
mode](https://developers.google.com/youtube/v3/docs/videos/insert). The lock
belongs to the API project, not to the video — you do not unlock it from Studio,
you unlock it by passing Google's compliance audit. So this command does not
publish. It puts the video on the channel with the metadata already filled in
and hands back the Studio link; a person reviews and presses publish. That is
the same shape as `/publish`, which only ever writes Ghost drafts.
One-time setup, in [console.cloud.google.com](https://console.cloud.google.com):
1. Enable **YouTube Data API v3** on a project.
2. OAuth consent screen → External → **publish it to "In production"**. Leaving
it in "Testing" makes Google revoke the refresh token after seven days, and
the bot dies on its own the following Tuesday.
3. Credentials → OAuth client ID → **Desktop app**.
4. `python scripts/youtube_oauth.py --client-id … --client-secret …`, which
opens a browser, catches the redirect on localhost and prints the refresh
token.
5. Put `youtube-client-id`, `youtube-client-secret` and `youtube-refresh-token`
into Infisical (they arrive as `researchowl-secrets-infisical`).
The scope requested is `youtube.upload` only: a leaked token cannot read or
delete anything on the channel — the worst it can do is upload. Quota is not a
concern (1 unit per upload, 100 uploads a day). `YOUTUBE_ENABLED=false` is the
kill switch.
## Local Development
```bash
# 1. Clone and setup
git clone https://git.chemavx.xyz/chemavx/researchowl
cd researchowl
# 2. Create virtualenv
python3 -m venv venv && source venv/bin/activate
pip install -r requirements.txt
# 3. Configure
cp .env.example .env
# Edit .env with your values
# 4. Run
python main.py
```
## Deploy to k3s
```bash
# 1. Create namespace and secrets
kubectl create namespace researchowl
kubectl create secret generic researchowl-secrets \
--from-literal=telegram-bot-token=YOUR_TOKEN \
--from-literal=telegram-allowed-users=YOUR_USER_ID \
-n researchowl
# 2. Copy manifests to your k8s-manifests repo
cp k8s/*.yaml /path/to/k8s-manifests/researchowl/
# 3. Apply ArgoCD app
kubectl apply -f k8s/argocd-app.yaml
# 4. Push to Gitea → Gitea Actions builds → ArgoCD deploys
git add . && git commit -m "feat: add researchowl" && git push
```
## Tuning
| Variable | Default | Description |
|----------|---------|-------------|
| `MAX_SOURCES` | 150 | Hard cap on sources |
| `MAX_DEPTH` | 3 | Link recursion depth |
| `QUALITY_THRESHOLD` | 0.4 | Min chunk quality (0-1) |
| `REQUEST_DELAY` | 1.0s | Delay between requests |
**Want more thoroughness?**
- Increase `MAX_SOURCES` to 300+
- Increase `MAX_DEPTH` to 4-5
- Lower `QUALITY_THRESHOLD` to 0.3
**Want faster results?**
- Lower `MAX_SOURCES` to 50
- Set `MAX_DEPTH` to 1-2
- Higher `QUALITY_THRESHOLD` to 0.6
## Bot avatar
The profile picture of `@chemavx_researchowl_bot` is not an opaque binary
checked into the repo: `assets/make_avatar.py` draws it with PIL at 4× and
scales it down, so the emblem can be retouched without hunting for an original.
Telegram crops avatars to a **circle**, so everything that matters lives inside
the inscribed circle; verified legible at 48 px.
```bash
python3 assets/make_avatar.py # writes assets/avatar.png
```
It is applied **over the API with the token from the secret, no BotFather**.
Watch out for `setMyProfilePhoto`: its `photo` parameter is not the file, it is
an `InputProfilePhoto` object pointing at the attachment. Posting the file on
its own gets you a baffling `photo isn't specified`.
```bash
TOK=$(kubectl get secret researchowl-secrets-infisical -n researchowl \
-o jsonpath='{.data.telegram-bot-token}' | base64 -d)
curl -s -F 'photo={"type":"static","photo":"attach://av"}' \
-F "av=@assets/avatar.png" \
"https://api.telegram.org/bot$TOK/setMyProfilePhoto"
unset TOK
```
## Notes
- Uses **qwen2.5:7b** (scoring) and **bge-m3** (embeddings) on your existing Ollama — zero API cost
- Optionally add `ANTHROPIC_API_KEY` for Claude fallback on generation
- SQLite database stored in `/data/researchowl.db`
- All outputs saved to DB and available via `/outputs`