# 🦉 ResearchOwl **Exhaustive research engine with Telegram interface.** Recursively discovers, scrapes, and processes sources from across the web, then generates podcast scripts, blog posts, reports, or social threads using Ollama. ## Architecture ``` Telegram (/research ) ↓ ExhaustiveScraper ├── DuckDuckGo (8 queries × 5 results) ├── Wikipedia + recursive internal links ├── Reddit (top posts + top comments) ├── YouTube (transcripts) ├── PDFs (public documents) └── Web scraping (trafilatura) ↓ recursive expansion (depth 1-3) ContentProcessor (Ollama qwen2.5:7b + bge-m3 embeddings) ├── Chunking (800 token chunks, 100 overlap) ├── Quality scoring (0-10 per chunk) ├── Embeddings (cosine similarity RAG) └── Deduplication ↓ OutputGenerator (Ollama) ├── 🎙️ Podcast script (20-30 min) ├── 📝 Blog post (1500-2500 words) ├── 📊 Research report (structured) └── 🐦 Social thread (15-25 tweets) ``` ## Telegram Commands | Command | Description | |---------|-------------| | `/research ` | Start exhaustive research | | `/status` | Check progress | | `/finish` | Stop early, proceed to generation | | `/generate podcast\|blog\|report\|thread` | Generate output | | `/generate short_en` | Vertical Short: shot spec → grounding check → MP4 | | `/short_spec` | Last shot spec as a JSON file; edit it and send it back to re-render free | | `/upload_short` | Upload the rendered Short to YouTube (private, for review) | | `/sources` | List all sources found | | `/cancel` | Cancel current research | ## Shorts (`/generate short_en`) Claude writes a **shot spec** — typed JSON, not prose — which [shortsmith](https://git.chemavx.xyz/chemavx/shortsmith) renders into a 1080×1920 MP4. The bot sends the video and, in a separate message, a **claims report**. ``` /research JAL 1628 Alaska 1986 … /generate blog en → Ghost draft, article URL stored on the output row /generate short_en → spec → grounding → render → video + claims report /upload_short → uploads to YouTube as PRIVATE, with metadata filled in; publishing stays a human click in Studio ``` Three things make this different from generating text, and each has its own mitigation: - **It is a contract, not prose.** The template schemas are fetched live from `GET /templates` and never copied here, so a template added to shortsmith is available immediately. A spec is validated locally against those schemas before anything renders, and the exact error paths (`shots.0.radar_sweep.props.sweeeps`) go back to the model verbatim — up to 3 attempts. - **It contains figures and quotes.** `grounding.py` extracts every quote, figure, date and proper noun and checks it against the exact chunks the model was given. No LLM in that path: normalisation plus substring, deterministic and free. Whatever is not in the chunks is checked against the worked example that travels in the prompt, so a figure lifted from it is reported as a **prompt leak**, not as an invention — different diagnosis, different fix. Neither ever blocks the render: both are surfaced next to the video and a human decides. - **It becomes a published video.** `/generate short_en` uploads nothing: the MP4 lands in Telegram for review and in `/data/shorts/{session_id}.mp4`. Getting it onto the channel is a separate, explicit `/upload_short`. Fallbacks hold throughout: if shortsmith is unreachable, the job errors, or the spec never validates, the spec JSON comes back as a file. The expensive part is the generation, not the render. **Hand-editing loop:** `/short_spec` hands you the spec as `short_{session_id}_spec.json`; edit it and send the file back to the bot. It validates against the live contract (errors come back with their exact paths), **re-runs the grounding check** — your edit may have introduced a new figure — saves the edited spec as a new output, and renders. No LLM in that path: it is free. The session comes from the filename, so it works even if the chat has researched something else since. Full spec of the phase: `docs/shortsmith-phase2-spec.md`. ## YouTube (`/upload_short`) Uploads `/data/shorts/{session_id}.mp4` to the channel with the title from the spec, a description carrying the article link and the sources the Short cites on screen, and tags derived from the topic. The YouTube URL is written back to the output row, so a second `/upload_short` on the same session refuses unless you say `/upload_short force`. It also refuses if the MP4 on disk is **older than the latest saved spec** — that happens when a spec regeneration's render fails, and uploading would put the new metadata on the old video. **Read this before setting it up.** Videos uploaded through `videos.insert` from an **unaudited API project** are [restricted to private viewing mode](https://developers.google.com/youtube/v3/docs/videos/insert). The lock belongs to the API project, not to the video — you do not unlock it from Studio, you unlock it by passing Google's compliance audit. So this command does not publish. It puts the video on the channel with the metadata already filled in and hands back the Studio link; a person reviews and presses publish. That is the same shape as `/publish`, which only ever writes Ghost drafts. One-time setup, in [console.cloud.google.com](https://console.cloud.google.com): 1. Enable **YouTube Data API v3** on a project. 2. OAuth consent screen → External → **publish it to "In production"**. Leaving it in "Testing" makes Google revoke the refresh token after seven days, and the bot dies on its own the following Tuesday. 3. Credentials → OAuth client ID → **Desktop app**. 4. `python scripts/youtube_oauth.py --client-id … --client-secret …`, which opens a browser, catches the redirect on localhost and prints the refresh token. 5. Put `youtube-client-id`, `youtube-client-secret` and `youtube-refresh-token` into Infisical (they arrive as `researchowl-secrets-infisical`). The scope requested is `youtube.upload` only: a leaked token cannot read or delete anything on the channel — the worst it can do is upload. Quota is not a concern (1 unit per upload, 100 uploads a day). `YOUTUBE_ENABLED=false` is the kill switch. ## Local Development ```bash # 1. Clone and setup git clone https://git.chemavx.xyz/chemavx/researchowl cd researchowl # 2. Create virtualenv python3 -m venv venv && source venv/bin/activate pip install -r requirements.txt # 3. Configure cp .env.example .env # Edit .env with your values # 4. Run python main.py ``` ## Deploy to k3s ```bash # 1. Create namespace and secrets kubectl create namespace researchowl kubectl create secret generic researchowl-secrets \ --from-literal=telegram-bot-token=YOUR_TOKEN \ --from-literal=telegram-allowed-users=YOUR_USER_ID \ -n researchowl # 2. Copy manifests to your k8s-manifests repo cp k8s/*.yaml /path/to/k8s-manifests/researchowl/ # 3. Apply ArgoCD app kubectl apply -f k8s/argocd-app.yaml # 4. Push to Gitea → Gitea Actions builds → ArgoCD deploys git add . && git commit -m "feat: add researchowl" && git push ``` ## Tuning | Variable | Default | Description | |----------|---------|-------------| | `MAX_SOURCES` | 150 | Hard cap on sources | | `MAX_DEPTH` | 3 | Link recursion depth | | `QUALITY_THRESHOLD` | 0.4 | Min chunk quality (0-1) | | `REQUEST_DELAY` | 1.0s | Delay between requests | **Want more thoroughness?** - Increase `MAX_SOURCES` to 300+ - Increase `MAX_DEPTH` to 4-5 - Lower `QUALITY_THRESHOLD` to 0.3 **Want faster results?** - Lower `MAX_SOURCES` to 50 - Set `MAX_DEPTH` to 1-2 - Higher `QUALITY_THRESHOLD` to 0.6 ## Bot avatar The profile picture of `@chemavx_researchowl_bot` is not an opaque binary checked into the repo: `assets/make_avatar.py` draws it with PIL at 4× and scales it down, so the emblem can be retouched without hunting for an original. Telegram crops avatars to a **circle**, so everything that matters lives inside the inscribed circle; verified legible at 48 px. ```bash python3 assets/make_avatar.py # writes assets/avatar.png ``` It is applied **over the API with the token from the secret, no BotFather**. Watch out for `setMyProfilePhoto`: its `photo` parameter is not the file, it is an `InputProfilePhoto` object pointing at the attachment. Posting the file on its own gets you a baffling `photo isn't specified`. ```bash TOK=$(kubectl get secret researchowl-secrets-infisical -n researchowl \ -o jsonpath='{.data.telegram-bot-token}' | base64 -d) curl -s -F 'photo={"type":"static","photo":"attach://av"}' \ -F "av=@assets/avatar.png" \ "https://api.telegram.org/bot$TOK/setMyProfilePhoto" unset TOK ``` ## Notes - Uses **qwen2.5:7b** (scoring) and **bge-m3** (embeddings) on your existing Ollama — zero API cost - Optionally add `ANTHROPIC_API_KEY` for Claude fallback on generation - SQLite database stored in `/data/researchowl.db` - All outputs saved to DB and available via `/outputs`