Files
researchowl/README.md
T
ChemaVXandClaude Opus 5 a43bbc752d
Build & Deploy ResearchOwl / build-and-push (push) Successful in 8s
avatar del bot: búho con gafas, generado y aplicado por API
Mismo patrón que hermes-bot y roswell-corpus: la foto de perfil se genera con
PIL (dibujado a 4× y reducido) en vez de ser un PNG opaco en el repo. Telegram
recorta en círculo y lo enseña a 48 px en la lista de chats, así que el dibujo
es una silueta y no una ilustración.

Las gafas son el único guiño a la investigación que sobrevive al tamaño chico
—una lupa o un libro se vuelven una mancha—, y la montura va en tono oscuro:
en ámbar sobre ámbar el aro se fundía con el plumaje y a 48 px desaparecía.

Documentado en el README, con el gotcha de `setMyProfilePhoto`: el parámetro
`photo` no admite el fichero suelto (responde `photo isn't specified`), hay que
pasarle un InputProfilePhoto que apunte al adjunto.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 19:57:18 +00:00

135 lines
4.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 🦉 ResearchOwl
**Exhaustive research engine with Telegram interface.**
Recursively discovers, scrapes, and processes sources from across the web,
then generates podcast scripts, blog posts, reports, or social threads using Ollama.
## Architecture
```
Telegram (/research <topic>)
ExhaustiveScraper
├── DuckDuckGo (8 queries × 5 results)
├── Wikipedia + recursive internal links
├── Reddit (top posts + top comments)
├── YouTube (transcripts)
├── PDFs (public documents)
└── Web scraping (trafilatura)
↓ recursive expansion (depth 1-3)
ContentProcessor (Ollama qwen2.5:7b + bge-m3 embeddings)
├── Chunking (800 token chunks, 100 overlap)
├── Quality scoring (0-10 per chunk)
├── Embeddings (cosine similarity RAG)
└── Deduplication
OutputGenerator (Ollama)
├── 🎙️ Podcast script (20-30 min)
├── 📝 Blog post (1500-2500 words)
├── 📊 Research report (structured)
└── 🐦 Social thread (15-25 tweets)
```
## Telegram Commands
| Command | Description |
|---------|-------------|
| `/research <topic>` | Start exhaustive research |
| `/status` | Check progress |
| `/finish` | Stop early, proceed to generation |
| `/generate podcast\|blog\|report\|thread` | Generate output |
| `/sources` | List all sources found |
| `/cancel` | Cancel current research |
## Local Development
```bash
# 1. Clone and setup
git clone https://git.chemavx.xyz/chemavx/researchowl
cd researchowl
# 2. Create virtualenv
python3 -m venv venv && source venv/bin/activate
pip install -r requirements.txt
# 3. Configure
cp .env.example .env
# Edit .env with your values
# 4. Run
python main.py
```
## Deploy to k3s
```bash
# 1. Create namespace and secrets
kubectl create namespace researchowl
kubectl create secret generic researchowl-secrets \
--from-literal=telegram-bot-token=YOUR_TOKEN \
--from-literal=telegram-allowed-users=YOUR_USER_ID \
-n researchowl
# 2. Copy manifests to your k8s-manifests repo
cp k8s/*.yaml /path/to/k8s-manifests/researchowl/
# 3. Apply ArgoCD app
kubectl apply -f k8s/argocd-app.yaml
# 4. Push to Gitea → Gitea Actions builds → ArgoCD deploys
git add . && git commit -m "feat: add researchowl" && git push
```
## Tuning
| Variable | Default | Description |
|----------|---------|-------------|
| `MAX_SOURCES` | 150 | Hard cap on sources |
| `MAX_DEPTH` | 3 | Link recursion depth |
| `QUALITY_THRESHOLD` | 0.4 | Min chunk quality (0-1) |
| `REQUEST_DELAY` | 1.0s | Delay between requests |
**Want more thoroughness?**
- Increase `MAX_SOURCES` to 300+
- Increase `MAX_DEPTH` to 4-5
- Lower `QUALITY_THRESHOLD` to 0.3
**Want faster results?**
- Lower `MAX_SOURCES` to 50
- Set `MAX_DEPTH` to 1-2
- Higher `QUALITY_THRESHOLD` to 0.6
## Bot avatar
The profile picture of `@chemavx_researchowl_bot` is not an opaque binary
checked into the repo: `assets/make_avatar.py` draws it with PIL at 4× and
scales it down, so the emblem can be retouched without hunting for an original.
Telegram crops avatars to a **circle**, so everything that matters lives inside
the inscribed circle; verified legible at 48 px.
```bash
python3 assets/make_avatar.py # writes assets/avatar.png
```
It is applied **over the API with the token from the secret, no BotFather**.
Watch out for `setMyProfilePhoto`: its `photo` parameter is not the file, it is
an `InputProfilePhoto` object pointing at the attachment. Posting the file on
its own gets you a baffling `photo isn't specified`.
```bash
TOK=$(kubectl get secret researchowl-secrets-infisical -n researchowl \
-o jsonpath='{.data.telegram-bot-token}' | base64 -d)
curl -s -F 'photo={"type":"static","photo":"attach://av"}' \
-F "av=@assets/avatar.png" \
"https://api.telegram.org/bot$TOK/setMyProfilePhoto"
unset TOK
```
## Notes
- Uses **qwen2.5:7b** (scoring) and **bge-m3** (embeddings) on your existing Ollama — zero API cost
- Optionally add `ANTHROPIC_API_KEY` for Claude fallback on generation
- SQLite database stored in `/data/researchowl.db`
- All outputs saved to DB and available via `/outputs`