- Priority: correctness > simplicity > speed
- Before any code change:
git diff --stat+git log --oneline -5 - On any user correction: codify a rule before resuming work
OpenSeed is an AI-powered research workflow CLI. It manages a local paper library (ArXiv fetch, PDF extraction) and provides Claude-powered summarization, review, and Q&A.
- Max function body: 15 lines. Extract or redesign if exceeded.
- No comments that restate code. Only "why" comments for non-obvious decisions.
- Prefer composition over inheritance. Prefer data transforms over mutation.
- Every abstraction must justify itself: used <2 places → inline it.
- No TODOs in committed code. Delete dead code paths immediately.
- Type signatures are documentation. Verbose names > comments.
- When two approaches are equally correct, pick the one with fewer moving parts.
- Lint:
ruff check src/ tests/andruff format src/ tests/must pass before commit. - Tests:
pytest tests/ -vmust pass before commit. - Line length: 100 chars max.
- Python 3.11+ — use
X | Yunion syntax, notUnion[X, Y]. - Delete dead code outright. No
# deprecatedor commented-out blocks. - Modify only files relevant to the task.
- Runtime: Python 3.11+
- AI:
anthropicSDK — client viaauth.make_anthropic_client()(supports bothANTHROPIC_API_KEYandCLAUDE_CODE_SETUP_TOKEN) - CLI: Click + Rich
- Models: Pydantic v2
- Layout: src-layout (
src/openseed/)
| Module | Purpose |
|---|---|
cli/ |
Click groups: paper, experiment, agent, alerts |
models/ |
Pydantic models: Paper, Author, Tag, Experiment, ExperimentRun, Claim, ClaimEdge, Alert |
storage/library.py |
SQLite-backed CRUD for papers + experiments + knowledge graph + claims |
services/arxiv.py |
ArXiv metadata fetch + search (sync + async) |
services/pdf.py |
PDF text extraction via PyMuPDF |
services/scholar.py |
Semantic Scholar API client (citations, references, recommendations) |
services/watch.py |
Watch execution service (run watches, return results) |
services/cron.py |
Crontab management (install/remove/status) |
services/digest.py |
Digest generation (markdown summary of watch results) |
storage/migrate.py |
JSON → SQLite auto-migration |
agent/reader.py |
PaperReader — structured summarize/analyze via Claude |
agent/discovery.py |
Paper discovery — Claude search + S2 enrichment |
agent/compare.py |
Paper comparison — structured side-by-side analysis |
agent/latex.py |
LaTeX related-work export with BibTeX |
agent/claims.py |
Claim extraction — atomic claims from papers via Claude |
agent/matcher.py |
Claim matching — FTS5 retrieval + Claude classification + alerts |
agent/assistant.py |
ResearchAssistant — freeform ask/review via Claude |
services/rss.py |
RSS/Atom feed discovery |
services/sharing.py |
Research session export/import for collaboration |
web/app.py |
FastAPI web dashboard |
auth.py |
make_anthropic_client(), has_anthropic_auth(), run_claude_setup_token() |
doctor.py |
Environment health checks with CheckResult + fix hints |
config.py |
OpenSeedConfig, paths, default model |
Only read when the trigger matches — do not bulk-load.
| Trigger | Read |
|---|---|
| ArXiv fetch / search issues | src/openseed/services/arxiv.py |
| PDF extraction issues | src/openseed/services/pdf.py |
| Agent AI features | src/openseed/agent/reader.py, assistant.py |
| Auth / API key issues | src/openseed/auth.py |
| Storage / data bugs | src/openseed/storage/library.py |
| CLI command issues | src/openseed/cli/<command>.py |
| Config / paths | src/openseed/config.py |
make install # pip install -e ".[dev]"
make test # pytest -v
make lint # ruff check src/ tests/
make format # ruff format src/ tests/
openseed setup # Configure auth + model
openseed doctor # Environment health check
openseed paper add # Add paper by ArXiv URL
openseed paper list # List library
openseed agent ask # Ask research question