A research assistant that treats the citation network as a first-class signal. Most related-work tools rank papers by embedding similarity alone; this one fuses semantic similarity with co-citation and bibliographic coupling drawn from a citation graph, so papers that are structurally central to a literature surface even when their abstracts read differently.
Seven modules behind one ResearchAgent class, driven from a single CLI.
| # | Module | Responsibility |
|---|---|---|
| 1 | PaperSearchModule |
Search papers via Semantic Scholar and ArXiv |
| 2 | CitationExtractorModule |
Parse citations from a PDF or raw reference text |
| 3 | CitationGraphModule |
Build and query the citation knowledge graph |
| 4 | RelatedWorkModule |
Multi-signal recommendation (four signals fused) |
| 5 | GapAnalyzerModule |
Detect missing citations in a reference list |
| 6 | PlagiarismDetectorModule |
Citation-aware plagiarism detection |
| 7 | ReferenceFormatterModule |
APA / IEEE / MLA / Chicago formatting |
Related-work scoring fuses four independent signals with fixed weights (config.SIGNAL_WEIGHTS):
| Signal | Weight | Basis |
|---|---|---|
| Semantic | 0.40 | Sentence-transformer embedding cosine similarity (all-MiniLM-L6-v2) |
| Co-citation | 0.25 | How often two papers are cited together |
| Bibliographic coupling | 0.20 | How much two papers' reference lists overlap |
| TF-IDF | 0.15 | Surface keyword overlap |
Graph centrality uses PageRank (α = 0.85). Pairs scoring below 0.15 similarity are discarded.
Combines TF-IDF similarity (0.60) with Winnowing fingerprinting (0.40) over 5-character shingles in a 4-token window, then buckets the result:
| Risk level | Combined score |
|---|---|
| Clean | 0.00 – 0.20 |
| Borderline | 0.20 – 0.40 |
| Suspicious | 0.40 – 0.60 |
| High risk | 0.60 – 1.00 |
It is citation-aware: correctly attributed quotations are not counted as overlap.
cd research_agent
pip install -r requirements.txtpython main.py search --query "transformer attention" --max 5
python main.py extract --text "References\n[1] ..."
python main.py related --abstract "your abstract here" --max 5
python main.py gaps --query "transformer NLP"
python main.py plagiarism --text "submitted text"
python main.py format --style IEEE
python main.py pipeline --query "BERT NLP" --style APApipeline runs search → extract → graph → related → gaps → format in one pass.
Everything tunable lives in config.py — endpoints, thresholds, fusion weights, and the
shared Paper / Author / Citation / PlagiarismResult dataclasses.
Set SEMANTIC_SCHOLAR_API_KEY for higher rate limits (optional):
export SEMANTIC_SCHOLAR_API_KEY="your-key"Without network access the agent falls back to config.MOCK_PAPERS, so every command is
runnable offline for testing.
Python · sentence-transformers · scikit-learn · pdfplumber · feedparser · requests · PyTorch (CPU)
Architecture overview: research_agent/system_design_diagram.html