Skip to content

About

Research agent fusing semantic similarity with co-citation and bibliographic coupling — search, gap analysis, plagiarism detection, and citation formatting

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

Graph-Augmented Intelligent Research Agent

A research assistant that treats the citation network as a first-class signal. Most related-work tools rank papers by embedding similarity alone; this one fuses semantic similarity with co-citation and bibliographic coupling drawn from a citation graph, so papers that are structurally central to a literature surface even when their abstracts read differently.

Seven modules behind one ResearchAgent class, driven from a single CLI.

Modules

# Module Responsibility
1 PaperSearchModule Search papers via Semantic Scholar and ArXiv
2 CitationExtractorModule Parse citations from a PDF or raw reference text
3 CitationGraphModule Build and query the citation knowledge graph
4 RelatedWorkModule Multi-signal recommendation (four signals fused)
5 GapAnalyzerModule Detect missing citations in a reference list
6 PlagiarismDetectorModule Citation-aware plagiarism detection
7 ReferenceFormatterModule APA / IEEE / MLA / Chicago formatting

The recommendation signal

Related-work scoring fuses four independent signals with fixed weights (config.SIGNAL_WEIGHTS):

Signal Weight Basis
Semantic 0.40 Sentence-transformer embedding cosine similarity (all-MiniLM-L6-v2)
Co-citation 0.25 How often two papers are cited together
Bibliographic coupling 0.20 How much two papers' reference lists overlap
TF-IDF 0.15 Surface keyword overlap

Graph centrality uses PageRank (α = 0.85). Pairs scoring below 0.15 similarity are discarded.

Plagiarism detection

Combines TF-IDF similarity (0.60) with Winnowing fingerprinting (0.40) over 5-character shingles in a 4-token window, then buckets the result:

Risk level Combined score
Clean 0.00 – 0.20
Borderline 0.20 – 0.40
Suspicious 0.40 – 0.60
High risk 0.60 – 1.00

It is citation-aware: correctly attributed quotations are not counted as overlap.

Usage

cd research_agent
pip install -r requirements.txt
python main.py search     --query "transformer attention" --max 5
python main.py extract    --text "References\n[1] ..."
python main.py related    --abstract "your abstract here" --max 5
python main.py gaps       --query "transformer NLP"
python main.py plagiarism --text "submitted text"
python main.py format     --style IEEE
python main.py pipeline   --query "BERT NLP" --style APA

pipeline runs search → extract → graph → related → gaps → format in one pass.

Configuration

Everything tunable lives in config.py — endpoints, thresholds, fusion weights, and the shared Paper / Author / Citation / PlagiarismResult dataclasses.

Set SEMANTIC_SCHOLAR_API_KEY for higher rate limits (optional):

export SEMANTIC_SCHOLAR_API_KEY="your-key"

Without network access the agent falls back to config.MOCK_PAPERS, so every command is runnable offline for testing.

Stack

Python · sentence-transformers · scikit-learn · pdfplumber · feedparser · requests · PyTorch (CPU)

Architecture overview: research_agent/system_design_diagram.html

About

Research agent fusing semantic similarity with co-citation and bibliographic coupling — search, gap analysis, plagiarism detection, and citation formatting

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages