Skip to content
CAQ0107Public

About

Design and Implementation of an Intelligent Q&A System for Longevity and Aging Based on Retrieval-Augmented Generation

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 

Repository files navigation

Lumi

Design and Implementation of an Intelligent Q&A System for Longevity and Aging Based on Retrieval-Augmented Generation

长寿抗衰领域检索增强生成(RAG)智能问答系统

项目简介

本项目面向长寿抗衰垂直领域,基于检索增强生成(RAG)技术构建智能问答系统,解决通用大模型在专业领域知识过时、事实幻觉、证据不足等问题。系统外接PubMed、Nature Aging等权威长寿知识库,实现精准问答、引文溯源、多轮交互等核心功能,为科研人员、健康管理从业者及大众提供可溯源、低幻觉的长寿抗衰知识服务。

核心特性

  • 📚 领域专属知识库:基于生物医学顶刊/指南构建长寿抗衰标准化知识库
  • 🎯 精准检索增强:融合BM25+稠密检索+重排模型,提升专业文献匹配精度
  • 📝 引文溯源:生成回答自动关联权威文献来源,支持知识可验证
  • 🚫 幻觉抑制:通过检索阈值过滤、事实校验等机制降低生成幻觉率
  • 🔧 本地化部署:支持知识库与系统全流程本地化运行,保障数据安全
  • 💻 友好交互:基于Gradio搭建轻量化可视化交互界面

技术栈

核心框架/工具

模块 技术选型 选型说明
编程语言 Python 3.9+ 生态丰富,适配RAG开发
RAG框架 LangChain 一站式RAG流程编排
向量数据库 Milvus 高效存储/检索Embedding向量
关系型数据库 PostgreSQL(生产)/SQLite(开发) 存储文献元数据、系统配置
交互界面 Gradio 轻量化可视化交互,易部署
文本嵌入 BioBERT/Zhipu Embedding 适配生物医学领域文本特征
重排模型 ColBERT/MonoT5 提升检索结果相关性
生成模型 Qwen/Llama(开源)/商用大模型API 适配垂直领域提示工程

环境依赖

python >= 3.9
langchain >= 0.1.0
pymilvus >= 2.3.0
psycopg2-binary >= 2.9.9  # PostgreSQL依赖
gradio >= 4.0.0
transformers >= 4.35.0
sentence-transformers >= 2.2.0
pandas >= 2.1.0
numpy >= 1.24.0

快速开始 / Quick Start

1. 环境搭建 / Environment Setup

# 克隆项目 / Clone repository
git clone https://github.com/your-username/longevity-rag-qa.git
cd longevity-rag-qa

# 创建虚拟环境 / Create conda environment
conda create -n longevity-rag python=3.9
conda activate longevity-rag

# 安装依赖 / Install requirements
pip install -r requirements.txt

2. 数据库配置 / Database Configuration

2.1 Milvus 向量库部署 / Milvus Vector Store

参考 Milvus 官方文档完成单机版部署。

config/milvus_config.py:

MILVUS_HOST = "localhost"
MILVUS_PORT = 19530
COLLECTION_NAME = "longevity_knowledge_base"

2.2 关系型数据库配置 / Relational Database

  • 开发环境(SQLite):无需额外部署,直接使用 data/longevity_dev.db
  • 生产环境(PostgreSQL):

创建数据库:

CREATE DATABASE longevity_db;

config/db_config.py:

DB_TYPE = "postgresql"  # 开发时改为 sqlite
PG_CONFIG = {
    "host": "localhost",
    "port": 5432,
    "user": "your_username",
    "password": "your_password",
    "dbname": "longevity_db"
}
SQLITE_PATH = "data/longevity_dev.db"

3. 知识库构建 / Knowledge Base Build

# 数据采集(爬取/导入权威文献)/ Data collection
data_crawl.py --source pubmed --keywords "longevity anti-aging"

# 数据预处理(清洗/切块/向量化)/ Data preprocessing
python scripts/data_process.py --input data/raw_literature --output data/processed_data

# 导入向量库 / Load into Milvus
python scripts/load_to_milvus.py --data_path data/processed_data

4. 启动系统 / Start the System

# 本地启动交互界面 / Start local UI
python app/main.py

启动后访问终端输出的本地链接(默认:http://localhost:7860)。

系统架构 / System Architecture

分层设计 / Layered Architecture

  • 数据层 / Data layer:构建长寿抗衰知识图谱,完成文献数据抽取、清洗、向量化,存储至 PostgreSQL(结构化元数据)+ Milvus(向量数据)。
  • 应用层 / Application layer:基于 LangChain 设计 RAG 工作流,包含检索(BM25 + 稠密检索)、重排、生成(提示工程 + CoT)、幻觉抑制模块。
  • 展示层 / Presentation layer:Gradio 交互界面,支持自然语言提问、答案展示、引文溯源、多轮对话。

核心流程 / Core Workflow

  1. 收集文献 / Collect literature
  2. 清洗、切块、向量化 / Clean, chunk, vectorize
  3. 存储元数据与向量 / Store metadata + vectors
  4. 检索、重排、生成 / Retrieve, rerank, generate
  5. UI 展示与溯源 / UI display + provenance

About

Design and Implementation of an Intelligent Q&A System for Longevity and Aging Based on Retrieval-Augmented Generation

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors