Table of Contents
The Vaccine Adverse Event Reporting System (VAERS) contains detailed reports of adverse events following vaccine administration. However, efficiently and accurately searching for specific infor- mation from VAERS poses significant challenges, especially for medical experts. Natural language querying (NLQ) methods tackle the challenge by translating the input questions into executable queries, allowing for the exploration of complex databases with large amounts of information. Most existing studies focus on the relational database and solve the Text-to-SQL task. However, the capability of full-text for Text-to-SQL is greatly limited by the data structures and functionality of the SQL databases. In addition, the potential of natural language querying has not been comprehen- sively explored in the healthcare domain. To overcome these limita- tions, we investigate the potential of NoSQL databases, specifically Elasticsearch, and forge a new research direction for NLQ, which we refer to as Text-to-ESQ generation. This exploration requires us to re-design various aspects of NLQ, such as the target applica- tion and the advantages of NoSQL database. In our approach, we develop a two-stage controllable (TSC) framework consisting of a question-to-question (Q2Q) translation module and an ESQ condition extraction (ECE) module. These modules are carefully designed to efficiently retrieve information from the VEARS data stored in a NoSQL database. Additionally, we construct a dedicated question-ESQ pair dataset called VAERSESQ, to support the task in the healthcare domain. Extensive experiments were conducted on the VAERSESQ dataset to evaluate the proposed methods. The results, both quantitative and qualitative, demonstrate the accuracy and efficiency of our approach in generating queries for NoSQL databases, thus enabling efficient retrieval of VEARS data.
Use the BLANK_README.md to get started.
All the experiments were performed using NVIDIA Quadro RTX 5000 GPUs. The proposed TSC model is implemented with PyTorch. We adopt the SGD with momentum optimizer during the training of the model parameters. The learning rate is set to 0.01. The experiments for all the models are obtained by running 16 epochs with the mini-batch size 32. The development set is used to select the best model.
- M2M
- Bart
- LSTM
- RoBERTa
- DistilBERT
- RoBERTa+Bi-LSTM
This is an example of how you may give instructions on setting up your project locally. To get a local copy up and running follow these simple example steps.
All the experiments were performed using NVIDIA Quadro RTX 5000 GPUs. The proposed TSC model is implemented with PyTorch. We adopt the SGD with momentum optimizer during the training of the model parameters. The learning rate is set to 0.01. The experiments for all the models are obtained by running 16 epochs with the mini-batch size 32. The development set is used to select the best model.
Below is an example of how you can instruct your audience on installing and setting up your app. This template doesn't rely on any external dependencies or services.
- Get a free API Key at https://example.com
- Clone the repo
git clone https://github.com/your_username_/Project-Name.git
- Install NPM packages
npm install
- Enter your API in
config.jsconst API_KEY = 'ENTER YOUR API';
Use this space to show useful examples of how a project can be used. Additional screenshots, code examples and demos work well in this space. You may also link to more resources.
For more examples, please refer to the Documentation
- Add Changelog
- Add back to top links
- Add Additional Templates w/ Examples
- Add "components" document to easily copy & paste sections of the readme
- Multi-language Support
- Chinese
- Spanish
See the open issues for a full list of proposed features (and known issues).
Our major contributions can be summarized as follows
- Formally propose and formulate the Text-to-ESQ task to support NLQ on NoSQL database. To the best of our knowledge, this is the initial comprehensive investigation of NLQ on the NoSQL database
- Propose a two-stage controllable (TSC) framework consisting of two modules for Text-to-ESQ: (1) Question-to-question transla- tion module for translating natural language questions into the corresponding template questions, and (2) ESQ condition extrac- tion module for parsing name entities about condition fields and values from template questions for further populating the query templates
- Create a large-scale dataset VAERSESQ for Text-to-ESQ task for retrieving information from VAERS data. For each question, we include both template-based and natural language forms
- Conduct an extensive experimental analysis of the VAERSESQ dataset and demonstrate the effectiveness of the proposed two- stage controllable method
Distributed under the MIT License. See LICENSE.txt for more information.
Ping Wang - pwang44@stevens.edu
Wenlong Zhang - wzhang71@stevens.edu
Project Link: (https://github.com/LEAF-Lab-Stevens/Text2ESQ)
Use this space to list resources you find helpful and would like to give credit to. I've included a few of my favorites to kick things off!